Image decoding device, image encoding device, image decoding method, and image encoding method

The image decoding apparatus ensures compatibility and extensibility by decoding a specific profile indicator and additional syntax elements, resolving challenges in extending image encoding profiles.

JP7712801B2Active Publication Date: 2025-07-24SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021108681
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-07-24
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Existing image encoding and decoding technologies face challenges in maintaining backward compatibility while extending profiles, and there is a need for forward compatibility in new profile definitions, particularly in the context of VVC's profile_tier_level() syntax.

Method used

An image decoding apparatus that decodes a general_profile_idc and, when it is a specific value, further decodes a syntax element indicating a restricted syntax length and a restricted syntax group from the encoded data, ensuring compatibility and future extensibility.

Benefits of technology

This approach effectively addresses the issues of backward and forward compatibility, enabling seamless integration of new profiles without disrupting existing operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712801000001
    Figure 0007712801000001
  • Figure 0007712801000002
    Figure 0007712801000002
  • Figure 0007712801000003
    Figure 0007712801000003
Patent Text Reader

Abstract

To provide an image decoding device and an image coding device having a bit stream structure which performs tool restriction while having interchangeability.SOLUTION: The image decoding device comprises a header decoding unit which decodes general_profile_idc indicating a profile from coded data. When the general_profile_idc has a specific value, the header decoding unit further decodes from the coded data a syntax element indicating a restriction syntax length and a restriction syntax group of a length equal to or longer than a length indicated by the syntax element indicating the restriction syntax length.SELECTED DRAWING: Figure 21
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an image decoding apparatus and an image encoding apparatus.

Background Art

[0002] To efficiently transmit or record an image, an image encoding apparatus that generates encoded data by encoding the image, and an image decoding apparatus that generates a decoded image by decoding the encoded data are used.

[0003] Specific image encoding methods include, for example, H.264 / AVC and HEVC (High-Efficiency Video Coding) methods and the like.

[0004] In such an image encoding method, an image (picture) constituting a moving image is managed by a hierarchical structure including a slice obtained by dividing the image, a coding tree unit (CTU: Coding Tree Unit) obtained by dividing the slice, a coding unit (sometimes called a Coding Unit: CU) obtained by dividing the coding tree unit, and a transform unit (TU: Transform Unit) obtained by dividing the coding unit, and is encoded / decoded for each CU. and, Also, in such an image encoding method, usually, a prediction image is generated based on a local decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Examples of the method for generating the prediction image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).

[0005] In addition, Non-Patent Document 1 (VVC) can be cited as a recent image encoding and decoding technology. Non-Pat

[0006] ​ Non-Patent Document 1 discloses conversion accuracy and Rice parameter extension techniques.

[0007] Non-Patent Document 2 describes techniques for defining new profiles in HEVC, known as Format Range Extension profiles (RExt). and is known for defining new profiles.

Prior Art Documents

Non-Patent Documents

[0008]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0009] Since Non-Patent Document 1 is an extension technique, it is necessary to define it to be backward compatible so as not to affect the operation in existing profiles and to be available in new profiles. Non-Patent Document 2 defines the definition of extended profiles in the syntax of HEVC's profile_tier_level(), but in the syntax of VVC's profile_tier_level() there are no appropriate reserved bits defined, so there is a problem that it is difficult to extend while maintaining backward compatibility. Also, for new profile definitions, a configuration considering future extensions (forward compatibility) is necessary.

Means for Solving the Problems

[0010] To solve the above problems, an image decoding apparatus according to one aspect of the present invention includes: a header decoding unit that decodes a general_profile_idc indicating a profile from encoded data and, when the general_profile_idc is a specific value, further decodes a syntax element indicating a restricted syntax length and a restricted syntax group having a length equal to or greater than the length indicated by the syntax element indicating the restricted syntax length from the encoded data. It is characterized by the above.

Advantages of the Invention

[0011] According to the above configuration, any of the above problems can be solved.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Mode for Carrying Out the Invention

[0013] 〔Embodiment 1〕 Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0014] FIG. 1 is a schematic diagram showing the configuration of an image transmission system 1 according to the present embodiment.

[0015] The image transmission system 1 is a system that transmits an encoded stream obtained by encoding an image to be encoded and decodes the transmitted encoded stream to display an image. The image transmission system 1 includes a moving image encoding device (image encoding device) 11, a network 21, a moving image decoding device (image decoding device) 31, and an image display device (image display device) 41.

[0016] An image T is input to the moving image encoding device 11.

[0017] The network 21 transmits the encoded stream Te generated by the moving image encoding device 11 to the moving image decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Further, the network 21 may be replaced by a storage medium that records the encoded stream Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).

[0018] The moving image decoding device 31 decodes each of the encoded streams Te transmitted by the network 21 and generates one or more decoded images Td obtained by decoding.

[0019] The image display device 41 displays all or part of the one or more decoded images Td generated by the moving image decoding device 31. The image display device 41 includes, for example, a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the form of the display include a stationary type, a mobile type, and an HMD. Further, when the moving image decoding device 31 has high processing power, an image with high image quality is displayed, and when it has only low processing power, an image that does not require high processing power and display ability is displayed.

[0020] <Operator> The operators used in this specification are described below.

[0021] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR. |= is an OR assignment operator, and || represents a logical OR.

[0022] x?y:z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).

[0023] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive), returning a if c < a, b if c > b, and c otherwise (where a <= b).

[0024] abs(a) is a function that returns the absolute value of a.

[0025] Int(a) is a function that returns the integer value of a.

[0026] floor(a) is a function that returns the smallest integer less than or equal to a.

[0027] ceil(a) is a function that returns the largest integer greater than or equal to a.

[0028] a / d represents the division of a by d (truncating the fractional part).

[0029] <Structure of the encoded stream Te> Prior to the detailed description of the moving image encoding device 11 and the moving image decoding device 31 according to this embodiment, the data structure of the encoded stream Te generated by the moving image encoding device 11 and decoded by the moving image decoding device 31 will be described.

[0030] Figure 4 is a diagram showing the hierarchical structure of the data in the encoded stream Te. The encoded stream Frame Te includes, by way of example, a sequence and a plurality of pictures that make up the sequence. FIG. 4 shows, respectively, an encoded video sequence that defines sequence SEQ, an encoded picture that defines picture PICT, an encoded slice that defines slice S, encoded slice data that defines slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit. In the encoded video sequence, a set of data that the moving image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in the encoded video sequence of FIG. 4, sequence SEQ includes a Video Parameter Set, a Sequence Parameter Set SPS, a Picture Parameter Set PPS, a picture header, picture PICT, and Supplemental Enhancement Information SEI.

[0031] (Encoded video sequence) In the encoded video sequence, a set of data that the moving image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in the encoded video sequence of FIG. 4, sequence SEQ includes a Video Parameter Set, a Sequence Parameter Set SPS, a Picture Parameter Set PPS, a picture header, picture PICT, and Supplemental Enhancement Information SEI. The Video Parameter Set VPS defines a set of encoding parameters common to a plurality of images, a plurality of layers included in the image, and a set of encoding parameters related to individual layers in an image composed of a plurality of layers.

[0032] In the Sequence Parameter Set SPS, a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the target sequence is defined. For example, the width and height of a picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS. The Sequence Parameter Set SPS defines a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS.

[0033] In the Sequence Parameter Set SPS, a set of encoding parameters that the moving image decoding device 31 refers to in order to decode the target sequence is defined. For example, the width and height of a picture are defined. Note that there may be a plurality of SPSs. In that case, one of the plurality of SPSs is selected from the PPS.

[0034] ​In the picture parameter set PPS, a set of encoding parameters that the moving picture decoding device 31 refers to in order to decode each picture in the target sequence is defined. For example, it includes a reference value of the quantization width (pic_init_qp_minus26) used for decoding a picture, a flag (weighted_pred_flag) indicating the application of weighted prediction, and a scaling list (quantization matrix). Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence.

[0035] In the picture header, encoding parameters common to all slices included in one encoded picture are defined. For example, it includes POC (Picture Order Count) and encoding parameters related to partitioning.

[0036] (Encoded Picture) In the encoded picture, a set of data that the moving picture decoding device 31 refers to in order to decode the picture PICT to be processed is defined. The picture PICT includes slices 0 to NS - 1 as shown in the encoded picture of FIG. 4 (NS is the total number of slices included in the picture PICT).

[0037] Note that hereinafter, when there is no need to distinguish each of slices 0 to NS - 1, the subscript of the symbol may be omitted in the description. The same applies to the data included in the encoding stream Te described below and other data with subscripts.

[0038] (Encoded Slice) In the encoded slice, a set of data that the moving picture decoding device 31 refers to in order to decode the slice S to be processed is defined. As shown in the encoded slice of FIG. 4, the slice includes a slice header and slice data.

[0039] ​​​​​ The slice header includes a set of encoding parameters that the moving image decoding device 31 refers to in order to determine the decoding method of the target slice. The slice type specification information (slice_type) that specifies the slice type is an example of the encoding parameters included in the slice header.

[0040] Examples of slice types that can be specified by the slice type specification information include: (1) an I slice that uses only intra prediction during encoding; (2) a P slice that uses either uni-directional prediction or intra prediction during encoding; (3) a B slice that uses uni-directional prediction, bi-directional prediction, or intra prediction during encoding, etc. Note that inter prediction is not limited to uni-prediction and bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it refers to a slice including a block that can use inter prediction.

[0041] Note that the slice header may include a reference (pic_parameter_set_id) to the picture parameter set PPS.

[0042] (Encoded slice data) The encoded slice data defines a set of data that the moving image decoding device 31 refers to in order to decode the slice data to be processed. The slice data includes CTUs as shown in the encoded slice header of FIG. 4. As shown in the slice header, the CTU is included. The CTU is a block of a fixed size (for example, 64x64) that constitutes a slice, and is sometimes also referred to as the largest coding unit (LCU).

[0043] (Coding tree unit) The coding tree unit in FIG. 4 defines a set of data that the moving image decoding device 31 refers to in order to decode the target CTU. The CTU is recursively quad-tree divided (QT (Quad Tree) division), binary-tree divided (BT (Binary Tree) division), or ternary-tree divided (TT (Ternary Tree) The BT and TT partitions are collectively called multi-tree partitions (MT (Multi Tree) partitions). The nodes of the tree structure obtained by partitioning are called coding nodes. The intermediate nodes of a quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the top coding node.

[0044] A CU is a terminal node of a coding node and is not divided any further. A CU is the basic unit of coding processing.

[0045] (Encoding Unit) As shown in the coding unit of FIG. 4, a video is generated to decode the coding unit to be processed. The CU defines a set of data to be referenced by the image decoding device 31. Specifically, a CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header defines a prediction mode, etc.

[0046] The prediction process may be performed in units of CUs, or in units of sub-CUs that are obtained by further dividing a CU. When the sizes of a CU and a sub-CU are equal, there is one sub-CU in a CU. If the size of the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, if the CU is 8x8 and the sub-CU is 4x4, the CU is divided into 2 horizontal divisions and 2 vertical divisions, into 4 sub-CUs.

[0047] The prediction type (prediction mode CuPredMode) includes at least two types: intra prediction (MODE_INTRA) and inter prediction (MODE_INTER). In addition, intra block copy prediction (MODE_IBC) Intra prediction and intra block copy prediction are predictions within the same picture, and inter prediction refers to prediction processing performed between different pictures (for example, between display times or between layer images).

[0048] The conversion and quantization processes are performed in units of CUs, but the quantized conversion coefficients may be entropy-coded in units of sub-blocks such as 4x4.

[0049] The predicted image is derived by prediction parameters associated with the block. The prediction parameters include prediction parameters for intra prediction and inter prediction.

[0050] (Configuration of the moving image decoding device) The configuration of the moving image decoding device 31 (Fig. 6) according to this embodiment will be described.

[0051] The moving image decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (predicted image decoding device ) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation device) 308, an inverse quantization and inverse transformation unit 311 (scaling unit), and an addition unit 312. Note that, in accordance with the moving image encoding device 11 described later, there is also a configuration in which the moving image decoding device 31 does not include the loop filter 305.

[0052] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS, and slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes CT from the encoded data. The CU decoding unit 3022 decodes CU from the encoded data. The TU decoding unit 3024 decodes the transform skip flag, QP update information (quantization correction value), and quantized conversion coefficients (residual coding) from the encoded data.

[0053] ​The header decoding unit 3020 decodes the sps_lfnst_enabled_flag, a flag indicating whether to use non-separable transformation from the SPS. Also, when the sps_lfnst_enabled_flag is 1 the header decoding unit 3020 decodes the ph_lfnst_enabled_flag from the picture header (PH). When the ph_lfnst_enabled_flag does not appear, the ph_lfnst_enabled_flag is inferred to be 0. Alternatively, when the ph_lfnst_enabled_flag does not appear, the value of the sps_lfnst_enabled_flag may be set as the value of the ph_lfnst_enabled_flag.

[0054] The TU decoding unit 3024 decodes the index mts_idx indicating the transform basis from the encoded data. Also, the TU decoding unit 3024 decodes the parameter lfnst_idx indicating the presence or absence of use of non-separable transformation and the transform basis from the encoded data. Specifically, the TU decoding unit 3024 decodes the lfnst_idx when the width and height of the CU are 4 or more and the prediction mode is the intra prediction mode. Note that when lfnst_idx is 0, it indicates non-application of non-separable transformation, and when it is 1, it indicates one of the sets (pairs) of non-separable transformation matrices (transform bases), and when it is 2, it indicates the other transformation matrix of the above pair.

[0055] The TU decoding unit 3024 decodes the transform_skip_flag[x0][y0][cIdx] when the size (tbWidth and tbHeight) of the transform unit is less than or equal to a predetermined maximum size (tbWidth <= MaxTsSize && tbHeight <= MaxTsSize).

[0056] The TU decoding unit 3024, when the TU contains prediction error (e.g., tu_cbf_luma[x0][y0] is ​​​In case 1), QP update information and quantization conversion coefficients are decoded from the encoded data. Multiple modes (e.g., RRC mode and TSRC mode) may be provided in the derivation of the quantization conversion coefficients. Specifically , the derivation of the normal prediction error using transformation (RRC: Regular Residual Coding) and the derivation of the prediction error in the transform skip mode without using transformation (TSRC: Transform Skip Residual Coding) may perform different processes. The QP update information is a difference value from the quantization parameter predicted value qPpred which is a predicted value of the quantization parameter QP. Also, hereinafter, an example using CTU and CU as processing units will be described, but this is not limited to this example

[0057] , and processing may be performed in units of sub-CUs. Alternatively, CTU and CU may be read as blocks, and sub-CU may be read as sub-blocks , and processing may be performed in units of blocks or sub-blocks. The entropy decoding unit 301 performs entropy decoding on the encoded stream Te input from the outside and parses individual codes (syntax elements). For entropy coding, there are a method of variable-length coding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and the surrounding situation, and a method of variable-length coding of syntax elements using a predetermined table or calculation formula. An example of the former is CABAC (Context Adaptive Binary Arithmetic Coding). The parsed codes include prediction information for generating a predicted image, a prediction error for generating a difference image, and the like.

[0058] The entropy decoding unit 301 outputs the separated codes to the parameter decoding unit 302. The separated codes are, for example, the prediction mode CuPredMode. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302. The entropy decoding unit 301 outputs the separated codes to the parameter decoding unit 302. The separated codes are, for example, the prediction mode CuPredMode. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.

[0059] The entropy decoding unit 301 outputs the separated codes to the parameter decoding unit 302. The separated codes are, for example, the prediction mode CuPredMode. The control of which code to decode is performed based on the instruction of the parameter decoding unit 302.

[0060] (Basic Flow) FIG. 7 is a flowchart for explaining the schematic operation of the moving image decoding apparatus 31.

[0061] (S1100: Decoding of Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, SEI, and PH from the encoded data.

[0062] (S1200: Decoding of Slice Information) The header decoding unit 3020 decodes a slice header (slice information) from the encoded data.

[0063] Hereinafter, the moving image decoding apparatus 31 derives a decoded image of each CTU by repeating the processes from S1300 to S5000 for each CTU included in the target picture.

[0064] (S1300: Decoding of CTU Information) The CT information decoding unit 3021 decodes a CTU from the encoded data.

[0065] (S1400: Decoding of CT Information) The CT information decoding unit 3021 decodes a CT from the encoded data.

[0066] (S1500: Decoding of CU) The CU decoding unit 3022 performs S1510 and S1520 to decode a CU from the encoded data. Also, the CU decoding unit 3022 decodes a quantization parameter difference CuQpDeltaVal in units of CU from the encoded data and derives a quantization parameter.

[0067] (S1510: Decoding of CU Information) The CU decoding unit 3022 decodes CU information, prediction information, TU split (S1510: Decoding of CU Information) The CU decoding unit 3022 decodes CU information, prediction information, TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.

[0068] (S1520: TU Information Decoding) When the TU contains prediction errors, the TU decoder 3024 decodes the QP update information and quantization transform coefficients from the encoded data.

[0069] (S2000: Predicted Image Generation) The predicted image generation unit 308 generates a predicted image for each block included in the target CU based on the prediction information.

[0070] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation processing for each TU included in the target CU.

[0071] (S4000: Decoded Image Generation) The addition unit 312 adds the predicted image supplied from the predicted image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transformation unit 311 to generate the decoded image of the target CU.

[0072] (S5000: Loop Filter) The loop filter 305 applies loop filters such as a deblocking filter, SAO (Sample Adaptive Filter), and ALF (Adaptive Loop Filter) to the decoded image to generate the decoded image.

[0073] (Derivation of Quantization Transform Coefficients, Residual Encoding) In the case of lossless encoding or when the pixel correlation of the original image is small, it may be more efficient in terms of encoding efficiency not to perform the transformation. The technique of not performing the transformation is called transform skip. Transform skip is also called Identical Transform and only performs scaling of the transform coefficients according to the quantization parameter. Whether it is transform skip is notified using the syntax element transform_skip_flag. transform_skip_flag may be notified for each color component (cIdx) of Y, Cb, and Cr.

[0074] Derivation of normal prediction error using transformation (RRC: Regular Residual Coding) and derivation of prediction error in the transformation skip mode (TSRC: Transform Skip Residual Coding) differ in both the prediction error encoding method and the decoding method.

[0075] FIG. 8 is a block diagram of the TU decoding unit 3024, which includes an RRC unit 30241 and a TSRC unit 30242. The RRC unit 30241 is a processing unit that derives a normal prediction error using transformation, and the TSRC unit 30242 is a processing unit that derives a prediction error in the transformation skip mode.

[0076] The transform_skip_flag[x0][y0][cIdx] in FIG. 10 indicates whether transformation is applied to the block with the upper left coordinates (x0, y0) and the color component cIdx. When transform_skip_flag = 1 (transformation skip mode), transformation is not applied to this block. When transform_skip_flag = 0, whether transformation is applied to this block depends on other parameters.

[0077] FIGS. 11 to 13 and FIGS. 14 to 15 are each syntax tables showing the structure (encoding method) of the encoded data of the transformation coefficients (prediction errors) in RRC (Regular Residual Coding) and TSRC (Transform Skip Residual Coding).

[0078] (RRC unit, RRC mode) In the normal prediction error coding method without conversion skipping (from FIGS. 11 to 13), the RRC unit 30241 decodes a syntax element (not shown) indicating the LAST position and derives the LAST position (LastSignificantCoeffX, LastSignificantCoeffY). The LAST position is the position of the last non-zero coefficient when scanning the transform coefficients of the TU from the low-frequency component to the high-frequency component direction. When encoding and decoding the transform coefficients (coefficients) in order from the high-frequency component, the LAST position indicates the position of the quantized transform coefficient to be decoded first. Next, the RRC unit 30241 decodes sb_coded_flag with reference to the LAST position. sb_coded_flag is a flag indicating whether the sub-block contains non-zero coefficients. The sub-block is an area obtained by dividing the TU into 4x4 units. If sb_coded_flag = 1 (the sub-block contains non-zero coefficients), the RRC unit 30241 decodes sig_coeff_flag. sig_coeff_flag is a flag indicating whether the coefficient value is non-zero. If sig_coeff_flag = 1 (the coefficient value is non-zero), the RRC unit 30241 decodes abs_level_gtx_flag, par_level_flag, abs_remainder, and dec_abs_level. These are syntax elements indicating the absolute value of the coefficient. abs_level_gtx_flag[n][j] is a flag indicating whether the absolute value of the coefficient at the scan position n is greater than (j << 1)+1. If abs_level_gtx_flag[n][j] is not notified, it is assumed to be 0. par_level_flag[n] is the parity of the coefficient at the scan position n. If par_level_flag[n][j] is not notified, it is assumed to be 0. abs_remainder[n] represents the remainder of the absolute value of the coefficient at the scan position n and is decoded by the Golomb-Rice code. If abs_remainder[n] is not notified, it is assumed to be 0. dec_abs_level[n] is the absolute value of the residual for deriving the absolute value of the coefficient at the scan position n. If dec_abs_level[n] is not notified, it is assumed to be 0.The RRC unit 30241 derives the absolute value of the coefficient from these syntax elements.

[0079] As described above, the RRC unit 30241 is characterized by decoding the LAST position in the sub-block of the TU.

[0080] (TSRC unit, TSRC mode) Prediction error coding method in the case of not performing conversion (conversion skip mode) (FIG. 14, FIG. 15, FIG. 9) In this case, the TSRC unit 30242 decodes the sb_coded_flag of each sub-block. If sb_coded_flag = 1 (the sub-block contains non-zero coefficients), the TSRC unit 30242 decodes the sig_coeff_flag[xC][yC] of the transform coefficient at the position (xC, yC) within the sub-block. If sig_coeff_flag = 1 (the coefficient value is non-zero), the TSRC unit 30242 decodes the coeff_sign_flag, abs_level_gtx_flag, par_level_flag, and abs_remainder. These are the syntax elements indicating the absolute value of the coefficient, and the definitions are as described above. The TSRC unit 30242 derives the absolute value of the coefficient from these syntax elements and outputs it.

[0081] As described above, the TSRC unit 30242 is characterized by not decoding the LAST position in the sub-block of the TU and outputs it.

[0082] (Decoding of abs_remainder and dec_abs_level) The TU decoding unit 3024 decodes the syntax value of the absolute value of the residual (abs_remainder or dec_abs_level, hereinafter referred to as the residual) from the encoded data using the Rice parameter cRiceParam. For the binaryization of the syntax, a code composed of a prefix and a suffix is used. As the prefix and the suffix, the alpha code of the Golomb-Rice code and the fixed-length binaryization may be used. The suffix may use the EG(k) (Exponential-Golomb code). The "absolute value of the residual of the transform coefficient (absolute residual value)", which is encoded as a syntax element, is simply expressed as the "residual".

[0083] In this embodiment, as the Golomb-Rice code, the Truncated Rice (TR) code is used, and an example will be described in which the prefix is divided into prefixValTR and suffixValTR for encoding, and in the EG(k) code, the suffix is divided into the exp part expVal and the escape part escapeVal for encoding. The quotients obtained by dividing the residual (abs_remainder or dec_abs_level) by 1<<cRiceParam and (((1<<expVal))-1)<<cRiceParam are encoded using prefixValTR and expVal, respectively. The remainders obtained by dividing the residual (abs_remainder or dec_abs_level) by 1<<cRiceParam and (1<<expVal<<cRiceParam) are encoded using suffixValTR and escapeVal, respectively. That is, the residual is encoded as follows.

[0084] ((prefixValTR + ((1<<expVal-1))) << cRiceParam) + suffixValTR + escapeVal When the residual is less than or equal to a predetermined value cMax, only prefixValTR and suffixValTR are encoded.

[0085] prefixVal = (prefixValTR << cRiceParam) + suffixValTR When the residual is greater than a predetermined value (when suffixVal exists), in addition to prefixValTR indicating prefixVal equal to cMax, expVal and escapeVal are encoded.

[0086] prefixVal = (prefixValTR << cRiceParam) = cMax suffixVal = (((1 << expVal) - 1) << cRiceParam) + escapeVal Here, the residual can be expressed as follows. The residual when it can be represented by only the prefix is given by the following formula.

[0087] Residual = ((prefixValTR) << cRiceParam) + suffixValTR Otherwise (when it cannot be represented by only the prefix, or when prefixValTR = 6), the residual is given by the following formula.

[0088] Residual = ((prefixValTR + (1 << (expVal) - 1)) << cRiceParam) + escapeVal Also, prefixVal and suffixVal may be defined as follows.

[0089] prefixVal = (prefixValTR << cRiceParam) + suffixValTR suffixVal = (((2 << expVal) - 2) << cRiceParam) + escapeVal (Flow) The TU decoding unit 3024 first derives the maximum value cMax of the prefix prefixVal from cRiceParam to do.

[0090] cMax = 6 << cRiceParam Here, 6 is the maximum value of prefixValTR, also called maxPrefixValTR.

[0091] The TU decoder 3024 decodes the prefix value (prefixVal) from the encoded data based on the binaryization of the Truncated Rice (TR) code described below. Note that the relationship between the prefix and the remainder is as follows. The relationship between the difference and the remainder is as follows.

[0092] prefixVal = Min(cMax, abs_remainder[n]) Alternatively, the following may also be used.

[0093] prefixVal = Min(cMax, dec_abs_level) Subsequently, when the remainder is greater than cMax (when the remainder cannot be represented by prefixVal alone), the TU decoder 3024 further decodes suffixVal that satisfies the following relational expression.

[0094] suffixVal = abs_remainder[n] - cMax Alternatively, the following may also be used.

[0095] suffixVal = dec_abs_level - cMax Specifically, when suffixVal exists, the TU decoder 3024 decodes suffixVal from the encoded data based on the binaryization of Limited k-th order Exp-Golomb (EG(k)). The case where suffixVal exists is when the bit sequence of the prefixVal of the TR code has a length of 6 (= maxPrefixValTR), that is, "111111". Here, let the order k be cRiceParam, the variable maxPreExtLen be 11, and the variable truncSuffixLen be 15. maxPreExtLen may not be a fixed value (11) and may be derived according to the range of the conversion coefficient log2TransformRange.

[0096] In this embodiment, for example, maxPreExtLen may be derived by the following formula.

[0097] maxPreExtLen = (32 - 6) - log2TransformRange = 26 - log2TransformRange For example, when log2TransformRange is as in the following (Equation R-1), it may be derived as follows.

[0098] maxPreExtLen = extended_precision_processing_flag? 26 - Max(15, BitDepth + BDOFFSET) : 11 Here, maxPreExtLen and escapeLength respectively indicate the maximum values of the exp part and the escape part in the EG(k) code. log2TransformRange is a variable indicating the range of coefficients in the conversion processed by the inverse quantization / inverse transform unit 311. in the coefficient.

[0099] (Binaryization of Truncated Rice (TR) Code) The input symbolVal of the TR binaryization is prefixVal. In the TR binaryization, the quotient of symbolVal (==prefixVal) divided by (1<<cRiceParam) is prefixValTR, and the remainder is suffixValTR, which are then encoded. The relationship between the input symbolVal of the TR code and prefixValTR is as follows.

[0100] prefixValTR = symbolVal>>cRiceParam prefixValTR is converted into a bit string according to the following definition. prefixValTR < (cMax>>cRiceParam) In the case of, it is a bit string indexed by binIdx of length prefixValTR + 1. When binIdx < prefixValTR, the bins are "1", when binIdx == prefixValTR, the bin is "0". Otherwise (when prefixValTR == (cMax >> cRiceParam), it is a column of "1"s of length cMax >> cRiceParam. Note that a bin is a bit string consisting of "0" or "1". binIdx is a value indicating the position from the beginning of the bit string, and binIdx = 0 corresponds to the beginning (the left end of the binary string). The following shows examples of bit strings (bins) when prefixValTR is from 0 to 5.

[0101] prefixValTR == 0 bins = 0 prefixValTR == 1 bins = 10 prefixValTR == 2 bins = 110 prefixValTR == 3 bins = 1110 prefixValTR == 4 bins = 11110 prefixValTR == 5 bins = 111110 … However, when symbolVal >= cMax (i.e., prefixVal == (6 << cRiceParam)), the following formula applies.

[0102] prefixValTR == 6 bins = 111111 The bins of prefixValTR of the TR code are truncated alpha codes (unary codes). Here, being truncated means omitting the last 0 when it matches the maximum value.

[0103] When cMax is greater than symbolVal (symbolVal < cMax) and cRiceParam > 0, there is a suffixTR in the TR bin string, which is derived as follows. suffixValTR is the suffix of the TR bit string.

[0104] suffixValTR = symbolVal - (prefixValTR< <cRiceParam) In addition, the binarization of the suffix of the TR bit string is performed by setting the maximum value cMaxFL (1< <cRiceParam)-1とする、固定長バイナリゼーション(後述)を用いる。

[0105] In addition, when symbolVal==cMax, the suffix of the residual is the same as the above TR binary quantization. Use the binarization of the EG(k) code instead of the suffixValTR of the

[0106] (Binarization of Limited k-th order Exp-Golomb / EG(k) codes) The residual suffixVal may use the binarization of the EG(k) code. The EG(k) code is It is a binarization obtained by inputting the number k (=cRiceParam) and maxPreExtLen, and is shown in the following pseudo code. The EG(k) code consists of an exp part with a length of maxPreExtLen or less, and an escape part of the fixed-length binarization with a length of escapeLength. preExtLen is a variable that counts the number of put(1)s in the exp part, and is equal to the value espVal coded in the exp part. In other words, expVal=preExtLen. The corresponding residual value is expVal< <cRiceParamである。escapeValはescape部で符号化される値である。 / / exp part codeValue = symbolVal >> k preExtLen = 0 while ((preExtLen<maxPreExtLen) && (codeValue > ((2< <preExtLen) - 2)))) { preExtLen++ put(1) } if(preExtLen == maxPreExtLen) escapeLength = truncSuffixLen else { escapeLength = preExtLen + k put(0) } / / escape part symbolVal = symbolVal - (((1<<preExtLen)-1)<<k) while((escapeLength--)>0) put((symbolVal>>escapeLength) & 1) Here, symbolVal is the input value, which is the value obtained by subtracting prefixVal (== cMax) from the residual. put(X) indicates an operation of appending X ("0" or "1") to the end of the bit sequence to create the bit sequence. The initial value of the bit sequence is empty ({}). escapeLength is a variable indicating the bit position output from the escape part in symbolVal. Note that the escape part may be as follows.

[0107] symbolVal = symbolVal - (((2<<preExtLen)-2)<<k) while((escapeLength--)>0) put((symbolVal>>escapeLength) & 1) (Fixed - length binary serialization) Fixed - length binary serialization is performed for the maximum value cMaxFL and the value symbolVal using an unsigned integer sequence (binary sequence) of length fixedLength of Ceil(Log2(cMaxFL + 1)). Note that 0 at position binIdx corresponds to the most significant bit (MSB), and as binIdx increases, it approaches the least significant bit (LSB).

[0108] (Summary of residual decoding) The TU decoding unit 3024 decodes the residual composed of prefixVal and suffixVal.

[0109] The TU decoder 3024 derives cMax from cRiceParam, decodes prefixVal from the encoded data with cMax as the upper limit, and decodes suffixVal if it exists.

[0110] Specifically, the TU decoder 3024 decodes the prefixValTR of the TR binaryization. When symbolVal < cMax, cRiceParam > 0 (i.e., prefixVal < cMax, prefixVal < 6), it further decodes the suffixValTR of the fixed-length binaryization from the encoded data. In this case, there is no suffixVal.

[0111] When the above suffixVal does not exist, the TU decoder 3024 derives abs_remainder[n] and dec_abs_level using the following formula.

[0112] abs_remainder[n] = (prefixValTR << cRiceParam) + suffixValTR dec_abs_level[n]= (prefixValTR << cRiceParam) + suffixValTR In other cases (when prefixVal == cMax, prefixValTR == 6), the TU decoder 3024 decodes suffixVal using the EG(k) code and derives abs_remainder[n] and dec_abs_level[n] using the following formula.

[0113] abs_remainder[n] = cMax + suffixVal dec_abs_level[n] = cMax + suffixVal Here, when suffixVal consists of the preExtLen of the exp part and the escapeVal of the escape part of the EG(k) code the following holds.

[0114] suffixVal = (preExtLen << cRiceParam) + escapeVal To summarize, the TU decoding unit 3024 decodes the prefix of the TR binaryization from the encoded data prefixVal, and the suffixVal consisting of preExtLen of the exp part and escapeVal of the escape part of the EG(k) code, and derives the residual (abs_remainder[n] or dec_abs_level[n]).

[0115] abs_remainder[n] = prefixVal + suffixVal = (prefixValTR << cRiceParam) + (prefixValTR < 6)? suffixValTR : (1 << expVal << cRiceParam) + escapeVal Here, when preExtLen = escapeVal = 0 in the case of (prefixValTR < 6), abs_remainder[n] = (prefixValTR + preExtLen) << cRiceParam + (prefixValTR < 6)? suffixValTR : escapeVal The derivation of dec_abs_level[n] is the same. Note that since the code obtained by concatenating the bins of predfixValTR of the truncated alpha code and the bins of predExtLen of the alpha code is also an alpha code, the prefixVal of the TR code and the exp part of the EG(k) code may be decoded as a single entity. In this case, an alpha code with a maximum length of maxPreExtLen + 6 is read from the encoded data, and the length prefixtmp (= prefixValTR + expVal) of the alpha code is decoded. Here, maxPreExtLen + 6 is the sum of the maximum length 6 of prefixValTR and the maximum length maxPreExtLen of the exp part of the EG(k) code.

[0116] The TU decoder unit 3024 may also derive the residual as follows by dividing 1 << cRiceParam into the quotient and the remainder. This is acceptable.

[0117] Residual = (prefixRes << cRiceParam) + suffixRes Assuming that the length of the decoded alpha code is prefixtmp, the exponent part prefixRes of the residual is expressed by the following formula.

[0118] prefixRes = (prefixVal < 6)? prefixValTR : prefixValTR + (1 << expVal) - 1 When prefixtmp < 6, prefixValTR = prefixtmp; when prefixtmp >= 6, prefixValTR = 6 and expVal = prefixtmp - 6. Therefore, it can be derived as follows.

[0119] prefixRes = (prefixtmp < 6)? prefixtmp : 6 + (1 << (prefixtmp - 6)) - 1 = (prefixtmp < 6)? prefixtmp : (1 << (prefixtmp - 6)) + 5 Also, prefixRes = (prefixVal < 6)? prefixValTR : prefixValTR + (2 << expVal) - 2 In the case of the configuration using this, the following is also acceptable.

[0120] prefixRes = (prefixtmp < 6)? prefixtmp : 6 + (2 << (prefixtmp - 6)) - 2 = (prefixtmp < 6)? prefixtmp : (2 << (prefixtmp - 6)) + 4 = (prefixtmp < 6)? prefixtmp : (1 << (prefixtmp - 5)) + 4 Here, prefixtmp<6 can also be processed as prefixtmp<5. Also, since both suffixValTR and escapeVal are fixed-length binaryizations, the former can be regarded as a fixed length of cMaxFL and the latter as a fixed length of escapeLength. Similarly, suffixRes can also be decoded.

[0121] (Quantized transform coefficient) The TU decoder 3024 (residual decoder 30241) sequentially decodes abs_remainder from the encoded data for position n, and derives the quantized transform coefficient AbsLevel from AbsLevelPass1 and abs_remainder using the following formula. n is either the position up to the last coefficient position in the scan order within the sub-block or the position up to which a predetermined number of bits are decoded.

[0122] AbsLevel[xC][yC] = AbsLevelPass1[xC][yC] + 2 * abs_remainder[n] Here, AbsLevelPass1[xC][yC] = sig_coeff_flag[xC][yC] + par_level_flag[n] + abs _level_gtx_flag[n][0] + 2 * abs_level_gtx_flag[n][1].

[0123] In the case other than the above (when n is the position after decoding a predetermined number of bits), the TU decoder 3024 decodes dec_abs_level from the encoded data and derives AbsLevel using the following formula.

[0124] if (dec_abs_level[n] == ZeroPos[n]) AbsLevel[xC][yC] = 0 else if (dec_abs_level[n] < ZeroPos) AbsLevel[xC][yC] = dec_abs_level[n] + 1 else / * dec_abs_level[n] > ZeroPos[n] * / AbsLevel[xC][yC] = dec_abs_level[n] Here, ZeroPos is derived by the following formula.

[0125] ZeroPos[n] = (QState < 2? 1 : 2) << cRiceParam Here, QState indicates the state of dependent quantization.

[0126] (Rice parameter derivation) Among the syntax elements representing the absolute value of the coefficient, abs_remainder and dec_abs_level are values binary-coded by Golomb-Rice code (or Rice code) and decoded by bypass code (equiprobable code of CABAC). abs_remainder is the difference with respect to the offset of the absolute value of the transform coefficient, and dec_abs_level is the absolute value of the transform coefficient.

[0127] Note that the Golomb-Rice code encodes by dividing the syntax element into two parts: the first half (prefix) and the second half (remainder, suffix). For the prefix, a unary code, which can encode values close to 0 with a short code amount, is used. For the suffix, fixed-length binaryization using the Rice parameter cRiceParam or EG(k) binaryization is used. The larger the Rice parameter, the longer the bit length of small values and the shorter the bit length of large values. By adjusting the Rice parameter according to the occurrence probability of the magnitude of the syntax element value, it is possible to reduce the code amount.

[0128] In the moving image decoding device and moving image encoding device of this embodiment, the value of the syntax element to be decoded is predicted from the already derived absolute value of the coefficient. Then, an appropriate Rice parameter is derived from the predicted value. Details are shown below with reference to FIG. 17.

[0129] Let the upper left coordinate of the current TU be (x0, y0), the current scan position be (xC, yC), the log2 values of the width and height of the TU be log2TbWidth and log2TbHeight respectively. Also, an array storing the absolute values of the coefficients is represented by AbsLevel[x][y] using the position (x, y). Figure 17(a) shows the positions of the decoded coefficients around the coefficient when decoding abs_remainder or dec_abs_level. Let the variable representing the sum of the absolute values of the surrounding coefficients be locSumAbs. Then, the TU decoding unit 3024 derives locSumAbs in the RRC as follows The variable representing the sum of the absolute values of the surrounding coefficients is locSumAbs. Then, the TU decoding unit 3024 derives locSumAbs in the RRC as follows as follows.

[0130] locSumAbs = 0 if(xC < (1<<log2TbWidth)-1) { locSumAbs += AbsLevel[xC+1][yC] if(xC < (1<<log2TbWidth)-2) locSumAbs += AbsLevel[xC+2][yC] if(yC < (1<<log2TbHeight)-1) locSumAbs += AbsLevel[xC+1][yC+1] } if(yC < (1<<log2TbHeight)-1) { locSumAbs += AbsLevel[xC][yC+1] if(yC < (1<<log2TbHeight)-2) locSumAbs += AbsLevel[xC][yC+2] } locSumAbs0 = locSumAbs locSumAbs = Clip3(0, 31, locSumAbs-baseLevel*5) 1) When the coefficient position (xC + 1, yC) is within the block (xC < (1<<log2TbWidth)-1), add the value of t1(AbsLevel[xC + 1][yC]) to locSumAbs. 2) When in case 1), and further when the position (xC + 2, yC) is within the block (xC < (1 << log2TbWidth) - 2), add the value of t2(AbsLevel[xC + 2][yC]) to locSumAbs. 3) When in case 1), and further when the position (xC + 1, yC + 1) is within the block (yC < (1 << log2TbHeigh) - 1), add the value of t3(AbsLevel[xC + 1][yC + 1]) to locSumAbs. 4) When the position (xC, yC + 1) is within the block (yC < (1 << log2TbHeight) - 1), add the value of t4(AbsLevel[xC][yC + 1]) to locSumAbs. 5) When in case 4), and further when the position (xC, yC + 2) is within the block (yC < (1 << log2TbHeight) - 2) add the value of t5(AbsLevel[xC][yC + 2]) to locSumAbs. Update the locSumAbs calculated from the above processing with the variable baseLevel and clip processing as follows.

[0131] locSumAbs = Clip3(0, 31, locSumAbs - baseLevel * 5) Here, baseLevel may be 4 when decoding abs_remainder (tcoeff == 0), and may be 0 when decoding dec_abs_level (tcoeff == 1). tcoeff is a variable indicating whether the decoding target is abs_remainder or dec_abs_level.

[0132] Using the table shown in FIG. 17(b), derive the Rice parameters used for decoding abs_remainder and dec_abs_level from the derived locSumAbs.

[0133] Although not shown in the figure, the TU decoder 3024 may derive the Rice parameter cRiceParam using an equation without using a table as follows.

[0134] if (baseLevel != 0) cRiceParam = Max(0, floor(log2(15 * locSumAbs)) - 7) else cRiceParam = Max(0, floor(log2(9 * locSumAbs) + 20) - 5) As an equation, as described above, the product-sum with the sum of the surrounding conversion coefficients locSumAbs and the logarithm with base 2 can be used. Here, baseLevel!=0 indicates the case when used for decoding abs_remainder, and baseLevel==0 under else indicates the case when used for decoding dec_abs_level. Here, baseLevel!=0 indicates the case when used for decoding abs_remainder, and baseLevel==0 under else indicates the case when used for decoding dec_abs_level. Here, baseLevel!=0 indicates the case when used for decoding abs_remainder, and baseLevel==0 under else indicates the case when used for decoding dec_abs_level.

[0135] The TU decoding unit 3024 does not use the decoded coefficients around in TSRC and always sets the Rice parameter cRiceParam to 1.

[0136] (Derivation of Quantized Conversion Coefficient, Dependent Quantization) The inverse quantization and inverse transformation unit 311 inverse quantizes the conversion coefficients. As quantization methods, it has two methods: scalar quantization and dependent quantization. In the case of dependent quantization, a part of the inverse quantization process may be further performed in the RRC unit 30241. In the case of dependent quantization, a part of the inverse quantization process may be further performed in the RRC unit 30241.

[0137] The inverse quantization and inverse transformation unit 311 derives the linear scale value ls[x][y] as follows according to the quantization parameter qP, rectNonTsFlag, and the values of the quantization matrix m[][]. The inverse quantization and inverse transformation unit 311 switches the derivation method of ls[][] depending on whether dependent quantization is effective and transform skip is invalid or not.

[0138] if (sh_dep_quant_used_flag &&!transform_skip_flag) ls[x][y] = (m[x][y]*levelScale[rectNonTsFlag][(qP+1)%6]) << ((qP+1) / 6) else ls[x][y] = (m[x][y]*levelScale[rectNonTsFlag][qP%6]) << (qP / 6) rectNonTsFlag = (transform_skip_flag==0 && (((Log2(nTbW)+Log2(nTbH)) & 1)==1)) ? 1 : 0 bdShift1 = transform_skip_flag==1? 10 : BitDepth+rectNonTsFlag+((Log2(nTbW)+Log2(nTbH)) / 2)-5+sh_dep_quant_used_flag In addition, when expanding the coefficient range as described later, the scaling shift value bdShift1 may be derived by the following formula. bdShift1 = BitDepth + rectNonTsFlag +(((Log2(nTbW)+Log2(nTbH)) / 2) + 10 - log2TransformRange + sh_dep_quant_used_flag (When it is not dependent quantization) When it is not dependent quantization, the inverse quantization unit 3111 (scaling unit 31111, transform coefficient clip unit 31112) uniquely derives the transform coefficient from the quantized transform coefficient and the quantization parameter. For example, the transform coefficient value d is derived.

[0139] TransCoeffLevel[x0][y0][cIdx][xC][yC] = AbsLevel[xC][yC] * (1-2*coeff_sign_flag[n]) <Formula TC_REGULAR> dz[xC][yC] = TransCoeffLevel[x0][y0][cIdx][xC][yC] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][qP%6]) << (qP / 6) dnc[xC][yC] = (dz[xC][yC]*ls[xC][yC]+((1<<bdShift1)>>1)) >> bdShift1 d[xC][yC] = Clip3(CoeffMin, CoeffMax, dnc[xC][yC]) Here, AbsLevel is the quantization conversion coefficient value, and ls and bdShift1 are variables derived from the quantization parameter qP. CoeffMin and CoeffMax are the minimum and maximum values for clipping, and are derived by the following equations.

[0140] CoeffMin = -(1 << log2TransformRange) CoeffMax = (1 << log2TransformRange) - 1 log2TransformRange indicates the range of the transform coefficients. When log2TransformRange = 15, CoeffMin = -(1<<15) and CoeffMax = (1<<15)-1.

[0141] The range of the transform coefficients log2TransformRange may be set as follows depending on BitDepth.

[0142] log2TransformRange = extended_precision_processing_flag? Max(15, BitDepth + BDOFFSET) : 15 (Equation R-1) Also, the maximum length of log2TransformRange may be restricted as follows. log2TransformRange = extended_precision_processing_flag? Min(Max(15, BitDepth + BDOFFSET), 20) : 15 (Equation R-1A) It may be further restricted when bitDepth>10. log2TransformRange = extended_precision_processing_flag? Min(Max(15, BitDepth + BDOFFSET), 20) : 15 (Equation R-1B) Here, BDOFFSET is a fixed value for calculating the range of conversion coefficients, and 4, 5, 6, etc. are appropriate extended_precision_processing_flag is a flag indicating whether to use an extended range as the range of conversion coefficients.

[0143] (In the case of dependent quantization) On the other hand, as shown in FIG. 11, dependent quantization includes two quantizers with different levels. The TU decoding unit 3024 (residual decoding unit 30242) switches four states QState using the parity of the intermediate values (AbsLevelPass1, AbsLevel) of the quantization conversion coefficients. Then, quantization and inverse quantization are performed according to QState.

[0144] QState = QStateTransTable[QState][AbsLevelPass1[xC][yC] & 1] TransCoeffLevel[x0][y0][cIdx][xC][yC] = (2*AbsLevel[xC][yC]-(QState>1? 1 : 0)) * (1-2*coeff_sign_flag[n]) <Equation TC_DQ> dz[xC][yC] = TransCoeffLevel[x0][y0][cIdx][xC][yC] The inverse quantization unit 3111 (scaling unit 31111) performs scaling using the variable ls derived using the quantization parameter derived using the quantization parameter

[0145] ls[x][y] = (m[x][y] * levelScale[rectNonTsFlag][(qP+1) % 6]) << ((qP+1) / 6 dnc[xC][yC] = (dz[xC][yC]*ls[xC][yC]+((1<<bdShift1)>>1)) >> bdShift1 The inverse quantization unit 3111 (transformation coefficient clipping unit 31112) clips the transformation coefficient between CoeffMin and CoeffMax. to clip.

[0146] d[xC][yC] = Clip3(CoeffMin, CoeffMax, dnc[xC][yC]) Here, QState is the state, QStateTransTable[][] is the table used for state transition, for example, QStateTransTable[][] = {{0, 2}, {2, 0}, {1, 3}, {3, 1}}.

[0147] QState may be derived by the following formula without using QStateTransTable[][].

[0148] QState =(32040 >>((QState << 2)+((AbsLevelPass1[xC][yC] & 1)<< 1)))& 3 Depending on the value of QState, different TransCoeffLevel (or d) can be derived even if AbsLevel is the same. Since QState is derived by referring to the quantized transformation coefficient value decoded one before, compared with general scalar (inverse) quantization, (inverse) quantization with good coding efficiency using the correlation between coefficients can be achieved.

[0149] (Inverse transformation) The inverse quantization and inverse transformation unit 311 is composed of a scaling unit 31111, an inverse non-separable transformation unit 31121, and an inverse core transformation unit 31123. to be composed of.

[0150] The inverse quantization and inverse transformation unit 311 scales (inverse quantizes) the quantized transformation coefficients qd[][] input from the entropy decoding unit 301 by the scaling unit 31111 to obtain the transformation coefficients d[][]. These quantized transformation coefficients qd[][] are coefficients obtained by performing transformations such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error and then quantizing in the encoding process, or coefficients obtained by further non-separable transformation of the transformed coefficients. When lfnst_idx!= 0, the inverse quantization and inverse transformation unit 311 performs an inverse transformation by the inverse non-separable transformation unit 31121. Further, an inverse frequency transformation such as inverse DCT or inverse DST is performed on the transformation coefficients to calculate the prediction error. When lfnst_idx == 0, without performing the processing in the inverse non-separable transformation unit 31121, an inverse transformation such as inverse DCT or inverse DST is performed on the transformation coefficients scaled by the scaling unit 31111 to calculate the prediction error. The inverse quantization and inverse transformation unit 311 outputs the prediction error to the addition unit 312.

[0151] Note that since the inverse transformation and the transformation are paired processes, the transformation and the inverse transformation may be interpreted by replacing each other. Alternatively, when the inverse transformation is called the transformation, the transformation may be called the forward transformation. For example, when the inverse non-separable transformation is called the non-separable transformation, the non-separable transformation may be called the forward non-separable transformation. Also, the core transformation may simply be called the transformation.

[0152] The inverse core transformation unit 31123 (vertical transformation unit 311231) performs a vertical transformation on the obtained residual d[][].

[0153] e[x][y] = Σ(transMatrix[y][j]×d[x][j]) (j = 0..nTbS - 1) The inverse core transformation unit 31123 (intermediate value clipping unit 311232) shifts and clips the first intermediate value e[x][y] to derive the second intermediate value g[x][y].

[0154] g[x][y] = Clip3(CoeffMin, CoeffMax, (e[x][y] + 64) >> 7) Also, CoeffMin and CoeffMax are the minimum and maximum values for clipping.

[0155] Inverse core conversion unit 31123 (horizontal conversion unit 311233) performs conversion on the conversion coefficient d[ ][ ] or the modified conversion coefficient d [ ] to derive the prediction error r[][](S5).

[0156] r[x][y] = Σ(TransMatrix[x][j]×g[j][y]) (j = 0..nTbS - 1) Then, for r[][], a shift is performed according to the bit depth (BitDepth) to derive the error resSamples[][] with the same accuracy as the predicted image derived by the predicted image generation unit 308. For example, the shift is expressed as follows.

[0157] resSamples[x][y] = (r[x][y] + (1 << (bdShift2 - 1))) >> bdShift2 (Equation BD-1) Here, the shifted value bdShift2 = 5 + log2TransformRange - BitDepth after conversion is derived. In the case of (Equation R-1), bdShift2 = extended_precision_processing_flag? (5 + BDOFFSET) : Max(20 - BitDepth, 0) Note that the above shifted value bdShift2 is a configuration used when both the conversion sizes nTbW and nTH are greater than 1, and in other cases (where either the size of nTbW or nTH of the conversion size is 1), the shifted value may be increased by 1.

[0158] bdShift2 = (nTbH>1 && nTbW>1) ? 5 + log2TransformRange - BitDepth : 6 + log2TransformRange - BitDepth (Non-separable transform) Inverse non-separable transform is applied to the transform coefficients of a part or the entire region of a TU in the moving image decoding apparatus 31. After the inverse non-separable transform is applied, an inverse separable transform (such as DCT2 and DST7) is applied to the transform coefficients. Also, the TU is divided into 4x4 sub-blocks, and non-separable transform and inverse non-separable transform are applied only to a predetermined sub-block in the upper left. Examples of the size of a TU where one of the width W and height H is 4 include 4×4, 8×4, 4×8, L×4, and 4×L (L is a natural number of 16 or more).

[0159] Also, a technique of transmitting only some low-frequency components among the transform coefficients after separable transform is called RST (Reduced Secondary Transform) transform or LFNST (Low Frequency Non-Separable-Transform). When the number nonZeroSize of non-zero transform coefficients of non-separable transform is less than or equal to the size of separable transform ((1<<log2StSize)x(1<<log2StSize)), it is called LFNST.

[0160] (S2000: Decoding of non-separable transform index) The TU decoding unit 3024 decodes an index mts_idx indicating the transform matrix of separable transform from the encoded data. mts_idx may be decoded after lfnst_idx, and may be configured to decode mtx_idx only when lfnst_idx is 0. That is, a configuration may be adopted in which a transform matrix other than DCT2 indicated by mtx_idx!=0 is used only when non-separable transform is not used (lfnst_idx is 0).

[0161] ​Also, the TU decoding unit 3024 decodes the index lfnst_idx from the encoded data. lfnst_idx is an index indicating the presence or absence of the use of non-separable transformation and the transformation matrix. The TU decoding unit 3024 derives the flag LfnstDcOnly and the flag LfnstZeroOutSigCoeffFlag. LfnstDcOnly is a flag indicating whether the transform coefficient is only DC, and LfnstZeroOutSigCoeffFlag is a flag indicating whether there is a transform coefficient in a predetermined high-frequency region (zero-out region). The TU decoding unit 3024 decodes lfnst_idx when LfnstDcOnly == 0 and LfnstZeroOutSigCoeffFlag == 1. Here, LfnstDcOnly == 0 indicates that there are transform coefficients other than the DC coefficient. LfnstZeroOutSigCoeffFlag == 1 indicates that there are no non-zero transform coefficients in the zero-out region. The TU decoding unit 3024 sets LfnstDcOnly = 1 and LfnstZeroOutSigCoeffFlag = 1 before decoding the residual of the TU. When the position of the last coefficient is other than DC (lastSubBlock == 0 && lastScanPos>0), set LfnstDcOnly = 0. When the last position is in the high-frequency region, set LfnstZeroOutSigCoeffFlag = 0. When the last position is in the high-frequency region, for example, it is the case where (lastScanPos>7 && (log2TbWidth == 2 || log2TbWidth == 3) is satisfied. If the TU decoding unit 3024 does not decode lfnst_idx, set lfnst_idx = 0.

[0162] The TU decoding unit 3024 decodes lfnst_idx when the prediction mode is the intra prediction mode and sps_lfnst_enabled_flag is 1. Note that when lfnst_idx is 0, it indicates the non-application of non-separable transformation and when it is 1, it indicates the use of one of the non-separable transformation matrix sets (pairs), and when it is 2, it indicates the use of the other transformation of the pair.

[0163] (Derivation of Rice parameters) Among the syntax elements representing the absolute value of the coefficient, abs_remainder and dec_abs_level are decoded by bypass decoding (equiprobable decoding of CABAC) for the values binary-coded by the Golomb-Rice code (or Rice code). abs_remainder is the difference with respect to the offset of the absolute value of the transform coefficient, and dec_abs_level is the absolute value of the transform coefficient.

[0164] Note that the Golomb-Rice code encodes by dividing the syntax element into two parts: the first half (prefix) and the second half (remainder, suffix). For the prefix, a unary code (alpha code) capable of encoding values close to 0 with a short code amount is used. For the suffix, fixed-length binaryization using the Rice parameter cRiceParam or EG(k) binaryization is used. The larger the Rice parameter, the longer the bit length of small values and the shorter the bit length of large values. By adjusting the Rice parameter according to the occurrence probability of the magnitude of the syntax element value, it is possible to reduce the code amount.

[0165] Let the upper left coordinate of the current TU be (x0, y0), the current scan position be (xC, yC), and the log2 values of the width and height of the TU be log2TbWidth and log2TbHeight respectively. Also, the array storing the absolute value of the coefficient is represented by AbsLevel[x][y] using the position (x, y). Figure 17(a) shows the positions of the decoded coefficients around the coefficient used when decoding abs_remainder or dec_abs_level. Let the variable representing the sum of the absolute values of the surrounding coefficients be locSumAbs. Then, the TU decoding unit 3024 derives locSumAbs in the RRC as follows.

[0166] locSumAbs = 0 if(xC < (1<<log2TbWidth)-1) { locSumAbs += AbsLevel[xC+1][yC] if (xC < (1 << log2TbWidth) - 2) locSumAbs += AbsLevel[xC + 2][yC] if (yC < (1 << log2TbHeight) - 1) locSumAbs += AbsLevel[xC + 1][yC + 1] } if (yC < (1 << log2TbHeight) - 1) { locSumAbs += AbsLevel[xC][yC + 1] if (yC < (1 << log2TbHeight) - 2) locSumAbs += AbsLevel[xC][yC + 2] } locSumAbs0 = locSumAbs locSumAbs = Clip3(0, 31, locSumAbs - baseLevel * 5) The locSumAbs calculated from the above processing is updated as follows by the variable baseLevel and the clip processing.

[0167] locSumAbs = Clip3(0, 31, locSumAbs - baseLevel * 5) Here, baseLevel may be 4 when decoding abs_remainder (tcoeff == 0), and may be 0 when decoding dec_abs_level (tcoeff == 1). tcoeff is a variable indicating whether the decoding target is abs_remainder or dec_abs_level.

[0168] Using the table RiceParamTbl, derive the Rice parameter cRiceParam used for decoding abs_remainder and dec_abs_level from the derived locSumAbs.

[0169] cRiceParam = RiceParamTbl[locSumAbs] RiceParamTbl[] = 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3 The above is the case where Rice parameter extension (hereinafter referred to as Rice extension) is not performed, and the Rice parameter cRiceParam takes values from 0 to 3.

[0170] Note that when the flag sps_rrc_rice_extension_flag indicates validity, the following processing is performed. In this case, by adding shiftVal or the like, the Rice parameter cRiceParam can use values exceeding the values from 0 to 3. Note that when range extension is performed (extended_precision_processing_flag == 1) or non-separable transformation is performed (sps_lfnst_enabled_flag == 1, lfnst_idx is other than 0), the following derivation of StatCoeff, derivation of shiftVal using localSumAbs, and update processing of cRiceParam by shiftVal do not have to be performed. The sps_rrc_rice_extension_flag is a flag indicating whether the extension of the Rice parameter used for binning is valid.

[0171] The TU decoding unit 3024 decodes abs_reminder and dec_abs_level from the residual_coding of the encoded data using the following pseudo-code. Also, the TU decoding unit 3024 derives StatCoeff. updateHist is initialized to 0 or 1 in units of residuals (TUs) with reference to the sps_persistent_rice_adaptation_enabled_flag. When sps_persistent_rice_adaptation_enabled_flag = 1, updateHist is initialized to 1 and the following processing is performed. Save the logarithm of the first abs_remainder or dec_abs_level greater than 0 in the TU to StatCoeff (update the value of StatCoeff). When updating, average with the current value of StatCoeff[cIdx] to suppress fluctuations. Note that the sps_persistent_rice_adaptation_enabled_flag is a flag indicating whether to initialize the Rice parameter used for binning at the beginning of each TU.

[0172] lastScanPos = numSbCoeff lastSubBlock = (1 << (log2TbWidth + log2TbHeight - (log2SbW + log2SbH))) - 1 HistValue = sps_persistent_rice_adaptation_enabled_flag? (1 << StatCoeff[cIdx]) : 0 updateHist = sps_persistent_rice_adaptation_enabled_flag? 1 : 0 if(abs_level_gtx_flag[n][1]) { abs_remainder[n] if(updateHist && abs_remainder[n] > 0) { StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(abs_remainder[n])) + 2) >> 1 updateHist = 0 } } ... if(sb_coded_flag[xS][yS]) { dec_abs_level[n] if(updateHist && dec_abs_level[n] > 0) { StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[n]))) >> 1 updateHist = 0 } } StatCoeff[idx] is initialized as follows.

[0173] StatCoeff[idx] = (bitDepth > 10)? (2 * Floor(Log2(bitDepth - 10)) : 0 When sps_persistent_rice_adaptation_enabled_flag = 0, StatCoeff[idx] is set to 0 and not updated.

[0174] The TU decoder 3024 derives shiftVal according to the values of sps_rrc_rice_extension_flag and localSumAbs.

[0175] if(!sps_rrc_rice_extension_flag) { shiftVal = 0 }else { shiftVal = (localSumAbs < Tx[0])? Rx[0] : ((localSumAbs < Tx[1])? Rx[1] : ((localSumAbs < Tx[2])? Rx[2] : ((localSumAbs < Tx[3])? Rx[3] : Rx[4]))) } Here, Rx[] and Tx[] are the following tables.

[0176] Tx[ ] = {32, 128, 512, 2048} Rx[ ] = {0, 2, 4, 6, 8} The TU decoder 3024 derives cRiceParam from the following processing.

[0177] locSumAbs = locSumAbs >> shiftVal locSumAbs = Clip3(0, 31, locSumAbs - baseLevel * 5) cRiceParam = RiceParamTbl[locSumAbs] When the TU decoder 3024 performs Rice extension, it adds a value shiftVal that can be non-zero and updates cRiceParam.

[0178] cRiceParam = cRiceParam + shiftVal The TU decoding unit 3024 always sets the Rice parameter cRiceParam to 1 without using the decoded coefficients around in the TSRC. Also, when sps_ts_residual_coding_rice_present_in_sh_flag is 1, sh_ts_residual_coding_rice_idx_minus1 may be encoded and decoded in the slice header. In that case, it is derived by the following formula. Infer as 0 when sh_ts_residual_coding_rice_idx_minus1 is not encoded and decoded. sps_ts_residual_coding_rice_present_in_sh_flag is a flag indicating whether sh_ts_residual_coding_rice_idx_minus1 exists in the slice header. a lag. sh_ts_residual_coding_rice_idx_minus1 is a parameter that defines the Rice parameter to be used in the TSRC.

[0179] cRiceParam = sh_ts_residual_coding_rice_idx_minus1 + 1 The inverse quantization and inverse transformation unit 311 inverse quantizes the quantization transformation block of the size of tbWidth and tbHeight. Then, it transforms using the transformation basis specified by the transformation size.

[0180] (Decoding of abs_remainder and dec_abs_level) The TU decoding unit 3024 decodes the syntax value of the absolute residual value (abs_remainder or dec_abs_level, hereinafter referred to as the residual) from the encoded data using the Rice parameter cRiceParam. For the binaryization of the syntax, a code composed of a prefix and a suffix is used. As the prefix and suffix, the alpha code of the Golomb-Rice code and the fixed-length binaryization may be used. The suffix may use the EG(k) (Exponential-Golomb code) code. In the Golomb-Rice code, the prefix is divided into prefixValTR and suffixValTR for encoding, and in the EG(k) code, the suffix is divided into the exp part expVal and the escape part escapeVal for encoding. The quotients obtained by dividing the residual by 1<<cRiceParam and ((1<<expVal))-1)<<cRiceParam are encoded using prefixValTR and expVal respectively. The remainders obtained by dividing the residual by 1<<cRiceParam and (1<<expVal<<cRiceParam) are encoded using suffixValTR and escapeVal respectively. That is, the residual is encoded as follows.

[0181] ((prefixValTR + ((1<<(expVal) -1) ) << cRiceParam) + suffixValTR + escapeVal When the residual is less than or equal to a predetermined value cMax, only prefixValTR and suffixValTR are encoded.

[0182] prefixVal = (prefixValTR<<cRiceParam) + suffixValTR When the residual is greater than a predetermined value (when suffixVal exists), in addition to prefixValTR indicating prefixVal equal to cMax, expVal and escapeVal are encoded.

[0183] prefixVal = (prefixValTR<<cRiceParam) = cMax suffixVal = (((1<<expVal)-1)<<cRiceParam) + escapeVal Here, the residual can be expressed as follows. When it can be expressed only by the prefix, the residual is given by the following formula.

[0184] Residual = ((prefixValTR) << cRiceParam) + suffixValTR For other cases (when it cannot be expressed only by the prefix, or when prefixValTR = 6), the residual is given by the following formula.

[0185] Residual = ((prefixValTR + (1<<(expVal)-1 ) << cRiceParam) + escapeVal Also, prefixVal and suffixVal may be defined as follows.

[0186] prefixVal = (prefixValTR << cRiceParam) + suffixValTR suffixVal = (((2<<expVal)-2 << cRiceParam) + escapeVal According to the above configuration, when the range of the conversion coefficient becomes large (extended_precision_processing_flag == 1), in order to limit the conversion size to a predetermined size or less, the scale of the hardware circuit can be reduced when the range of the conversion coefficient is large. This effect can always be obtained without depending on the quantization parameter. Also, when the range of the conversion coefficient does not become large (extended_precision_processing_flag == 0), the conversion size is not limited to a predetermined size or less, and a large size can be used. Therefore, when the range of the conversion coefficient does not become large, there is an effect that the reduction in coding efficiency does not occur.

[0187] The inverse quantization and inverse transformation unit 311 scales (inverse quantizes) the quantization transformation coefficients input from the entropy decoding unit 301 to obtain transformation coefficients d[][]. These quantization transformation coefficients are coefficients obtained by performing transformations such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error and then quantizing them in the encoding process. When the transform_skip_flag is 0, the inverse quantization and inverse transformation unit 311 performs inverse frequency transformations such as inverse DCT and inverse DST on the scaled transformation coefficients d[][] to calculate the prediction error res[][]. When the transform_skip_flag is 1, the inverse quantization and inverse transformation unit 311 sets res[x][y] = d[x][y]. The inverse quantization and inverse transformation unit 311 outputs the prediction error to the addition unit 312. In the encoding process, these are coefficients obtained by performing transformations such as DCT (Discrete Cosine Transform) and DST (Discrete Sine Transform) on the prediction error and then quantizing them. When the transform_skip_flag is 0, the inverse quantization and inverse transformation unit 311 performs inverse frequency transformations such as inverse DCT and inverse DST on the scaled transformation coefficients d[][] to calculate the prediction error res[][]. When the transform_skip_flag is 1, the inverse quantization and inverse transformation unit 311 sets res[x][y] = d[x][y]. The inverse quantization and inverse transformation unit 311 outputs the prediction error to the addition unit 312.

[0188] Since inverse transformation and transformation are paired processes, they may be interpreted by replacing transformation and inverse transformation with each other. Alternatively, when inverse transformation is called transformation, transformation may be called forward transformation. For example, when inverse non-separable transformation is called non-separable transformation, non-separable transformation may be called forward non-separable transformation. Also, separable transformation is simply called transformation.

[0189] The addition unit 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization and inverse transformation unit 311 pixel by pixel to generate the decoded image of the block. The addition unit 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.

[0190] (Configuration of the moving image encoding device) Next, the configuration of the moving image encoding apparatus 11 according to the present embodiment will be described. FIG. 16 is a block diagram showing the configuration of the moving image encoding apparatus 11 according to the present embodiment. The moving image encoding apparatus 11 includes a prediction image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, a coding parameter determination unit 110, a parameter coding unit 111, and an entropy coding unit 104.

[0191] The prediction image generation unit 101 generates a prediction image for each CU, which is a region obtained by dividing each picture of the image T. The prediction image generation unit 101 operates in the same manner as the prediction image generation unit 308 described above, and thus the description thereof will be omitted.

[0192] The subtraction unit 102 subtracts the pixel value of the prediction image of the block input from the prediction image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the transform / quantization unit 103 for output.

[0193] The transform / quantization unit 103 calculates transform coefficients for the prediction error input from the subtraction unit 102 by frequency conversion, and derives quantized transform coefficients by quantization. The transform / quantization unit 103 outputs the quantized transform coefficients to the entropy coding unit 104 and the inverse quantization / inverse transform unit 105. for output.

[0194] The inverse quantization / inverse transform unit 105 is the same as the inverse quantization / inverse transform unit 311 in the moving image decoding apparatus 31, and thus the description thereof will be omitted. The calculated prediction error is output to the addition unit 106.

[0195] The entropy coding unit 104 receives the quantized transform coefficients from the transform / quantization unit 103 and the coding parameters from the parameter coding unit 111. The coding parameters are, for example It is predMode indicating the prediction mode. predMode may be either MODE_INTRA indicating intra prediction, MODE_INTER indicating inter prediction, or MODE_IBC indicating intra block copy prediction where a block within the screen is copied to form a predicted image.

[0196] The entropy encoding unit 104 entropy encodes split information, prediction parameters, quantized transform coefficients, etc. to generate and output an encoded stream Te.

[0197] The parameter encoding unit 111 includes a header encoding unit 1110 (not shown), a CT information encoding unit 1111, a CU encoding unit 1112 (prediction mode encoding unit), and an inter prediction parameter encoding unit 112 and an intra prediction parameter encoding unit 113. The CU encoding unit 1112 further includes a TU encoding unit 1114 therein.

[0198] The following briefly describes the operations of each module. The parameter encoding unit 111 performs encoding processing on parameters such as header information, split information, prediction information, and quantized transform coefficients.

[0199] The CT information encoding unit 1111 encodes QT, MT (BT, TT) split information, etc. from the encoded data.

[0200] The CU encoding unit 1112 encodes CU information, prediction information, TU split flag, CU residual flag, etc.

[0201] When there is prediction error included in the TU, the TU encoding unit 1114 encodes QP update information and quantized transform coefficients. The TU encoding unit 1114 encodes the residual using the TR code and the EG(k) code as described in the TU decoding unit 3024. That is, using the TR code and the EG(k) code, sint indicating the absolute value of the residual Encode the abs_remainder and dec_abs_level elements of the syntax element. For the method of deriving the Rice parameter used for encoding, the maximum value of the prefix of the TR code, the maximum length maxPreExtLen of the exp part, and the length truncSuffixLen of the escape part, use any of the methods described in the image decoding device. The conversion / quantization unit 103 and the inverse quantization / inverse conversion unit 105 can also use the values of CoeffMin, CoeffMax, bdShift1, and bdShift2 as described in the image decoding device.

[0202] That is, the TU encoding unit 1114 may adaptively derive bdShift1 and bdShift2 according to the actual range Log2ResidualRange of the values of the transform coefficients before inverse quantization / scaling in the decoding device (after quantization / scaling in the encoding device) in the transform block. Further, depending on the extended_precision_flag and the sizes of the transform block sizes nTbW and nTbH, the range log2TransformRange of the transform coefficients and bdShift1 and bdShift2 may be derived.

[0203] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters, intra prediction parameters, and quantized transform coefficients to the entropy encoding unit 104.

[0204] The addition unit 106 adds the pixel values of the predicted image of the block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization / inverse conversion unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.

[0205] The loop filter 107 performs a deblocking filter, SAO, and ALF on the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters and may be configured with only a deblocking filter, for example.

[0206] SAO is a filter that adds an offset according to the classification result in sample units, and ALF is a filter that uses the sum of products of the transmitted filter coefficients and the reference image (or the difference between the reference image and the target pixel).

[0207] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.

[0208] The reference picture memory 109 stores the decoded images generated by the loop filter 107 at predetermined positions for each target picture and CU.

[0209] The encoding parameter determination unit 110 selects one set from a plurality of sets of encoding parameters. The encoding parameters are the QT, BT or TT segmentation information, prediction parameters, or parameters to be encoded generated in relation to these as described above. The prediction image generation unit 101 generates a prediction image using these encoding parameters.

[0210] The encoding parameter determination unit 110 calculates the amount of information and the encoding RD cost value indicating the error. The encoding parameter determination unit 110 determines that the calculated cost value is The set of encoding parameters with the minimum value is selected. As a result, the entropy encoding unit 10 4 outputs the selected set of encoding parameters as the encoding stream Te. Encoding The encoding parameter determination unit 110 stores the determined encoding parameters in the prediction parameter memory 108.

[0211] (Encoding / Decoding Method for Extended Profile) In the above, the tool restrictions by extended_precision_processing_flag, sps_ts_residual_coding_rice_present_in_sh_flag, sps_rrc_rice_extension_flag, and sps_persistent_rice_adaptation_enabled_flag were described. However, since the above tools are not defined in the first version of VVC, it is necessary to prevent them from operating in a decoder that operates in the first version.

[0212] A profile is a mechanism that defines tools (decoding operations and the structure of encoded data) that can be used to ensure compatibility. Specifically, in a certain profile, the values of specific syntax elements are restricted. A profile is defined using general_profile_idc and other syntaxes described later.

[0213] Figure 5 is a syntax table that defines the tool restrictions of the present invention. general_profile_idc is a syntax that indicates a profile. In the first version of VVC, as shown in (a) of Figure 18, the profile (profile name) and value are defined. Here, a unique value is defined for each profile. For example, for general_profile_idc, Main 10 is 1, Main 10 4:4:4 is 33, Main 10 4:4:4 Still is 97, and Multilayer Main is 17. Since general_profile_idc is a value from 0 to 127 (128 values) as shown by u(7), it is difficult to assign a unique value to the general_profile_idc of the extended profile. Note that u(n) in Description indicates that it is a fixed length of n bits. general_profile_idc: A parameter that represents a profile. general_tier_flag: A parameter that represents a tier. When equal to 1 for ptl_frame_only_constraint_flag, it indicates that all sps_field_seq_flag are 1. That is, it indicates that field pictures are not used. When equal to 1 for ptl_multilayer_enabled_flag, it indicates that there is only 1 layer. ptl_sublayer_level_present_flag[i] indicates the level value when the TemporalId (ID added to each picture for temporal hierarchicalization) is i. ptl_reserved_zero_bit is a reserved bit within profile_level_tier that is always 0. ptl_num_sub_profiles is a syntax that indicates the number of sub-profiles. general_sub_profile_idc[i] is a value of the profile specified in Rec. ITU-T T.35.

[0214] Figure 18(b) is an example of base_profile_idc of the extended profile. In this example, since the same value is used for multiple profiles, the profile cannot be uniquely determined by base_profile_idc alone. Here, one of the values 4, 36, 68, or 100 is used according to the following rules. Base profile value (basic value, base_profile_idc): · Profile of Class 1: 1 · Profile of Class 2: 4 Feature profile value (feature value, feature value): · Multilayer: 16 · 4:4:4: 32 · StillPicture: 64 · That is, general_profile_idc = base_profile_idc + Σfeature_value.

[0215] ​Specifically, if it is the first type of profile as the basic value, 1 is used; if it is the second type of profile otherwise, 4 is used. Further, the feature values are added. The feature values are added if there are tools available in the profile. For example, when using "multi layer", 16 is added to base_profile_idc. Multi layer is a feature that can combine multiple video images at the same time using multiple layers (layer_idc) into one stream. 4:4:4 is a feature that can decode the 4:4:4 format of the color space, and 32 is added to base_profile_idc. Still Picture is a feature for only the first frame (i.e., a still picture), and 64 is added to base_profile_idc. When there are multiple features, the feature values are summed up. For example, when the basic value is of the second type and has the features of 4:4:4 and still picture, as general_profile_idc, 4 (profile value of the second type) + 32 (feature value of 4:4:4) + 64 (feature value of still picture) = 100 is used. ayer), 16 is added to base_profile_idc. Multi layer can combine multiple video images at the same time using multiple layers (layer_idc) into one stream. It is a feature. 4:4:4 is a feature that can decode the 4:4:4 format of the color space, and 32 is added to base_profile_idc. Still Picture is a feature for only the first frame (i.e., a still picture), and 64 is added to base_profile_idc. When there are multiple features, the feature values are summed up. For example, when the basic value is of the second type and has the features of 4:4:4 and still picture, as general_profile_idc, 4 (profile value of the second type) + 32 (feature value of 4:4:4) + 64 (feature value of still picture) = 100 is used. That is, according to the fact that general_profile_idc is a specific value and the values of the syntax elements of the above-mentioned restricted syntax group, the encoded data for which the profile is identified is decoded or encoded. The value of general_profile_idc is shown as the sum of the basic value and one or more feature values. The feature values include at least 16 indicating multi layer, 32 indicating 4:4:4, and 64 indicating still picture. The header decoding unit decodes the encoded data having at least the value 1 of the first type and the value of the second type (in the example of the embodiment, it is 4, but not limited to 4, and may be 2, 3, 5, 6, 7, etc.) as the above-mentioned basic value. Further, the header decoding unit determines that the sum of the basic value of the second type and the above-mentioned feature value is the value of a predetermined general_profile_idc.

[0216] That is, according to the fact that general_profile_idc is a specific value and the values of the syntax elements of the above-mentioned restricted syntax group, the encoded data for which the profile is identified is decoded or encoded. The value of general_profile_idc is shown as the sum of the basic value and one or more feature values. The feature values include at least 16 indicating multi layer, 32 indicating 4:4:4, and 64 indicating still picture. The header decoding unit decodes the encoded data having at least the value 1 of the first type and the value of the second type (in the example of the embodiment, it is 4, but not limited to 4, and may be 2, 3, 5, 6, 7, etc.) as the above-mentioned basic value. The value of general_profile_idc is shown as the sum of the basic value and one or more feature values. The feature values include at least 16 indicating multi layer, 32 indicating 4:4:4, and 64 indicating still picture. The header decoding unit decodes the encoded data having at least the value 1 of the first type and the value of the second type (in the example of the embodiment, it is 4, but not limited to 4, and may be 2, 3, 5, 6, 7, etc.) as the above-mentioned basic value. Also, the header decoding unit determines that the sum of the basic value of the second type and the above-mentioned feature value is the value of a predetermined general_profile_idc. The header decoding unit determines that the sum of the basic value of the second type and the above-mentioned feature value is the value of a predetermined general_profile_idc.

[0217] Figure 19 is a table that defines the tool limits for the profile unit of the present invention.

[0218] sps_chroma_format_idc, sps_bitdepth_minus8, extended_precision_processing_flag, sps_ts_residual_coding_rice_present_in_sh_flag, sps_rrc_rice_extension_flag, and sps_persistent_rice_adaptation_enabled_flag are each syntax that defines the decoding operation. - indicates that there is no restriction on the value.

[0219] sps_chroma_format_idc: It is syntax that specifies the color space. 0 represents 4:0:0, and 1 represents 4:2:2 and 2 represents the 4:4:4 color space.

[0220] sps_bitdepth_minus8: It is the value obtained by subtracting 8 from the bit depth value, and it is a common value for luminance and color difference.

[0221] extended_precision_processing_flag: It is a flag that extends the conversion range.

[0222] sps_ts_residual_coding_rice_present_in_sh_flag: It is a flag that extends the method for deriving the Rice parameter for transform skip.

[0223] sps_rrc_rice_extension_flag: Extends the method for deriving the Rice parameter other than transform skip and is a flag.

[0224] sps_persistent_rice_adaptation_enabled_flag: The Rice parameter other than transform skip It is a flag that extends the derivation of [[unknown character "タ"]] and uses the state across multiple TUs.

[0225] In the present invention, compatibility is ensured by restricting the values of available syntax elements as shown in FIG. 19 for each profile.

[0226] FIG. 19 is a table that defines the tool limitations for each profile of the present invention. general_max_12bit_constraint_flag, general_max_10bit_constraint_flag, general_max_8bit_constraint_flag, general_max_422chroma_constraint_flag, general_max_420chroma_constraint_flag, general_max_monochrome_constraint_flag, general_one_picture_only_c onstraint_flag, general_lower_bit_rate_constraint_flag are each syntax that defines the decoding operation.

[0227] general_max_12bit_constraint_flag: When it is 1, it is a flag that restricts the value of bitdepth to 12 or less (sps_bitdepth_minus8 <= 4). When it is 0, there is no such restriction (bitdepth can be 13 or more as well).

[0228] general_max_10bit_constraint_flag: When it is 1, it is a flag that restricts the value of bitdepth to 10 or less (sps_bitdepth_minus8 <= 2). When it is 0, there is no such restriction (bitdepth can be 11 or more as well).

[0229] When general_max_8bit_constraint_flag is 1, it is a flag that restricts the value of bitdepth to 8 (when sps_bitdepth_minus8 == 0). When it is 0, there is no such restriction (bitdepth can be 9 or more).

[0230] When general_max_422chroma_constraint_flag is 1, it is a flag that restricts the color space to 4:2:2 or 4:2:0. When it is 0, there is no such restriction (use of color space 4:4:4 is also possible).

[0231] When general_max_420chroma_constraint_flag is 1, it is a flag that restricts the color space to 4:2:0. When it is 0, there is no such restriction (use of color spaces other than 4:2:0 is also possible).

[0232] When general_max_monochrome_constraint_flag is 1, it is a flag that restricts the color space to 4:0:0 When it is 0, there is no such restriction (use of color spaces other than 4:0:0 is also possible).

[0233] When general_one_picture_only_constraint_flag is 1, it indicates that there is only one picture (i.e., a still image) in the bitstream. When it is 0, there is no such restriction (multiple pictures are also possible).

[0234] When general_lower_bit_rate_constraint_flag is 1, it is a flag that restricts the available bitrate. When it is 0, there is no such restriction (different bitrate restrictions are possible).

[0235] In the present invention, compatibility is ensured by restricting the values of syntax elements as shown in FIG. 20 for each profile.

[0236] Profiles having the value of general_profile_idc of Class 2 have the above color space and bit depth The operation of the tool is defined by the values of a group of syntactic elements that indicate picture type restrictions.

[0237] (Configuration example) FIG. 21 is an example of a syntax table that defines tool limitations of the present invention. As shown in FIG. 21, general_constraints_info, which is general restriction information, has a gci_present_flag, and when the gci_present_flag is 1, it includes general tool restriction information. The general tool restriction information is composed of first general tool restriction information such as gci_intra_only_constraint_flag, gci_all_layers_independent_constraint_flag, gci_one_au_only_constraint_flag, ..., gci_no_virtual_boundaries_constraint_flag, and a group of restriction syntaxes (second general tool restriction information) that are decoded and encoded depending on general_profile_idc. The group of restriction syntaxes shown in SYN_CONST3 has already been described. SYN_CONST1 indicates a determination as to whether general_profile_idc is a predetermined value (here, a profile value of the second type, which consists of the addition of a basic value of the second type and a feature value, 4, 36, 68, 100). Hereinafter, a determination as to whether general_profile_idc is a profile value of the second type is shown as general_profile_idc being a predetermined value. That is, it is determined whether to decode and encode the restriction syntax length shown in subsequent SYN_CONST2, the group of restriction syntaxes shown in SYN_CONST3 (second general restriction syntax group), and the spare bits shown in SYN_CONST4. When general_profile_idc is a predetermined value, the header decoder 3020 and the header encoder 1110 of the present invention decode and encode a predetermined group of restriction syntaxes.As the restricted syntax group, any one of general_max_12bit_constraint_flag, general_max_10bit_constraint_flag, general_max_8bit_constraint_flag, general_max_422chroma_constraint_flag, general_max_420chroma_constraint_flag, general_max_monochrome_constraint_flag, general_one_picture_only_constraint_flag, general_lower_bit_rate_constraint_flag may be included as shown in the figure. Further, the restricted syntax group may include general_max_12bit_constraint_flag for restricting at least the bit depth and general_max_422chroma_constraint_flag, general_max_420chroma_constraint_flag for restricting the color space. That is, in the moving image decoding apparatus and the moving image encoding apparatus of the present invention, while using the same general_profile_idc, when general_profile_idc is a predetermined value, a new constraint flag is defined. And, with that constraint flag, individual profiles are identified. By the above, while including the same general_profile_idc, bitstreams corresponding to different bit depths and color spaces are defined as profiles, and it is possible to provide a moving image decoding apparatus that supports decoding of the bitstream and a moving image encoding apparatus that supports encoding of the bitstream.

[0238] More specifically, the header decoding unit 3020 and the header encoding unit 1110 of the present invention determine whether the general_profile_idc is a predetermined value (SYN_CONST1 in FIG. 21). If the determination is true, the restricted syntax length shown in SYN_CONST2, the restricted syntax group shown in SYN_CONST3, and the spare bits shown in SYN_CONST4 are decoded and encoded. The gci_num_profile_constraint_bits in SYN_CONST2 indicates the restricted syntax length and the number of spare bits (since they are 1 bit each, the bit length). The moving image decoding apparatus and the moving image encoding apparatus decode and encode a bit stream having a restriction that "the value of gci_num_profile_constraint_bits is equal to or greater than the number (NCONST3) of restricted syntax groups" as bit stream compliance. Here, NCONST3 is 8, but it may be other values according to the number of restricted syntax groups SYN_CONST3. Also, the value may be changed in the future.

[0239] When the general_profile_idc is a predetermined value, the header decoding unit 3020 and the header encoding unit 1110 decode and encode encoded data in which the gci_present_flag is always 1 as a bit stream compliance restriction. When the general_profile_idc is a predetermined value, since the gci_present_flag is 1, general tool restriction information (the first general tool restriction information and the restricted syntax group) is always decoded and encoded. Thereby, the restriction items of the profile that cannot be uniquely identified only by the general_profile_idc can be decoded, encoded, and the profile can be identified. Furthermore, the tool restrictions defined by the profile can also be realized at the same time.

[0240] When the gci_present_flag is 1, the header decoding unit 3020 and the header encoding unit 1110 decode and encode the first general tool restriction information such as the gci_intra_only_constraint_flag, the gci_all_layers_independent_constraint_flag, the gci_one_au_only_constraint_flag, ..., the gci_no_virtual_boundaries_constraint_flag. When the gci_intra_only_constraint_flag is 1, it is a flag that restricts to intra slice I. When it is 0, there is no such restriction (inter slice P and B are also available).

[0241] When the gci_all_layers_independent_constraint_flag is 1, it indicates that all layers can be encoded and decoded independently (that is, they do not refer to other layers. Prediction that includes other inter-layer images in the reference picture list is not performed). When it is 0, there is no such restriction.

[0242] When the gci_one_au_only_constraint_flag is 1, it indicates that there is only one access unit in this case. When it is 0, there is no such restriction.

[0243] When the gci_no_virtual_boundaries_constraint_flag is 1, the tool that does not perform the loop filter called the virtual boundary on specific horizontal and vertical lines is not used. When it is 0, there is no such restriction. in this case. When it is 0, there is no such restriction.

[0244] Furthermore, the header decoding unit 3020 and the header encoding unit 1110 may decode and encode gci_num_profile_constraint_bits after the gci_no_virtual_boundaries_constraint_flag. gci_num_profile_constraint_bits is a fixed-length 8-bit value with the same length as gci_num_reserved_bits. Also, as shown in SYN_CONST4, the header decoding unit 3020 and the header encoding unit 1110 may decode and encode the gci_reserved_profile_zero_bit a predetermined number of times. The predetermined number of times is the number obtained by subtracting a fixed value NCONST3 from the bit length specified by gci_num_profile_constraint_bits. gci_reserved_profile_zero_bit is 0 in a certain version of the standard, but may be a value other than 0 in subsequent versions of the standard. That is, gci_reserved_profile_zero_bit may be extended in the future as another syntax element general_xxx_constraint_flag.

[0245] If it is not the above-mentioned predetermined general_profile_idc, the header decoding unit 3020 and the header encoding unit 1110 may decode the 8-bit fixed-length gci_num_reserved_bits as shown in SYN_CONST5, and decode and encode the gci_reserved_zero_bit indicated by num_reserved_bits.

[0246] According to the above configuration, the header decoding unit 3020 and the header encoding unit 1110 determine whether general_profile_idc is a predetermined value. If the determination is true, they decode and encode the syntax element indicating the length of the restricted syntax group from the encoded data. Also, by decoding and encoding the restricted syntax group equal to or more than the number equal to the syntax indicating the length of the restricted syntax group, it is possible to define the restrictions in the extended profile as syntax elements and achieve the effect of enabling decoding and encoding.

[0247] Moreover, by making gci_num_reserved_bits equal to gci_num_profile_constraint_bits and enabling SYN_CONST3 with a length specified by gci_num_profile_constraint_bits to be truncated in the same way as gci_num_reserved_bits of SYN_CONST4, an effect of providing a backward-compatible moving image decoding apparatus and a moving image encoding apparatus is achieved.

[0248] Furthermore, depending on the bit length of gci_num_profile_constraint_bits, a moving image decoding apparatus and a moving image encoding apparatus having forward compatibility for increasing future restricted syntax groups may be provided by configuring to skip gci_reserved_profile_zero_bit.

[0249] (Another configuration example) FIG. 22 shows another example of a syntax table that defines the tool limitations of the present invention.

[0250] More specifically, the header decoding unit 3020 and the header encoding unit 1110 of the present invention determine whether the general_profile_idc is a predetermined value (SYN_CONST1). If true, they decode and encode the restricted syntax length shown in SYN_CONST2B, the restricted syntax group shown in SYN_CONST3, and the leading bit length and leading bits of the restricted syntax group shown in SYN_CONST4B. The gci_num_profile_constraint_N_bits in SYN_CONST2B indicates the number of restricted syntax lengths (since it is 1 bit each, it is the bit length). The moving image decoding device and the moving image encoding device decode and encode the same number of syntaxes (restricted syntax groups) as the value of gci_num_profile_constraint_N_bits (here NCONST3) when the general_profile_idc is a predetermined value. The header decoding unit 3020 and the header encoding unit 1110 decode and encode gci_num_profile_constraint_N_bits equal to gci_num_reserved2_bits as shown in SYN_CONST2B when the general_profile_idc is a predetermined value. Also, the header decoding unit 3020 and the header encoding unit 1110 decode and encode gci_reserved_profile_zero_bit by the bit length specified by the leading bit length gci_num_reserved2_bits of the restricted syntax group as shown in SYN_CONST4B. That is, they decode and encode the number of leading bits indicated by the leading bit length. The gci_reserved_profile_zero_bit is 0 in a certain version of the standard, but may be a value other than 0 in subsequent versions of the standard. That is, this gci_reserved_profile_zero_bit may be extended in the future as another syntax element general_xxx_constraint_flag.

[0251] If it is not the above-mentioned predetermined general_profile_idc, the header decoding unit 3020 and the header encoding unit 1110 decode the 8-bit fixed-length gci_num_reserved_bits as shown in SYN_CONST5, and decode and encode the reserved bits indicated by num_reserved_bits. gci_reserved_profile_zero_bit and gci_reserved_zero_bit are 1-bit flags.

[0252] According to the above configuration, the header decoding unit 3020 and the header encoding unit 1110 determine whether the general_profile_idc is a predetermined value. If the determination is true, the syntax element indicating the length of the restricted syntax group is decoded and encoded from the encoded data, and the same number of restricted syntax groups as the syntax indicating the length of the restricted syntax group are decoded and encoded. Thereby, the restriction in the extended profile is defined as a syntax value, and the effect of enabling decoding and encoding is achieved. Also, in the first version of the moving image decoding apparatus where the bit lengths of gci_num_reserved_bits and gci_num_profile_constraint_N_bits are made equal and the gci_num_reserved_bits in SYN_CONST5 are read, the gci_num_profile_constraint_N_bits of the same length can be read without problems.

[0253] Furthermore, by making the value of gci_num_profile_constraint_N_bits equal to the sum of the bit length of SYN_CONST3, the bit length (8 bits) of gci_num_reserved2_bits, and the value of gci_num_reserved2_bits, SYN_CONST3, gci_num_reserved2_bits, and SYN_CONST4B can be discarded in the same way as the gci_num_reserved_bits in SYN_CONST5. Therefore, there is an effect of providing a moving image decoding apparatus and a moving image encoding apparatus having backward compatibility.

[0254] Furthermore, by decoding, encoding gci_num_reserved2_bit, and skipping gci_reserved_profile_zero_bit according to its bit length, a video decoding device and a video encoding device having forward compatibility that can increase future restricted syntax groups may be provided.

[0255] Note that a part of the video encoding device 11 and the video decoding device 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the addition unit 312, the predicted image generation unit 101, the subtraction unit 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, and the parameter encoding unit 111 may be realized by a computer. In that case, a program for realizing this control function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to be realized. Here, the "computer system" refers to a computer system built in either the video encoding device 11 or the video decoding device 31 and includes hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built in a computer system. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that holds a program dynamically for a short time, and something that holds a program for a certain time, like a volatile memory inside a computer system serving as a server or a client in that case. Also, the above program may be for realizing a part of the above-described functions, and may further be realized in combination with a program already recorded in the computer system for realizing the above-described functions.

[0256] Further, part or all of the moving image encoding device 11 and the moving image decoding device 31 in the above-described embodiments may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the moving image encoding device 11 and the moving image decoding device 31 may be individually made into a processor, or part or all of them may be integrated and made into a processor. Also, the method of integrating into a circuit is not limited to LSI, and it may be realized by a dedicated circuit or a general-purpose processor. Further, when a technology for integrating into a circuit that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit using such technology may be used.

[0257] As described above, an embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.

[0258] 〔Application Example〕 The above-described moving image encoding device 11 and moving image decoding device 31 can be mounted and used in various devices that perform transmission, reception, recording, and reproduction of moving images. Note that the moving image may be a natural moving image captured by a camera or the like, or an artificial moving image (including CG and GUI) generated by a computer or the like.

[0259] First, the fact that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for transmission and reception of moving images will be described with reference to FIG. 2.

[0260] FIG. 2 shows a block diagram showing the configuration of a transmission device PROD_A equipped with the moving image encoding device 11. As shown in the figure, the transmitting device PROD_A includes an encoding unit PROD_A1 that obtains encoded data by encoding a moving image, a modulation unit PROD_A2 that obtains a modulated signal by modulating a carrier wave with the encoded data obtained by the encoding unit PROD_A1, and a transmitting unit PROD_A3 that transmits the modulated signal obtained by the modulation unit PROD_A2. The above-mentioned moving image encoding device 11 is used as this encoding unit PROD_A1.

[0261] The transmission device PROD_A captures moving images as a supply source of moving images to be input to the encoding unit PROD_A1. The apparatus may further include a camera PROD_A4 for recording moving images, a recording medium PROD_A5 for recording moving images, an input terminal PROD_A6 for inputting moving images from the outside, and an image processing unit A7 for generating or processing images. In the figure, a configuration in which the transmitting device PROD_A includes all of these components is illustrated, but some of them may be omitted.

[0262] The recording medium PROD_A5 may be one in which unencoded moving images are recorded. Alternatively, the recording medium PROD_A5 may be a recording medium that has recorded thereon a moving image encoded by a recording encoding method different from the transmission encoding method. In the latter case, a decoding unit ( It is advisable to use a conductor (not shown) between the two.

[0263] FIG. 2 is a block diagram showing the configuration of a receiving device PROD_B equipped with a video decoding device 31. FIG. 1 shows a receiving device PROD_B. As shown in the figure, the receiving device PROD_B includes a receiving unit PROD_B1 that receives a modulated signal, a demodulation unit PROD_B2 that obtains coded data by demodulating the modulated signal received by the receiving unit PROD_B1, and a decoding unit PROD_B3 that obtains a moving image by decoding the coded data obtained by the demodulation unit PROD_B2. The above-mentioned moving image decoding device 31 is used as this decoding unit PROD_B3.

[0264] As a destination for the moving image output by the decoding unit PROD_B3, the receiving device PROD_B may further include a display PROD_B4 for displaying the moving image, a recording medium PROD_B5 for recording the moving image, and an output terminal PROD_B6 for outputting the moving image externally. In the figure, a configuration in which the receiving device PROD_B includes all of these is illustrated, but a part thereof may be omitted. Note that the recording medium PROD_B5 may be for recording an unencoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5. Note that the recording medium PROD_B5 may be for recording an unencoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.

[0265] Note that the recording medium PROD_B5 may be for recording an unencoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5. Note that the recording medium PROD_B5 may be for recording an unencoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5. Note that the recording medium PROD_B5 may be for recording an unencoded moving image, or may be encoded using a recording encoding method different from the transmission encoding method. In the latter case, an encoding unit (not shown) for encoding the moving image acquired from the decoding unit PROD_B3 according to the recording encoding method may be interposed between the decoding unit PROD_B3 and the recording medium PROD_B5.

[0266] Note that the transmission medium for transmitting the modulation signal may be wireless or wired. Also, the transmission mode for transmitting the modulation signal may be broadcasting (here, referring to a transmission mode where the transmission destination is not specified in advance) or communication (here, referring to a transmission mode where the transmission destination is specified in advance). That is, the transmission of the modulation signal may be realized by any of wireless broadcasting, wired broadcasting, wireless communication, and wired communication.

[0267] For example, a broadcast station (broadcast equipment, etc.) / reception station (television receiver, etc.) for terrestrial digital broadcasting is an example of a transmission device PROD_A / receiving device PROD_B that transmits and receives a modulation signal by wireless broadcasting. Also, a broadcast station (broadcast equipment, etc.) / reception station (television receiver, etc.) for cable television broadcasting is an example of a transmission device PROD_A / receiving device PROD_B that transmits and receives a modulation signal by wired broadcasting.

[0268] In addition, servers (workstations, etc.) / clients (television receivers, personal computers, smartphones, etc.) for VOD (Video On Demand) services and video sharing services using the Internet are examples of a transmission device PROD_A / receiving device PROD_B that transmits and receives modulated signals through communication (usually, either wireless or wired is used as the transmission medium in a LAN, and wired is used as the transmission medium in a WAN). Here, personal computers include desktop PCs, laptop PCs, and tablet PCs. Also, smartphones include multifunctional mobile phone terminals.

[0269] Note that in addition to the function of decrypting the encoded data downloaded from the server and displaying it on the display, the client of the video sharing service has a function of encoding the moving images captured by the camera and uploading them to the server. That is, the client of the video sharing service functions as both the transmission device PROD_A and the receiving device PROD_B.

[0270] Next, it will be described with reference to FIG. 3 that the above-described moving image encoding device 11 and moving image decoding device 31 can be used for recording and playing back moving images.

[0271] FIG. 3 shows a block diagram showing the configuration of a recording device PROD_C equipped with the above-described moving image encoding device 11. As shown in the figure, the recording device PROD_C includes an encoding unit PROD_C1 that obtains encoded data by encoding a moving image, and a writing unit PROD_C2 that writes the encoded data obtained by the encoding unit PROD_C1 to a recording medium PROD_M. The above-described moving image encoding device 11 is used as this encoding unit PROD_C1.

[0272] ​Note that the recording medium PROD_M may be of a type built into the recording device PROD_C, such as (1) an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., or (2) a type connected to the recording device PROD_C, such as an SD memory card, a USB (Universal Serial Bus) flash memory, etc., or (3) a type loaded into a drive device (not shown) built into the recording device PROD_C, such as a DVD (Digital Versatile Disc: registered trademark), a BD (Blu-ray Disc: registered trademark), etc. It may also be of a type connected to the recording device PROD_C, or (3) a type loaded into a drive device (not shown) built into the recording device PROD_C.

[0273] In addition, the recording device PROD_C may further include a camera PROD_C3 that captures a moving image, an input terminal PROD_C4 for inputting a moving image from the outside, a receiving unit PROD_C5 for receiving a moving image, and an image processing unit PROD_C6 for generating or processing an image, as a supply source of the moving image input to the encoding unit PROD_C1. In the figure, a configuration in which the recording device PROD_C includes all of these is illustrated, but a part of them may be omitted. Note that the receiving unit PROD_C5 may receive an unencoded moving image, or may receive encoded data encoded by an encoding method for transmission different from the encoding method for recording. In the latter case, a transmission decoder (not shown) for decoding the encoded data encoded by the encoding method for transmission may be interposed between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0274] Note that the receiving unit PROD_C5 may receive an unencoded moving image, or may receive encoded data encoded by a transmission encoding method different from the recording encoding method. In the latter case, it is advisable to interpose a transmission decoder (not shown) for decoding the encoded data encoded by the transmission encoding method between the receiving unit PROD_C5 and the encoding unit PROD_C1.

[0275] Examples of such a recording device PROD_C include a DVD recorder, a BD recorder, an HDD (Hard Disk Drive) recorder, etc. (in this case, the input terminal PROD_C4 or the receiving unit PROD_C5 becomes the main supply source of the moving image). Also, a camcorder (in this case, the camera PROD_C3 is the main supply source of the moving image). a personal computer (in this case, the receiving unit PROD_C5 or the image processing unit C6 is the main source of the moving image), a smartphone (in this case, the camera PROD_C3 or the receiving unit PROD_C5 is the main source of the moving image), etc. are also examples of such a recording device PROD_C.

[0276] Also, FIG. 3 shows a block diagram of a playback device PROD_D equipped with the above-described moving image decoding device 31. As shown in the figure, the playback device PROD_D includes a reading unit PROD_D1 that reads encoded data written on a recording medium PROD_M, and a decoding unit PROD_D2 that obtains a moving image by decoding the encoded data read by the reading unit PROD_D1. The above-described moving image decoding device 31 is used as this decoding unit PROD_D2.

[0277] Note that the recording medium PROD_M may be of a type built into the playback device PROD_D, such as (1) an HDD or an SSD, or may be of a type connected to the playback device PROD_D, such as (2) an SD memory card or a USB flash memory, or may be loaded into a drive device (not shown) built into the playback device PROD_D, such as (3) a DVD or a BD.

[0278] Also, the playback device PROD_D may further include a display PROD_D3 that displays the moving image, an output terminal PROD_D4 for outputting the moving image to the outside, and a transmitting unit PROD_D5 that transmits the moving image, as destinations for the moving image output by the decoding unit PROD_D2. In the figure, all of these are illustrated as configurations included in the playback device PROD_D, but some may be omitted.

[0279] Note that the transmitting unit PROD_D5 may transmit an unencoded moving image.​​​​​ It may be configured to transmit encoded data encoded by an encoding method for transmission that is different from the encoding method for recording. In the latter case, an encoding unit (not shown) that encodes the moving image by the encoding method for transmission may be interposed between the decoding unit PROD_D2 and the transmission unit PROD_D5.

[0280] Examples of such a playback device PROD_D include a DVD player, a BD player, an HDD player, etc. (in this case, the output terminal PROD_D4 to which a television receiver or the like is connected is the main supply destination of the moving image) Also, a television receiver (in this case, the display PROD_D3 is the main supply destination of the moving image), digital signage (also referred to as an electronic billboard or an electronic bulletin board, etc., where the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the moving image), a desktop PC (in this case, the output terminal PROD_D4 or the transmission unit PROD_D5 is the main supply destination of the moving image), a laptop or tablet PC (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the moving image), a smartphone (in this case, the display PROD_D3 or the transmission unit PROD_D5 is the main supply destination of the moving image), etc. are also examples of such a playback device PROD_D There is.

[0281] (Hardware implementation and software implementation) In addition, each block of the above-described moving image decoding device 31 and moving image encoding device 11 may be realized hardware-wise by a logic circuit formed on an integrated circuit (IC chip), or may be realized software-wise using a CPU (Central Processing Unit).

[0282] In the latter case, each of the above devices includes a CPU that executes instructions of a program for realizing each function, the above a ROM (Read Only Memory) that stores the program, a RAM (Random It is provided with a storage device (recording medium) such as a memory for storing the access memory, the above program, and various data. The object of the embodiment of the present invention is to supply a recording medium in which program codes (executable program, intermediate code program, source program) of the control programs of the above devices, which are software for realizing the above-described functions, are recorded in a computer-readable manner to each of the above devices, and it is also achievable by a computer (or CPU or MPU) reading and executing the program codes recorded on the recording medium.

[0283] As the above recording medium, for example, tapes such as magnetic tapes and cassette tapes, magnetic disks such as floppy (registered trademark) disks / hard disks, and optical disks including optical disks such as CD-ROM (Compact Disc Read-Only Memory) / MO disks (Magneto-Optical disc) / MD (Mini Disc) / DVD (Digital Versatile Disc: registered trademark) / CD-R (CD Recordable) / Blu-ray Disc (Blu-ray Disc: registered trademark), cards such as IC cards (including memory cards) / optical cards, semiconductor memories such as mask ROM / EPROM (Erasable Programmable Read-Only Memory) / EEPROM (Electrically Erasable and Programmable Read-Only Memory: registered trademark) / flash ROM, or logic circuits such as PLD (Programmable logic device) and FPGA (Field Programmable Gate Array) can be used.

[0284] In addition, each of the above devices may be configured to be connectable to a communication network, and the above program code may be supplied via the communication network. This communication network only needs to be capable of transmitting the program code and is not particularly limited. For example, the Internet, intranet, extranet, LAN (Local Area Network), ISDN (Integrated Services Digital Network), VAN (Value-Added Network), CATV (Community Antenna television / Cable Television) communication network, virtual private network, telephone line network, mobile communication network, satellite communication network, etc. can be used. Also, the transmission medium constituting this communication network only needs to be a medium capable of transmitting the program code and is not limited to a specific configuration or type. For example, it can be wired such as IEEE (Institute of Electrical and Electronic Engineers) 1394, USB, power line carrier, cable TV line, telephone line, ADSL (Asymmetric Digital Subscriber Line) line, etc., or wireless such as infrared rays like IrDA (Infrared Data Association) and remote controls, Bluetooth (registered trademark), IEEE802.11 wireless, HDR (High Data Rate), NFC (Near Field Communication), DLNA (Digital Living Network Alliance: registered trademark), mobile phone network, satellite line, terrestrial digital broadcast network, etc. Note that the embodiments of the present invention can also be realized in the form of a computer data signal embedded in a carrier wave, in which the above program code is embodied by electronic transmission.

[0285] The embodiments of the present invention are not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. That is, embodiments obtained by combining appropriately modified technical means within the scope shown in the claims are also included in the technical scope of the present invention.

Industrial Applicability

[0286] Embodiments of the present invention are suitably applicable to a moving image decoding apparatus that decodes encoded data in which image data is encoded, and a moving image encoding apparatus that generates encoded data in which image data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding apparatus and referred to by the moving image decoding apparatus.

Explanation of Signs

[0287] 31 Moving image decoding apparatus 301 Entropy decoding unit 302 Parameter decoding unit 3020 Header decoding unit 308 Predicted image generation unit 311 Inverse quantization and inverse transformation unit 312 Addition unit 11 Moving image encoding apparatus 101 Predicted image generation unit 102 Subtraction unit 103 Transformation and quantization unit 104 Entropy encoding unit 105 Inverse quantization and inverse transformation unit (scaling unit) 107 Loop filter 110 Encoded parameter determination unit 111 Parameter encoding unit 1110 Header encoding unit 1111 CT information encoding unit 1112 CU encoding unit (prediction mode encoding unit) 1114 TU encoding unit 311 Inverse quantization and inverse transformation unit (scaling unit) 3111 Inverse quantization unit 3112 Inverse transformation unit 31123 Inverse core transformation unit​

Claims

1. An image decoding apparatus for decoding general restriction information, comprising: a header decoding unit that decodes a first flag indicating whether there is a syntax element of the general restriction information from encoded data; wherein the header decoding unit: when the first flag indicates the presence of a syntax element of the general restriction information, decodes first general tool restriction information from the encoded data, and further decodes second general tool restriction information and a first parameter regarding the length of padding bits; decodes the second general tool restriction information from the encoded data; skips the padding bits from the encoded data a predetermined number of times; wherein the predetermined number of times is a value obtained by subtracting a fixed value determined in advance from the first parameter, characterized in that the image decoding apparatus.

2. An image encoding apparatus for encoding general restriction information, comprising: a header encoding unit that encodes a first flag indicating whether there is a syntax element of the general restriction information; wherein the header encoding unit: when the first flag indicates the presence of a syntax element of the general restriction information, encodes first general tool restriction information, and further encodes second general tool restriction information and a first parameter regarding the length of padding bits; encodes the second general tool restriction information; encodes the padding bits a predetermined number of times; wherein the predetermined number of times is a value obtained by subtracting a fixed value determined in advance from the first parameter, characterized in that the image encoding apparatus.

3. An image decoding method for decoding general restriction information, comprising: decoding a first flag indicating whether there is a syntax element of the general restriction information from encoded data; when the first flag indicates the presence of a syntax element of the general restriction information, decoding first general tool restriction information from the encoded data, and further decoding second general tool restriction information and a first parameter regarding the length of padding bits; decoding the second general tool restriction information from the encoded data; skipping the padding bits from the encoded data a predetermined number of times; wherein the predetermined number of times is a value obtained by subtracting a fixed value determined in advance from the first parameter, characterized in that the image decoding method.

4. An image encoding method for encoding general restriction information, comprising: encoding a first flag indicating whether there is a syntax element of the general restriction information; ​ ​ ​ ​ ​ ​ ​ When the first flag indicates the presence of a syntax element of the general restriction information, first encode the first general tool restriction information, and further encode the first parameter regarding the second general tool restriction information and the length of the spare bits, encode the second general tool restriction information, encode the spare bits a predetermined number of times, wherein the predetermined number of times is a value obtained by subtracting a fixed value predetermined from the first parameter from an image encoding method characterized by being.