Image processing apparatus and method

By allocating predetermined context variables to the bin sequence of adaptive orthogonal transform identifiers in image encoding and performing context encoding, the problem of increased memory and processing load is solved, and a more efficient encoding and decoding process is achieved.

CN120281922APending Publication Date: 2025-07-08SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510484259.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-06-19
Filing Date
2020-05-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In image encoding, improper allocation of context variables in the prior art leads to an increase in memory usage and processing load, especially during encoding and decoding of adaptive orthogonal transform identifiers.

Method used

By assigning a predetermined context variable to the first bin in the bin sequence of the adaptive orthogonal transform identifier and performing context encoding on it, or assigning the context variable based on the block size parameters and performing context encoding on the first bin in the bin sequence, while applying bypass encoding to reduce the number of unnecessary context-encoded bins.

Benefits of technology

It effectively suppresses the increase in memory usage and encoding processing load, and improves encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281922A_ABST
    Figure CN120281922A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing apparatus and method. The image processing apparatus includes circuitry configured to assign a context variable to a first bin in a sequence of bins based on a parameter regarding a tree type, the sequence of bins being obtained by binarizing a quadratic orthogonal transformation identifier indicating a pattern of quadratic orthogonal transformation in image coding, and performing context coding for a first bin in the sequence of bins.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application with application number 202080042943.4 and invention title "Image Processing Apparatus and Method". The international filing date of the parent application is May 8, 2020, and the international application number is PCT / JP2020 / 018617. Technical Field

[0002] The present disclosure relates to an image processing apparatus and method, and more particularly to an image processing apparatus and method capable of suppressing an increase in load. Background Art

[0003] Conventionally, in image coding, an adaptive orthogonal transform identifier mts_idx has been signaled (encoded / decoded) as mode information regarding an adaptive orthogonal transform (multiple transform selection (MTS)). For the coding of the adaptive orthogonal transform identifier mts_idx, context coding is applied, in which the adaptive orthogonal transform identifier mts_idx is binarized, and a context variable ctx is assigned to each bin in a bin sequence bins to perform arithmetic coding. Further, context decoding corresponding to the context coding is applied to the decoding of the coded data of the adaptive orthogonal transform identifier mts_idx.

[0004] Citation List

[0005] Non-Patent Literature

[0006] Non-Patent Literature 1: Benjamin Bross, Jianle Chen, Shan Liu, "Versatile Video Coding (Draft 5)", JVET-N1001v8, Fourteenth Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019. Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] However, in the case of such a method, there is a possibility that the context variable increases unnecessarily and the memory usage increases unnecessarily. That is, there is a possibility that the load of the encoding process and the decoding process increases.

[0009] The present disclosure has been made in view of the above problems, and an object thereof is to suppress an increase in the load of the encoding process and the decoding process.

[0010] Solutions to the Technical Problems

[0011] An image processing apparatus according to one aspect of the present technology is an image processing apparatus including an encoding unit configured to assign a predetermined context variable to a first bin in a bin sequence and perform context encoding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0012] An image processing method according to one aspect of the present technology is an image processing method including assigning a predetermined context variable to a first bin in a bin sequence and performing context encoding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0013] An image processing apparatus according to another aspect of the present technology is an image processing apparatus including an encoding unit configured to assign a context variable based on a parameter regarding a block size to a first bin in a bin sequence and perform context encoding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0014] An image processing method according to another aspect of the present technology is an image processing method including assigning a context variable based on a parameter regarding a block size to a first bin in a bin sequence and performing context encoding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0015] In an image processing apparatus and method according to one aspect of the present technology, a predetermined context variable is assigned to a first bin in a bin sequence and context encoding is performed on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0016] In an image processing apparatus and method according to another aspect of the present technology, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence and context encoding is performed on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a diagram for describing an example of a state for encoding an adaptive orthogonal transform identifier.

[0018] Figure 2 It is a diagram for describing an example of the state for encoding an adaptive orthogonal transform identifier.

[0019] Figure 3 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0020] Figure 4 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0021] Figure 5 It is a diagram showing a comparison example of the number of contexts of each method, the number of context coding bins, and the number of bypass coding bins.

[0022] Figure 6 It is a block diagram showing an example of the main configuration of an image coding device.

[0023] Figure 7 It is a block diagram showing an example of the main configuration of a coding unit.

[0024] Figure 8 It is a flowchart for describing an example of the process of image coding processing.

[0025] Figure 9 It is a flowchart for describing an example of the process of image coding processing.

[0026] Figure 10 It is a block diagram showing an example of the main configuration of an image decoding device.

[0027] Figure 11 It is a block diagram showing an example of the main configuration of a decoding unit.

[0028] Figure 12 It is a flowchart for describing an example of the process of image decoding processing.

[0029] Figure 13 It is a flowchart for describing an example of the process of decoding processing.

[0030] Figure 14 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0031] Figure 15 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0032] Figure 16 It is a diagram showing a comparison example of the number of contexts of each method, the number of context coding bins, and the number of bypass coding bins.

[0033] Figure 17 It is a flowchart for an example describing the process of encoding processing.

[0034] Figure 18 It is a flowchart for an example describing the process of decoding processing.

[0035] Figure 19 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0036] Figure 20 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0037] Figure 21 It is a diagram showing a comparison example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.

[0038] Figure 22 It is a flowchart for an example describing the process of encoding processing.

[0039] Figure 23 It is a flowchart for an example describing the process of decoding processing.

[0040] Figure 24 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0041] Figure 25 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0042] Figure 26 It is a diagram showing a comparison example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.

[0043] Figure 27 It is a flowchart for an example describing the process of encoding processing.

[0044] Figure 28 It is a flowchart for an example describing the process of decoding processing.

[0045] Figure 29 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0046] Figure 30 It is a diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0047] Figure 31 A diagram showing an example of the ratio of the area of a coded block to the area of a CTU.

[0048] Figure 32 A diagram showing a comparison example of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each method.

[0049] Figure 33 A flowchart showing an example of the process for describing the encoding process.

[0050] Figure 34 A flowchart showing an example of the process for describing the decoding process.

[0051] Figure 35 A diagram showing an example of the state of binarization of an adaptive orthogonal transform identifier.

[0052] Figure 36 A diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0053] Figure 37 A diagram showing a comparison example of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each method.

[0054] Figure 38 A flowchart showing an example of the process for describing the encoding process.

[0055] Figure 39 A flowchart showing an example of the process for describing the decoding process.

[0056] Figure 40 A diagram showing an example of the syntax regarding a transform unit.

[0057] Figure 41 A diagram showing an example of the syntax regarding an orthogonal transform mode.

[0058] Figure 42 A diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0059] Figure 43 A diagram showing an example of assigning context variables to each bin in a bin sequence of an adaptive orthogonal transform identifier.

[0060] Figure 44 A diagram showing a comparison example of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each method.

[0061] Figure 45 A flowchart showing an example of the process for describing the encoding process.

[0062] Figure 46 is a flowchart showing an example of a process for describing decoding processing.

[0063] Figure 47 is a block diagram showing an example of the main configuration of a computer. Detailed implementation manners

[0064] Hereinafter, modes for implementing the present disclosure (hereinafter referred to as embodiments) will be described. Note that the description will be given in the following order.

[0065] 1. Encoding of adaptive orthogonal transform identifiers

[0066] 2. First embodiment

[0067] 3. Second embodiment

[0068] 4. Third embodiment

[0069] 5. Fourth embodiment

[0070] 6. Fifth embodiment

[0071] 7. Sixth embodiment

[0072] 8. Seventh embodiment

[0073] 9. Appendix

[0074] <1. Encoding of adaptive orthogonal transform identifiers>

[0075] <1-1. Literature etc. supporting technical content and technical terms>

[0076] The scope disclosed in the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent literature etc. known at the time of filing the application and the content of other literatures referred to in the following non-patent literature.

[0077] Non-patent literature 1: (as described above).

[0078] Non-patent literature 2: ITU-T Recommendation H.264 (04 / 2017) “Advanced video coding for generic audiovisual services”, April 2017.

[0079] Non-patent literature 3: ITU-T Recommendation H.265 (12 / 2016) “High efficiency video coding”, April 2016.

[0080] Non-Patent Document 4: J. Chen, E. Alshina, G. J. Sullivan, J. R. Ohm, and J. Boyce, "Algorithm Description of Joint Exploration Test Model (JEM7)", JVET-G1001, 7th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Turin, Italy, July 13, 2017 - July 21, 2017.

[0081] Non-Patent Document 5: Bross, J. Chen, and S. Liu, "Versatile Video Coding (Draft 3)", JVET-L1001, 12th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Macau, China, October 3, 2018 - October 12, 2018.

[0082] Non-Patent Document 6: J. Chen, Y. Ye, and S. Kim, "Algorithm description for Versatile Video Coding and Test Model 3 (VTM 3)", JVET-L1002, 12th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Macau, China, October 3, 2018 - October 12, 2018.

[0083] Non-Patent Document 7: Jianle Chen, Yan Ye, and Seung Hwan Kim: "Algorithm description for Versatile Video Coding and Test Model 5 (VTM 5)", JVET-N1002-v2, 14th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 - March 27, 2019.

[0084] Non-Patent Document 8: Moonmo Koo, Jaehyun Lim, Mehdi Salehifar, and Seung Hwan Kim, "CE6: Reduced Secondary Transform (RST) (CE6-3.1)", JVET-N0193, 14th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.

[0085] Non-Patent Document 9: Mischa Siekmann, Martin Winken, Heiko Schwarz, and Detlev Marpe, "CE6-related: Simplification of the Reduced Secondary Transform", JVET-N0555-v3, 14th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.

[0086] Non-Patent Document 10: C. Rosewarne and J. Gan, "CE6-related: RST binarization", JVET-N0105-v2, 14th Meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.

[0087] That is to say, the content described in the above non-patent documents is also used as the basis for determining the support requirements. For example, even if the quadtree block structure and the quadtree plus binary tree (QTBT) block structure described in the above non-patent documents are not directly described in the examples, these contents fall within the scope of the disclosure of the present technology and meet the support requirements of the claims. In addition, for example, even if technical terms such as parsing, syntax, and semantics are not directly described in the examples, these technical terms similarly fall within the scope of the disclosure of the present technology and meet the support requirements of the claims.

[0088] In addition, in this specification, unless otherwise specified, a "block" (not the block indicating a processing unit) used to describe a partial area or a processing unit of an image (picture) indicates an arbitrary partial area in the picture, and the size, shape, characteristics, etc. of the block are not restricted. For example, a "block" includes any partial area (processing unit) such as a transform block (TB), a transform unit (TU), a prediction block (PB), a prediction unit (PU), a minimum coding unit (SCU), a coding unit (CU), a maximum coding unit (LCU), a coding tree block (CTB), a coding tree unit (CTU), a transform block, a sub-block, a macro-block, a tile, or a slice described in the above non-patent documents.

[0089] In addition, when specifying the size of such a block, not only can the block size be directly specified, but the block size can also be indirectly specified. For example, the block size can be specified using identification information for identifying the size. In addition, for example, the block size can be specified by a ratio or difference from the size of a reference block (e.g., LCU, SCU, etc.). For example, when the information for specifying the block size is sent as a syntax element or the like, the information for indirectly specifying the size as described above can be used as this information. With this configuration, the amount of information can be reduced, and in some cases, the coding efficiency can be improved. In addition, the specification of the block size also includes the specification of the range of the block size (e.g., the specification of the allowable range of the block size, etc.).

[0090] In addition, in this specification, encoding includes not only the entire process of transforming an image into a bitstream, but also a part of this process. For example, encoding includes not only a process including prediction processing, orthogonal transformation, quantization, arithmetic coding, etc., but also a process collectively referred to as quantization and arithmetic coding, a process including prediction processing, quantization, and arithmetic coding, and so on. Similarly, decoding includes not only the entire process of transforming a bitstream into an image, but also a part of this process. For example, decoding includes not only a process including inverse arithmetic decoding, inverse quantization, inverse orthogonal transformation, prediction processing, etc., but also a process including inverse arithmetic decoding and inverse quantization, a process including inverse arithmetic decoding, inverse quantization, and prediction processing, and so on.

[0091] <1-2. Context Encoding / Context Decoding of Adaptive Orthogonal Transformation Identifier>

[0092] Traditionally, in image encoding and decoding, the adaptive orthogonal transformation identifier mts_idx has been signaled (encoded / decoded) as mode information regarding adaptive orthogonal transformation (multiple transform selection (MTS)). Context encoding using the following context has been applied to the encoding of the adaptive orthogonal transformation identifier mts_idx.

[0093] First, as Figure 1In the table shown in A, the adaptive orthogonal transform identifier mts_idx is binarized by a truncated unary (TU) code to obtain a bin sequence bins. Note that the TU code is equivalent to a truncated Rice (TR) code with a Rice parameter cRiceParam = 0.

[0094] Next, arithmetic coding is performed with reference to context variables ctx corresponding to each binIdx (index indicating the bin number) in the bin sequence bins obtained by TR. The index used to identify the context variable ctx is called ctxInc (or ctxIdx).

[0095] Specifically, as Figure 1 In the table shown in B, context variables ctx corresponding one-to-one to the values of the CQT partition depth cqtDepth are assigned to the first bin (binIdx = 0) in the bin sequence. The CQT partition depth cqtDepth represents the depth of the division of the CTU by the quadtree of the CU. In Figure 1 In the example in B, the smaller of the CQT partition depth cqtDepth and 5 is set as the context index ctxInc (ctxInc = min(cqtDepth, 5)) corresponding to the first bin (binIdx = 0) in the bin sequence. That is, since the output (frequency) of 1 of order 0 in the adaptive orthogonal transform may change according to the partition depth, for efficiency, the context is also variably corresponding.

[0096] In addition, context variables ctx corresponding one-to-one to each binIdx (in Figure 1 the example in B, ctxInc = 6 to 8) are assigned to the second bin (binIdx = 1) to the fourth bin (binIdx = 3) in the bin sequence bins.

[0097] Note that, as Figure 2 In the table shown in A, each bin in the bin sequence bins of the adaptive orthogonal transform identifier mts_idx can be interpreted as a flag corresponding to the transform type. In this example, the value of the first bin (binIdx = 0) corresponds to a flag indicating whether the transform type is DCT2×DCT2 (0 indicates "yes" and 1 indicates "no"), the value of the second bin (binIdx = 1) corresponds to a flag indicating whether the transform type is DST7×DST7 (0 indicates "yes" and 1 indicates "no"), the value of the third bin (binIdx = 2) corresponds to a flag indicating whether the transform type is DCT8×DST7 (0 indicates "yes" and 1 indicates "no"), and the value of the fourth bin (binIdx = 3) corresponds to a flag indicating whether the transform type is DST7×DCT8 (0 indicates "yes" and 1 indicates "no").

[0098] The encoded data of the adaptive orthogonal transform identifier mts_idx has been decoded by a method corresponding to such encoding. That is, context decoding using context has been applied.

[0099] However, in the case of such context encoding and context decoding, there is a possibility of an increase in processing load.

[0100] For example, the adaptive orthogonal transform identifier mts_idx does not appear in the case of a specific value of the CQT partition depth cqtDepth. Therefore, there are context variables ctx that are not used at all, and there is a possibility of an unnecessary increase in memory usage due to the context variables ctx (there is a possibility of an increase in the memory capacity required for processing).

[0101] For example, in the case where all CU partitions are performed using a quadtree, when the CTU size = 128×128, the CU sizes corresponding to each CQT partition depth cqtDepth are as Figure 2 shown in the table in B. Since the adaptive orthogonal transform is not applicable to blocks larger than 32×32, the adaptive orthogonal transform identifier mts_idx does not appear for 128×128 and 64×64 CUs. Therefore, in the context variables ctx corresponding to the first bin in the bin sequence of the adaptive orthogonal transform identifier mts_idx, ctxInc = 0 and 1 are not used at all. That is, due to these context variables ctx, there is a possibility of an unnecessary increase in memory usage.

[0102] Therefore, a predetermined context variable is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0103] By doing so, an increase in the number of contexts assigned to the bin sequence of the adaptive orthogonal transform identifier can be suppressed, and thus an increase in memory usage can be suppressed and an increase in the load of the encoding process and the encoding process can be suppressed.

[0104] In addition, a context variable based on a parameter regarding the block size is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0105] By doing so, an increase in the number of contexts assigned to the bin sequence of the adaptive orthogonal transform identifier can be suppressed, and thus an increase in memory usage can be suppressed and an increase in the load of the encoding process and the encoding process can be suppressed.

[0106] <1-3. Bypass Coding>

[0107] In addition, the selectivity of the transform types in descending order is Discrete Cosine Transform (DCT) 2×DCT2, Discrete Sine Transform (DST) 7×DST7, DCT8×DST7, DST7×DCT8, or DCT8×DCT8. That is, the selectivity of each transform type is not uniform. Therefore, it is inefficient to similarly assign the context variable ctx to all transform types, which may unnecessarily increase the total number of context coding bins. As described above, each bin in the bin sequence bins of the adaptive orthogonal transform identifier mts_idx can be interpreted as a flag corresponding to the transform type. That is, it is inefficient to similarly assign the context variable ctx to each bin in the bin sequence bins, which may unnecessarily increase the total number of context coding bins.

[0108] As the total number of context coding bins increases in this way, there is a possibility of an increase in the processing amount (throughput) of Context-based Adaptive Binary Arithmetic Coding (CABAC).

[0109] Therefore, bypass coding is applied to the bins corresponding to the transform types with relatively low selectivity. By doing so, it is possible to suppress an increase in the number of context coding bins and an increase in the processing amount (throughput) of CABAC while suppressing a decrease in coding efficiency. That is, it is possible to suppress an increase in the load of the encoding process and the decoding process.

[0110] <2. First Embodiment>

[0111] <2-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0112] In this embodiment, the context variable is assigned to each bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding as already performed in the table shown in Figure 1 B (Method 0).

[0113] That is, a predetermined context variable (fixed (one-to-one) context variable) ctx is assigned to the first bin in the bin sequence, and context coding is performed on the first bin (Method 1).

[0114] For example, as shown in the table in Figure 3 A, a predetermined context variable (index ctxInc for identifying the context variable ctx) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, context coding can be performed on the first bin, and bypass coding can be performed on the second to fourth bins in the bin sequence (Method 1-1).

[0115] InFigure 3 In the case of the example in A, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context encoding is performed on the first bin, and bypass encoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0116] In addition, for example, as Figure 3 shown in the table in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first bin and the second bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, context encoding can be performed on the first bin and the second bin, and bypass encoding can be performed on the third bin and the fourth bin in the bin sequence (Method 1-2).

[0117] In Figure 3 the case of the example in B, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), context encoding is performed on the second bin, and bypass encoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0118] In addition, for example, as Figure 4 shown in the table in A, different predetermined context variables ctx (index ctxInc) can be assigned to the first bin to the third bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, context encoding can be performed on the first bin to the third bin, and bypass encoding can be performed on the fourth bin in the bin sequence (Method 1-3).

[0119] In Figure 4 the case of the example in A, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), context encoding is performed on the second bin, the index ctxInc = B2 is assigned to the third bin (binIdx = 2), context encoding is performed on the third bin, and bypass encoding is performed on the fourth bin (binIdx = 3).

[0120] In addition, for example, as Figure 4In the table shown in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first to fourth bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding (methods 1-4) can be performed on the first to fourth bins.

[0121] In Figure 4 In the case of the example in B, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.

[0122] Note that in Figure 3 and Figure 4 In the table, non-overlapping unique values are set in the indexes A0, B1, B2, and B3.

[0123] Figure 5 The table shows examples of the number of contexts, the number of context-encoded bins, and the number of bypass-encoded bins for each of these methods. For example, in the case of method 0, the number of contexts is 9, the number of context-encoded bins is 4, and the number of bypass-encoded bins is 0. In contrast, in the case of method 1-1, the number of contexts is 1, the number of context-encoded bins is 1, and the number of bypass-encoded bins is 3. In addition, in the case of method 1-2, the number of contexts is 2, the number of context-encoded bins is 2, and the number of bypass-encoded bins is 2. In addition, in the case of method 1-3, the number of contexts is 3, the number of context-encoded bins is 3, and the number of bypass-encoded bins is 1. In addition, in the case of method 1-4, the number of contexts is 4, the number of context-encoded bins is 4, and the number of bypass-encoded bins is 0.

[0124] As described above, in any of the cases of methods 1-1 to 1-4, compared with the case of method 0, the number of contexts required for encoding can be reduced. That is, by applying method 1, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0125] In addition, in any of the cases of Method 1-1 to Method 1-3, the number of context coding bins required for coding can be reduced compared to the case of Method 0. Note that in the case of Method 1-4, the number of context coding bins required for coding is equivalent to the number of context coding bins required for coding in the case of Method 0. That is, by applying Method 1, bypass coding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in coding efficiency and suppressing an increase in the processing amount (throughput).

[0126] As described above, by applying Method 1, an increase in the load of the coding process can be suppressed.

[0127] <2-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0128] Similarly, in the case of decoding, context variables are assigned to each bin in the bin sequence of the binarized adaptive orthogonal transform identifier indicating the mode of inverse adaptive orthogonal transform in image decoding as already performed in the table shown in Figure 1 B.

[0129] That is, a predetermined context variable (a fixed (one-to-one) context variable) ctx is assigned to the first bin in the bin sequence, and context decoding (Method 1) is performed on the first bin.

[0130] For example, as shown in the table in Figure 3 A, a predetermined context variable (an index ctxInc for identifying the context variable ctx) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, context decoding can be performed on the first bin, and bypass decoding (Method 1-1) can be performed on the second to fourth bins in the bin sequence.

[0131] In Figure 3 the example in A, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context decoding is performed on the first bin, and bypass decoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0132] In addition, for example, as shown in the table in Figure 3 B, different predetermined context variables ctx (indexes ctxInc) can be assigned to the first and second bins in the bin sequence of the binarized adaptive orthogonal transform identifier, context decoding can be performed on the first and second bins, and bypass decoding (Method 1-2) can be performed on the third and fourth bins in the bin sequence.

[0133] In Figure 3 the case of the example in B of

[0133] , the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. And bypass decoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0134] In addition, for example, as shown in the table in A of Figure 4

[0133] , different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first to third bins, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 1-3).

[0135] In Figure 4 the case of the example in A of

[0133] , the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And bypass decoding is performed on the fourth bin (binIdx = 3).

[0136] In addition, for example, as shown in the table in B of Figure 4

[0133] , different predetermined context variables ctx (index ctxInc) can be assigned to the first to fourth bins in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first to fourth bins (Method 1-4).

[0137] In Figure 4 the case of the example in B of

[0133] , the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.

[0138] Note that, even in the case of decoding, similar to the case of encoding, non-overlapping unique values are also set in the indexes A0, B1, B2, and B3 in the tables of Figure 3 and Figure 4 .

[0139] The number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are similar to the case of encoding ( Figure 5 ).

[0140] As described above, in any of the cases of Method 1-1 to Method 1-4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 1, the number of contexts allocated to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0141] In addition, in any of the cases of Method 1-1 to Method 1-3, compared with the case of Method 0, the number of context encoding bins required for decoding can be reduced. Note that, in the case of Method 1-4, the number of context encoding bins required for decoding is comparable to the number of context encoding bins required for decoding in the case of Method 0. That is, by applying Method 1, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context encoding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.

[0142] As described above, by applying Method 1, an increase in the load of decoding processing can be suppressed.

[0143] <2-3. Encoding side>

[0144] <Image encoding device>

[0145] Next, the encoding side will be described. Figure 6 is a block diagram showing an example of the configuration of an image encoding device as a mode of an image processing device to which the present technology is applied. Figure 6 The image encoding device 100 shown in encodes the image data of a moving image. For example, the image encoding device 100 encodes the image data of a moving image by an encoding method described in any one of Non-Patent Documents 1 to 10.

[0146] Note that, Figure 6 shows the main processing units (blocks), data flows, etc., and Figure 6 these shown in are not necessarily all. That is, in the image encoding device 100, there may be processing units that are not shown as blocks in Figure 6 , or there may be processing or data flows that are not shown as arrows, etc. in Figure 6 .

[0147] As Figure 6 shown, the image encoding apparatus 100 includes a control unit 101, a rearrangement buffer 111, a calculation unit 112, an orthogonal transformation unit 113, a quantization unit 114, an encoding unit 115, an accumulation buffer 116, an inverse quantization unit 117, an inverse orthogonal transformation unit 118, a calculation unit 119, a loop filter unit 120, a frame memory 121, a prediction unit 122, and a rate control unit 123.

[0148] <Control Unit>

[0149] The control unit 101 divides the moving image data held by the rearrangement buffer 111 into blocks (CU, PU, transform block, etc.) of a processing unit based on the block size of an external or pre-specified processing unit. Further, the control unit 101 determines encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filtering information Finfo, etc.) to be provided to each block based on, for example, rate distortion optimization (RDO).

[0150] Details of these encoding parameters will be described below. After determining the above encoding parameters, the control unit 101 provides the encoding parameters to each block. Specifically, the encoding parameters are as follows.

[0151] The header information Hinfo is provided to each block.

[0152] The prediction mode information Pinfo is provided to the encoding unit 115 and the prediction unit 122.

[0153] The transform information Tinfo is provided to the encoding unit 115, the orthogonal transformation unit 113, the quantization unit 114, the inverse quantization unit 117, and the inverse orthogonal transformation unit 118.

[0154] The filtering information Finfo is provided to the loop filter unit 120.

[0155] <Rearrangement Buffer>

[0156] Each field (input image) of the moving image data is input to the image encoding apparatus 100 in the reproduction order (display order). The rearrangement buffer 111 acquires and holds (stores) each input image in its reproduction order (display order). The rearrangement buffer 111 rearranges the input images in the encoding order (decoding order) or divides the input images into blocks of a processing unit under the control of the control unit 101. The rearrangement buffer 111 provides the processed input images to the calculation unit 112. Further, the rearrangement buffer 111 also provides the input images (original images) to the prediction unit 122 and the loop filter unit 120.

[0157] <Calculation Unit>

[0158] The calculation unit 112 receives the image I corresponding to the block of the processing unit and the predicted image P provided from the prediction unit 122 as inputs, subtracts the predicted image P from the image I as shown in the following expression to derive the prediction residual D, and provides the prediction residual D to the orthogonal transformation unit 113.

[0159] D = I - P.

[0160] <Orthogonal transformation unit>

[0161] The orthogonal transformation unit 113 uses the prediction residual D provided from the calculation unit 112 and the transformation information Tinfo provided from the control unit 101 as inputs, and performs an orthogonal transformation on the prediction residual D based on the transformation information Tinfo to derive the transformation coefficient Coeff. Note that the orthogonal transformation unit 113 can perform an adaptive orthogonal transformation for adaptively selecting the type of orthogonal transformation (transformation coefficient). The orthogonal transformation unit 113 provides the obtained transformation coefficient Coeff to the quantization unit 114.

[0162] <Quantization unit>

[0163] The quantization unit 114 uses the transformation coefficient Coeff provided from the orthogonal transformation unit 113 and the transformation information Tinfo provided from the control unit 101 as inputs, and scales (quantizes) the transformation coefficient Coeff based on the transformation information Tinfo. Note that the quantization rate is controlled by the rate control unit 123. The quantization unit 114 provides the quantized transformation coefficient obtained by quantization (i.e., the quantization transformation coefficient level "level") to the encoding unit 115 and the inverse quantization unit 117.

[0164] <Encoding unit>

[0165] The encoding unit 115 uses the quantization transformation coefficient level "level" provided from the quantization unit 114, various encoding parameters (header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, filtering information Finfo, etc.) provided from the control unit 101, information about the filter (e.g., filter coefficients) provided from the loop filter unit 120, and information about the best prediction mode provided from the prediction unit 122 as inputs. The encoding unit 115 performs variable length encoding (e.g., arithmetic encoding) on the quantization transformation coefficient level "level" to generate a bit string (encoded data).

[0166] In addition, the encoding unit 115 derives the residual information Rinfo from the quantization transformation coefficient level "level", and encodes the residual information Rinfo to generate a bit string.

[0167] In addition, the encoding unit 115 includes the information about the filter provided by the loop filter unit 120 into the filtering information Finfo, and includes the information about the best prediction mode provided by the prediction unit 122 into the prediction mode information Pinfo. Then, the encoding unit 115 encodes the above various encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filtering information Finfo, etc.) to generate a bit string.

[0168] In addition, the encoding unit 115 multiplexes the bit strings of various types of information generated as described above to generate encoded data. The encoding unit 115 provides the encoded data to the accumulation buffer 116.

[0169] <Accumulation buffer>

[0170] The accumulation buffer 116 temporarily stores the encoded data obtained by the encoding unit 115. The accumulation buffer 116 outputs the stored encoded data as a bit stream or the like to the outside of the image encoding apparatus 100 at a predetermined timing. For example, the encoded data is sent to the decoding side via an arbitrary recording medium, an arbitrary transmission medium, an arbitrary information processing apparatus, or the like. That is, the accumulation buffer 116 is also a transmission unit that transmits the encoded data (bit stream).

[0171] <Inverse quantization unit>

[0172] The inverse quantization unit 117 performs a process of inverse quantization. For example, the inverse quantization unit 117 uses the quantization transform coefficient level "level" provided by the quantization unit 114 and the transform information Tinfo provided by the control unit 101 as inputs, and scales (inverse quantizes) the value of the quantization transform coefficient level "level" based on the transform information Tinfo. Note that inverse quantization is an inverse process of the quantization performed in the quantization unit 114.

[0173] The inverse quantization unit 117 provides the transform coefficient Coeff_IQ obtained by the inverse quantization to the inverse orthogonal transform unit 118. Note that since the inverse orthogonal transform unit 118 is similar to the inverse orthogonal transform unit on the decoding side (which will be described below), the description given for the decoding side (which will be described below) can be applied to the inverse quantization unit 117.

[0174] <Inverse orthogonal transform unit>

[0175] The inverse orthogonal transform unit 118 performs processing regarding the inverse orthogonal transform. For example, the inverse orthogonal transform unit 118 uses the transform coefficients Coeff_IQ provided from the inverse quantization unit 117 and the transform information Tinfo provided from the control unit 101 as inputs, and performs an inverse orthogonal transform on the transform coefficients Coeff_IQ based on the transform information Tinfo to derive a prediction residual D'. Note that the inverse orthogonal transform is the inverse process of the orthogonal transform performed in the orthogonal transform unit 113. That is, the inverse orthogonal transform unit 118 can perform an adaptive inverse orthogonal transform for adaptively selecting the type (transform coefficients) of the inverse orthogonal transform.

[0176] The inverse orthogonal transform unit 118 provides the prediction residual D' obtained through the inverse orthogonal transform to the calculation unit 119. Note that since the inverse orthogonal transform unit 118 is similar to the inverse orthogonal transform unit on the decoding side (which will be described below), the description given for the decoding side (which will be described below) can be applied to the inverse orthogonal transform unit 118.

[0177] <Calculation Unit>

[0178] The calculation unit 119 uses the prediction residual D' provided from the inverse orthogonal transform unit 118 and the predicted image P provided from the prediction unit 122 as inputs. The calculation unit 119 adds the prediction residual D' and the predicted image P corresponding to the prediction residual D' to derive a local decoded image Rlocal. The calculation unit 119 provides the derived local decoded image Rlocal to the loop filter unit 120 and the frame memory 121.

[0179] <Loop Filter Unit>

[0180] The loop filter unit 120 performs processing regarding loop filter processing. For example, the loop filter unit 120 uses the local decoded image Rlocal provided from the calculation unit 119, the filter information Finfo provided from the control unit 101, and the input image (original image) provided from the rearrangement buffer 111 as inputs. Note that the information input to the loop filter unit 120 is arbitrary and can include information other than the above-mentioned information. For example, information such as the prediction mode, motion information, code amount target value, quantization parameter QP, picture type, block (CU, CTU, etc.) can be input to the loop filter unit 120 as needed.

[0181] The loop filter unit 120 appropriately performs a filtering process on the local decoded image Rlocal based on the filter information Finfo. The loop filter unit 120 also uses the input image (original image) and other input information for the filtering process as needed.

[0182] For example, the loop filter unit 120 applies these four loop filters in the order of a bilateral filter, a deblocking filter (DBF), an adaptive offset filter (sample adaptive offset (SAO)), and an adaptive loop filter (adaptive loop filter (ALF)). Note that which filter to apply and in what order to apply the filters are arbitrary and can be appropriately selected.

[0183] Of course, the filtering process performed by the loop filter unit 120 is arbitrary and is not limited to the above example. For example, the loop filter unit 120 may apply a Wiener filter or the like.

[0184] The loop filter unit 120 provides the filtered local decoded image Rlocal to the frame memory 121. Note that in the case of sending information about the filter (e.g., filter coefficients) to the decoding side, the loop filter unit 120 provides the information about the filter to the encoding unit 115.

[0185] <Frame Memory>

[0186] The frame memory 121 performs a process of storing data related to the image. For example, the frame memory 121 uses the local decoded image Rlocal provided from the calculation unit 119 and the filtered local decoded image Rlocal provided from the loop filter unit 120 as inputs and holds (stores) the inputs. In addition, the frame memory 121 reconstructs and holds the decoded image R for each picture unit using the local decoded image Rlocal (stores the decoded image R in a buffer in the frame memory 121). The frame memory 121 provides the decoded image R (or a part thereof) to the prediction unit 122 in response to a request from the prediction unit 122.

[0187] <Prediction Unit>

[0188] The prediction unit 122 performs a process of generating a prediction image. For example, the prediction unit 122 uses the prediction mode information Pinfo provided from the control unit 101, the input image (original image) provided from the rearrangement buffer 111, and the decoded image R (or a part thereof) read from the frame memory 121 as inputs. The prediction unit 122 performs prediction processes such as inter-frame prediction and intra-frame prediction using the prediction mode information Pinfo and the input image (original image), performs prediction using the decoded image R as a reference image, performs motion compensation processing based on the prediction result, and generates a prediction image P. The prediction unit 122 provides the generated prediction image P to the calculation unit 112 and the calculation unit 119. In addition, the prediction unit 122 provides the prediction mode (i.e., information about the best prediction mode) selected through the above processing to the encoding unit 115 as needed.

[0189] <Rate control unit>

[0190] The rate control unit 123 performs processing related to rate control. For example, the rate control unit 123 controls the rate of the quantization operation of the quantization unit 114 based on the code amount of the encoded data accumulated in the accumulation buffer 116 so that overflows or underflows do not occur.

[0191] Note that these processing units (the control unit 101 and the rearrangement buffer 111 to the rate control unit 123) have arbitrary configurations. For example, each processing unit can be configured by a logic circuit that implements the above processing. In addition, each processing unit can include, for example, a central processing unit (CPU), a read-only memory (ROM), a random access memory (RAM), etc., and implements the above processing by executing a program using the above resources. Of course, each processing unit can have both of these configurations, and implements a part of the above processing by a logic circuit and implements another part of the processing by executing a program. The configurations of the processing units can be independent of each other. For example, some of the processing units can implement a part of the above processing by a logic circuit, some of the processing units can implement the above processing by executing a program, and some of the processing units can implement the above processing by both a logic circuit and the execution of a program.

[0192] <Encoding unit>

[0193] Figure 7 is a block diagram showing Figure 6 a main configuration example of the encoding unit 115 in. As Figure 7 shown, the encoding unit 115 includes a binarization unit 131, a selection unit 132, a context setting unit 133, a context encoding unit 134, and a bypass encoding unit 135.

[0194] Note that although the encoding of the adaptive orthogonal transform identifier is described here, the encoding unit 115 also encodes other encoding parameters, residual information Rinfo, etc. as described above. The encoding unit 115 encodes the adaptive orthogonal transform identifier by applying Method 1 described in <2-1. Encoding of Adaptive Orthogonal Transform Identifier>.

[0195] The binarization unit 131 performs processing related to the binarization of the adaptive orthogonal transform identifier. For example, the binarization unit 131 acquires the adaptive orthogonal transform identifier mts_idx provided from the control unit 101. In addition, the binarization unit 131 binarizes the adaptive orthogonal transform identifier using a truncated unary code (or a truncated Rice code) to generate a bin sequence. In addition, the binarization unit 131 provides the generated bin sequence to the selection unit 132.

[0196] The selection unit 132 performs processing for selecting a supply destination for each bin in the bin sequence of the adaptive orthogonal transform identifier. For example, the selection unit 132 acquires the bin sequence of the adaptive orthogonal transform identifier provided from the binarization unit 131.

[0197] In addition, the selection unit 132 selects, for each bin in the acquired bin sequence of the adaptive orthogonal transform identifier, whether to set the supply destination to the context setting unit 133 or the bypass encoding unit 135. The selection unit 132 performs the selection according to Method 1 described above in <2-1. Encoding of Adaptive Orthogonal Transform Identifier>. For example, the selection unit 132 may perform the selection according to Method 1-1 (i.e., Figure 3 the table shown in A of Figure 3 ). In addition, the selection unit 132 may perform the selection according to Method 1-2 (i.e., Figure 4 the table shown in B of Figure 4 ). In addition, the selection unit 132 may perform the selection according to Method 1-3 (i.e.,

[0198] the table shown in A of

[0199] ). In addition, the selection unit 132 may perform the selection according to Method 1-4 (i.e., Figure 3 the table shown in B of Figure 3 ). Figure 4 In the case of allocating a context variable (index ctxInc) and performing context encoding, the selection unit 132 supplies the bin to the context setting unit 133. In the case of bypass encoding, the selection unit 132 supplies the bin to the bypass encoding unit 135. Figure 4 The context setting unit 133 performs processing for context setting. For example, the context setting unit 133 acquires the bin provided from the selection unit 132. The context setting unit 133 allocates a context variable (index ctxInc) to the bin. The context setting unit 133 performs the allocation according to Method 1 described above in <2-1. Encoding of Adaptive Orthogonal Transform Identifier>. For example, the context setting unit 133 may perform the allocation according to Method 1-1 (i.e.,

[0200] In addition, the context setting unit 133 can obtain the encoding result from the context encoding unit 134. The context setting unit 133 can appropriately update the context variable (index ctxInc) using the encoding result. The context setting unit 133 provides the context variable (index ctxInc) derived in this way to the context encoding unit 134.

[0201] The context encoding unit 134 performs processing regarding arithmetic coding. For example, the context encoding unit 134 obtains the context variable (index ctxInc) provided from the context setting unit 133. In addition, the context encoding unit 134 performs arithmetic coding using the context variable (index ctxInc). That is, context encoding is performed. In addition, the context encoding unit 134 provides the encoding result as encoded data to the accumulation buffer 116.

[0202] The bypass encoding unit 135 performs processing regarding bypass coding. For example, the bypass encoding unit 135 obtains the bin provided from the selection unit 132. The bypass encoding unit 135 performs bypass coding (arithmetic coding) on the bin. The bypass encoding unit 135 provides the encoding result as encoded data to the accumulation buffer 116.

[0203] Since each of the processing units (the binarization unit 131 to the bypass encoding unit 135) performs the processing as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 1 (for example, any one of Method 1-1 to Method 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0204] <Flow of Image Encoding Processing>

[0205] Next, an example of the flow of the image encoding processing performed by the image encoding apparatus 100 having the above configuration will be described with reference to Figure 8 the flowchart.

[0206] When the image encoding processing starts, in step S101, the rearrangement buffer 111 is controlled by the control unit 101, and the frames of the input moving image data are rearranged from the display order to the encoding order.

[0207] In step S102, the control unit 101 sets the processing unit for the input image held by the rearrangement buffer 111 (performs block partitioning).

[0208] In step S103, the control unit 101 determines (sets) the encoding parameters for the input image held by the rearrangement buffer 111.

[0209] In step S104, the prediction unit 122 performs prediction processing and generates a prediction image of the best prediction mode, etc. For example, in the prediction processing, the prediction unit 122 performs intra prediction to generate a prediction image of the best intra prediction mode, performs inter prediction to generate a prediction image of the best inter prediction mode, and selects the best prediction mode from the prediction images based on the cost function value, etc.

[0210] In step S105, the calculation unit 112 calculates the difference between the input image and the prediction image of the best mode selected by the prediction processing in step S104. That is, the calculation unit 112 generates a prediction residual D between the input image and the prediction image. The prediction residual D obtained in this way has a reduced data volume compared to the original image data. Therefore, the data volume can be compressed compared to the case of encoding the image as it is.

[0211] In step S106, the orthogonal transformation unit 113 performs an orthogonal transformation process on the prediction residual D generated by the process in step S105 to derive the transformation coefficient Coeff.

[0212] In step S107, the quantization unit 114 quantizes the transformation coefficient Coeff obtained by the process in step S106 using the quantization parameter calculated by the control unit 101, etc., to derive the quantized transformation coefficient level "level".

[0213] In step S108, the inverse quantization unit 117 performs inverse quantization on the quantized transformation coefficient level "level" generated by the process in step S107 using the characteristics corresponding to the quantization in step S107 to derive the transformation coefficient Coeff_IQ.

[0214] In step S109, the inverse orthogonal transformation unit 118 performs an inverse orthogonal transformation on the transformation coefficient Coeff_IQ obtained by the process in step S108 by a method corresponding to the orthogonal transformation process in step S106 to derive the prediction residual D'. Note that since the inverse orthogonal transformation process is similar to the inverse orthogonal transformation process performed on the decoding side (which will be described below), the description of the decoding side (which will be given below) can be applied to the inverse orthogonal transformation process in step S109.

[0215] In step S110, the calculation unit 119 adds the prediction image obtained by the prediction processing in step S104 to the prediction residual D' derived by the process in step S109 to generate a locally decoded image.

[0216] In step S111, the loop filter unit 120 performs a loop filter process on the locally decoded image derived by the process in step S110.

[0217] In step S112, the frame memory 121 stores the locally decoded image derived through the processing in step S110 and the locally decoded image filtered in step S111.

[0218] In step S113, the encoding unit 115 encodes the quantization transform coefficient level "level" obtained through the processing in step S107. For example, the encoding unit 115 encodes the quantization transform coefficient level "level", which is information about the image, through arithmetic coding or the like to generate encoded data. In addition, at this time, the encoding unit 115 encodes various encoding parameters (header information Hinfo, prediction mode information Pinfo, and transform information Tinfo). Further, the encoding unit 115 derives residual information RInfo from the quantization transform coefficient level "level" and encodes the residual information RInfo.

[0219] In step S114, the accumulation buffer 116 accumulates the encoded data thus obtained and outputs the encoded data, for example, as a bitstream to the outside of the image encoding device 100. For example, the bitstream is sent to the decoding side via a transmission path or a recording medium. In addition, the rate control unit 123 performs rate control as needed.

[0220] When the processing in step S114 ends, the image encoding process ends.

[0221] <Flow of the encoding process>

[0222] In Figure 8 the encoding process of step S113, the encoding unit 115 encodes the adaptive orthogonal transform identifier mts_idx. At this time, the encoding unit 115 encodes the adaptive orthogonal transform identifier by applying Method 1 described in <2-1. Encoding of the adaptive orthogonal transform identifier>. An example of the flow for encoding the adaptive orthogonal transform identifier mts_idx will be described with reference to the flowchart in Figure 9 below.

[0223] When the encoding process starts, in step S131, the binarization unit 131 of the encoding unit 115 binarizes the adaptive orthogonal transform identifier mts_idx through a truncated unary code (or a truncated Rice code) to generate a bin sequence.

[0224] In step S132, the selection unit 132 sets the first bin (binIdx = 0) in the bin sequence as the bin to be processed. In this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (i.e., selects context coding). For example, the selection unit 132 selects according to Figure 3 and Figure 4Select context encoding as the encoding method for the bin from any one of the tables shown (i.e., by applying any one of methods 1-1 to 1-4).

[0225] In step S133, the context setting unit 133 assigns a predetermined context variable ctx (index ctxInc) to the bin. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.

[0226] In step S134, the selection unit 132 sets the unprocessed bins in the second and subsequent bins in the bin sequence as the bins to be processed. In step S135, the selection unit 132 determines whether to perform bypass encoding on the bins to be processed. For example, the selection unit 132 determines whether to perform bypass encoding on the bins to be processed according to Figure 3 and Figure 4 any table shown (by applying any one of methods 1-1 to 1-4).

[0227] In the case where it is determined to perform bypass encoding, the process proceeds to step S136. That is, in this case, the selection unit 132 selects the bypass encoding unit 135 as the bin supply destination. In step S136, the bypass encoding unit 135 performs bypass encoding (arithmetic encoding) on the bins to be processed. When the processing in step S136 is completed, the process proceeds to step S138.

[0228] In addition, in step S135, in the case where it is determined not to perform bypass encoding (perform context encoding), the process proceeds to step S137. That is, in this case, the selection unit 132 selects the context setting unit 133 as the bin supply destination. In step S137, the context setting unit 133 assigns a predetermined context variable ctx (index ctxInc) to the bin. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed. When the processing in step S137 is completed, the process proceeds to step S138.

[0229] In step S138, the encoding unit 115 determines whether to terminate the encoding of the adaptive orthogonal transform identifier mts_idx. In the case where it is determined not to terminate the encoding, the process returns to step S134, and the processing in step S134 and subsequent steps is repeated. In addition, in step S138, in the case where it is determined to terminate the encoding, the encoding process is terminated.

[0230] By performing each of the processes described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 1 (e.g., any one of Methods 1-1 to 1-4). Thus, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0231] <2-4. Decoding side>

[0232] <Image decoding device>

[0233] Next, the decoding side will be described. Figure 10 is a block diagram showing an example of the configuration of an image decoding device as a mode of an image processing device to which the present technology is applied. Figure 10 The image decoding device 200 shown in encodes encoded data of a moving image. For example, the image decoding device 200 decodes the encoded data by a decoding method described in any one of Non-Patent Documents 1 to 10 to generate moving image data. For example, the image decoding device 200 decodes the encoded data (bitstream) generated by the above-described image encoding device 100 to generate moving image data.

[0234] Note that Figure 10 shows main processing units, data flows, etc., and Figure 10 these shown in are not necessarily all. That is, in the image decoding device 200, there may be a processing unit that is not shown as a block in Figure 10 or there may be a process or data flow that is not shown as an arrow, etc. in Figure 10 In

[0235] In Figure 10 the image decoding device 200 includes an accumulation buffer 211, a decoding unit 212, an inverse quantization unit 213, an inverse orthogonal transform unit 214, a calculation unit 215, a loop filter unit 216, a rearrangement buffer 217, a frame memory 218, and a prediction unit 219. Note that the prediction unit 219 includes an intra-frame prediction unit and an inter-frame prediction unit (not shown).

[0236] <Accumulation buffer>

[0237] The accumulation buffer 211 acquires the bitstream input to the image decoding device 200 and holds (stores) the bitstream. For example, the accumulation buffer 211 provides the accumulated bitstream to the decoding unit 212 at a predetermined timing or when a predetermined condition is satisfied.

[0238] <Decoding unit>

[0239] The decoding unit 212 performs processing related to image decoding. For example, the decoding unit 212 uses the bitstream provided from the accumulation buffer 211 as input, and performs variable length decoding on the syntax value of each syntax element from the bitstring according to the definition of the syntax table to derive parameters.

[0240] Parameters derived from the syntax element and the syntax value of the syntax element include, for example, information such as header information Hinfo, prediction mode information Pinfo, transform information Tinfo, residual information Rinfo, and filtering information Finfo. That is, the decoding unit 212 parses (analyzes and obtains) such information from the bitstream. This information will be described below.

[0241] <Header information Hinfo>

[0242] The header information Hinfo includes, for example, header information such as video parameter set (VPS) / sequence parameter set (SPS) / picture parameter set (PPS) / slice header (SH). The header information Hinfo includes, for example, information defining the following: image size (width PicWidth and height PicHeight), bit depth (luminance bitDepthY and chrominance bitDepthC), chrominance array type ChromaArrayType, maximum CU size MaxCUSize / minimum CU size MinCUSize, maximum depth MaxQTDepth / minimum depth MinQTDepth of quadtree partitioning, maximum depth MaxBTDepth / minimum depth MinBTDepth of binary tree partitioning, maximum size MaxTSSize of transform skip blocks (also referred to as maximum transform skip block size), on / off flag for each coding tool (also referred to as enable flag), etc.

[0243] For example, examples of the on / off flag of the coding tool included in the header information Hinfo include on / off flags related to the following transform and quantization processing. Note that the on / off flag of the coding tool can also be interpreted as a flag indicating whether the syntax related to the coding tool exists in the coded data. In addition, when the value of the on / off flag is 1 (true), the value indicates that the coding tool is available. When the value of the on / off flag is 0 (false), the value indicates that the coding tool is not available. Note that the interpretation of the flag value can be swapped.

[0244] The inter-component prediction enable flag (ccp_enabled_flag) is flag information indicating whether inter-component prediction (cross-component prediction (CCP), also referred to as CC prediction) is available. For example, when the flag information is "1" (true), the flag information indicates that inter-component prediction is available. When the flag information is "0" (false), the flag information indicates that inter-component prediction is not available.

[0245] Note that this CCP is also known as Component - to - Component Linear Prediction (CCLM or CCLMP).

[0246] <Prediction mode information Pinfo>

[0247] The prediction mode information Pinfo includes, for example, information such as the size information PBSize (prediction block size) of the prediction block (PB) to be processed, the intra - prediction mode information IPinfo, and the motion prediction information MVinfo.

[0248] The intra - prediction mode information IPinfo includes, for example, prev_intra_luma_pred_flag, mpm_idx, and rem_intra_pred_mode in JCTVC - W1005, 7.3.8.5 coding unit syntax, the luma intra - prediction mode IntraPredModeY derived from the syntax, and so on.

[0249] In addition, the intra - prediction mode information IPinfo includes, for example, the component - to - component prediction flag (ccp_flag (cclmp_flag)), the multi - class linear prediction mode flag (mclm_flag), the chroma sample location type identifier (chroma_sample_loc_type_idx), the chroma MPM identifier (chroma_mpm_idx), the luma intra - prediction mode (IntraPredModeC) derived from these syntaxes, and so on.

[0250] The component - to - component prediction flag (ccp_flag (cclmp_flag)) is flag information indicating whether component - to - component linear prediction is applied. For example, ccp_flag == 1 indicates that component - to - component prediction is applied, while ccp_flag == 0 indicates that component - to - component prediction is not applied.

[0251] The multi - class linear prediction mode flag (mclm_flag) is information about the linear prediction mode (linear prediction mode information). More specifically, the multi - class linear prediction mode flag (mclm_flag) is flag information indicating whether the multi - class linear prediction mode is set. For example, "0" indicates a one - class mode (single - class mode) (e.g., CCLMP), and "1" indicates a two - class mode (multi - class mode) (e.g., MCLMP).

[0252] The chroma sample location type identifier (chroma_sample_loc_type_idx) is an identifier that identifies the type of pixel location of the chroma component (also referred to as the chroma sample location type). For example, when the chroma array type (ChromaArrayType), which is information about the color format, indicates the 420 format, the chroma sample location type identifier is assigned as in the following expressions.

[0253] chroma_sample_loc_type_idx == 0: Type 2

[0254] chroma_sample_loc_type_idx == 1: Type 3

[0255] chroma_sample_loc_type_idx == 2: Type 0

[0256] chroma_sample_loc_type_idx == 3: Type 1

[0257] Note that the chroma sample location type identifier (chroma_sample_loc_type_idx) is sent as information (chroma_sample_loc_info()) about the pixel location of the chroma component (by being stored in this information).

[0258] The chroma MPM identifier (chroma_mpm_idx) is an identifier that indicates which prediction mode candidate in the intra prediction mode candidate list (intraPredModeCandListC) for the chroma frame will be designated as the intra prediction mode for the chroma frame.

[0259] The motion prediction information MVinfo includes information such as, for example, merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X = {0,1}, mvd, etc. (see, for example, JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax).

[0260] Of course, the information included in the prediction mode information Pinfo is arbitrary and may include information other than the above information.

[0261] <Transformation information Tinfo>

[0262] The transformation information Tinfo includes, for example, the following information. Of course, the information included in the transformation information Tinfo is arbitrary and may include information other than the above information:

[0263] The width TBWSize and height TBHSize of the transform block to be processed (or it can be the base-2 logarithms log2TBWSize and log2TBHSize of TBWSize and TBHSize);

[0264] Transform skip flag (transform_ts_flag): A flag indicating whether to skip the (inverse) first transform and the (inverse) second transform.

[0265] Scan identifier (scanIdx);

[0266] Second transform identifier (st_idx);

[0267] Adaptive orthogonal transform identifier (mts_idx);

[0268] Quantization parameter (qp); and

[0269] Quantization matrix (scaling_matrix (e.g., JCTVC-W1005, 7.3.4 Scaling list data syntax)).

[0270] <Residual information Rinfo>

[0271] Residual information Rinfo (e.g., see 7.3.8.11 Residual coding syntax of JCTVC-W1005) includes, for example, the following syntax:

[0272] cbf (coded_block_flag): Residual data presence / absence flag;

[0273] last_sig_coeff_x_pos: The X coordinate of the last non-zero transform coefficient;

[0274] last_sig_coeff_y_pos: The Y coordinate of the last non-zero transform coefficient;

[0275] coded_sub_block_flag: Sub-block non-zero transform coefficient presence / absence flag;

[0276] sig_coeff_flag: Non-zero transform coefficient presence / absence flag;

[0277] gr1_flag: A flag indicating whether the level of the non-zero transform coefficient is greater than 1 (also known as the GR1 flag);

[0278] gr2_flag: A flag indicating whether the level of the non-zero transform coefficient is greater than 2 (also known as the GR2 flag);

[0279] sign_flag: A code indicating the sign of non-zero transform coefficients (also known as the sign code);

[0280] coeff_abs_level_remaining: The residual level of non-zero transform coefficients (also known as the non-zero transform coefficient residual level);

[0281] and so on.

[0282] Of course, the information included in the residual information Rinfo is arbitrary and may include information other than the above information.

[0283] <Filtering information Finfo>

[0284] The filtering information Finfo includes, for example, control information regarding the following filtering processes:

[0285] Control information regarding the deblocking filter (DBF)

[0286] Control information regarding the sample adaptive offset (SAO)

[0287] Control information regarding the adaptive loop filter (ALF)

[0288] Control information regarding other linear / nonlinear filters

[0289] More specifically, the filtering information Finfo includes, for example, the pictures to which each filter is applied, information for specifying regions in the pictures, filter on / off control information for each CU, filter on / off control information for slice and tile boundaries, etc. Of course, the information included in the filtering information Finfo is arbitrary and may include information other than the above information.

[0290] Returning to the description of the decoding unit 212. The decoding unit 212 refers to the residual information Rinfo and derives the quantized transform coefficient level "level" at each coefficient position in each transform block. The decoding unit 212 provides the quantized transform coefficient level "level" to the inverse quantization unit 213.

[0291] In addition, the decoding unit 212 provides the parsed header information Hinfo, prediction mode information Pinfo, quantized transform coefficient level "level", transform information Tinfo, and filtering information Finfo to each block. As described specifically below.

[0292] The header information Hinfo is provided to the inverse quantization unit 213, inverse orthogonal transform unit 214, prediction unit 219, and loop filtering unit 216.

[0293] The prediction mode information Pinfo is provided to the inverse quantization unit 213 and the prediction unit 219.

[0294] The transform information Tinfo is provided to the inverse quantization unit 213 and the inverse orthogonal transform unit 214.

[0295] The filtering information Finfo is provided to the loop filtering unit 216.

[0296] Of course, the above examples are examples, and this embodiment is not limited to this example. For example, each coding parameter can be provided to any processing unit. In addition, other information can be provided to any processing unit.

[0297] <Inverse Quantization Unit>

[0298] The inverse quantization unit 213 has at least a configuration required to perform processing regarding inverse quantization. For example, the inverse quantization unit 213 uses the transform information Tinfo provided from the decoding unit 212 and the quantization transform coefficient level "level" as inputs, and scales (inverse quantizes) the value of the quantization transform coefficient level "level" based on the transform information Tinfo to derive the transform coefficient Coeff_IQ after inverse quantization.

[0299] Note that this inverse quantization is performed as an inverse process of the quantization performed by the quantization unit 114 of the image coding device 100. In addition, the inverse quantization is a process similar to the inverse quantization performed by the inverse quantization unit 117 of the image coding device 100. That is, the inverse quantization unit 117 of the image coding device 100 performs a process (inverse quantization) similar to that of the inverse quantization unit 213.

[0300] The inverse quantization unit 213 provides the derived transform coefficient Coeff_IQ to the inverse orthogonal transform unit 214.

[0301] <Inverse Orthogonal Transform Unit>

[0302] The inverse orthogonal transform unit 214 performs processing regarding inverse orthogonal transform. For example, the inverse orthogonal transform unit 214 uses the transform coefficient Coeff_IQ provided from the inverse quantization unit 213 and the transform information Tinfo provided from the decoding unit 212 as inputs, and performs an inverse orthogonal transform process on the transform coefficient Coeff_IQ based on the transform information Tinfo to derive the prediction residual D'.

[0303] Note that this inverse orthogonal transform is performed as an inverse process of the orthogonal transform performed by the orthogonal transform unit 113 of the image encoding device 100. In addition, the inverse orthogonal transform is a process similar to the inverse orthogonal transform performed by the inverse orthogonal transform unit 118 of the image encoding device 100. That is, the inverse orthogonal transform unit 118 of the image encoding device 100 performs a process (inverse orthogonal transform) similar to that of the inverse orthogonal transform unit 214.

[0304] The inverse orthogonal transform unit 214 provides the derived prediction residual D' to the calculation unit 215.

[0305] <Calculation Unit>

[0306] The calculation unit 215 performs a process related to the addition of information about the image. For example, the calculation unit 215 uses the prediction residual D' provided from the inverse orthogonal transform unit 214 and the prediction image P provided from the prediction unit 219 as inputs. The calculation unit 215 adds the prediction residual D' and the prediction image P (prediction signal) corresponding to the prediction residual D' to derive a local decoded image Rlocal, as shown in the following expression.

[0307] Rlocal = D' + P

[0308] The calculation unit 215 provides the derived local decoded image Rlocal to the loop filter unit 216 and the frame memory 218.

[0309] <Loop Filter Unit>

[0310] The loop filter unit 216 performs a process related to loop filtering. For example, the loop filter unit 216 uses the local decoded image Rlocal provided from the calculation unit 215 and the filtering information Finfo provided from the decoding unit 212 as inputs. Note that the information input to the loop filter unit 216 is arbitrary and may include information other than the above-mentioned information.

[0311] The loop filter unit 216 appropriately performs a filtering process on the local decoded image Rlocal based on the filtering information Finfo.

[0312] For example, the loop filter unit 216 applies these four loop filters in the order of a bilateral filter, a deblocking filter (DBF), an adaptive offset filter (sample adaptive offset (SAO)), and an adaptive loop filter (adaptive loop filter (ALF)). Note that which filter to apply and in what order to apply the filters is arbitrary and can be appropriately selected.

[0313] The loop filter unit 216 performs a filtering process corresponding to the filtering process performed on the encoding side (e.g., by the loop filter unit 120 of the image encoding apparatus 100). Of course, the filtering process performed by the loop filter unit 216 is arbitrary and is not limited to the above example. For example, the loop filter unit 216 may apply a Wiener filter or the like.

[0314] The loop filter unit 216 supplies the filtered local decoded image Rlocal to the rearrangement buffer 217 and the frame memory 218.

[0315] <Rearrangement buffer>

[0316] The rearrangement buffer 217 uses the local decoded image Rlocal supplied from the loop filter unit 216 as an input and holds (stores) the local decoded image Rlocal. The rearrangement buffer 217 reconstructs the decoded image R for each picture unit using the local decoded image Rlocal and holds (stores) the decoded image R (in the buffer). The rearrangement buffer 217 rearranges the obtained decoded image R from the decoding order to the reproduction order. The rearrangement buffer 217 outputs the rearranged decoded image R group as moving image data to the outside of the image decoding apparatus 200.

[0317] <Frame memory>

[0318] The frame memory 218 performs a process of storing data related to an image. For example, the frame memory 218 uses the local decoded image Rlocal supplied from the calculation unit 215 as an input, reconstructs the decoded image R for each picture unit, and stores the decoded image R in a buffer in the frame memory 218.

[0319] In addition, the frame memory 218 uses the loop-filtered local decoded image Rlocal supplied from the loop filter unit 216 as an input, reconstructs the decoded image R for each picture unit, and stores the decoded image R in a buffer in the frame memory 218. The frame memory 218 appropriately supplies the stored decoded image R (or a part thereof) to the prediction unit 219 as a reference image.

[0320] Note that the frame memory 218 may store header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filtering information Finfo, etc. related to the generation of the decoded image.

[0321] <Prediction unit>

[0322] The prediction unit 219 performs processing related to the generation of a prediction image. For example, the prediction unit 219 uses the prediction mode information Pinfo provided from the decoding unit 212 as an input, and performs prediction by the prediction method specified by the prediction mode information Pinfo to derive a prediction image P. When deriving, the prediction unit 219 uses the decoded image R (or a part thereof) before or after filtering stored in the frame memory 218 as a reference image, and the decoded image R is specified by the prediction mode information Pinfo. The prediction unit 219 provides the derived prediction image P to the calculation unit 215.

[0323] Note that these processing units (accumulation buffer 211 to prediction unit 219) have an arbitrary configuration. For example, each processing unit may be configured by a logic circuit that implements the above processing. In addition, each processing unit may include, for example, a CPU, a ROM, a RAM, etc., and implements the above processing by executing a program using the above resources. Of course, each processing unit may have both of these configurations, and implements a part of the above processing by a logic circuit, and implements another part of the processing by executing a program. The configurations of the processing units may be independent of each other. For example, some of the processing units may implement a part of the above processing by a logic circuit, some of the processing units may implement the above processing by executing a program, and some of the processing units may implement the above processing by both a logic circuit and the execution of a program.

[0324] <Decoding Unit>

[0325] Figure 11 is a block diagram showing Figure 10 a main configuration example of the decoding unit 212 in Figure 11 As shown, the decoding unit 212 includes a selection unit 231, a context setting unit 232, a context decoding unit 233, a bypass decoding unit 234, and an inverse binarization unit 235.

[0326] Note that although the decoding of the encoded data of the adaptive orthogonal transform identifier is described here, the decoding unit 212 also decodes the encoded data of other encoding parameters, residual information Rinfo, etc. as described above. The decoding unit 212 decodes the encoded data of the adaptive orthogonal transform identifier by applying Method 1 described in <2-2. Decoding of Adaptive Orthogonal Transform Identifier>.

[0327] The selection unit 231 performs processing for selecting the supply destination of the encoded data of each bin in the bin sequence of the binarized adaptive orthogonal transform identifier. For example, the selection unit 231 acquires the encoded data of the bin sequence of the adaptive orthogonal transform identifier provided from the accumulation buffer 211.

[0328] In addition, the selection unit 231 selects whether to set the supply destination to the context setting unit 232 or the bypass decoding unit 234 for the encoded data of each bin in the bin sequence of the acquired adaptive orthogonal transform identifier. The selection unit 231 makes the selection according to Method 1 described above in <2-2. Decoding of Adaptive Orthogonal Transform Identifier>. For example, the selection unit 231 may make the selection according to Method 1-1 (i.e., Figure 3 the table shown in A of Figure 3 ). In addition, the selection unit 231 may make the selection according to Method 1-2 (i.e., Figure 4 the table shown in A of Figure 4 ). In addition, the selection unit 231 may make the selection according to Method 1-3 (i.e.,

[0329] the table shown in A of

[0330] ). In addition, the selection unit 231 may make the selection according to Method 1-4 (i.e., Figure 3 the table shown in B of Figure 3 ). In addition, the selection unit 231 may make the selection according to Method 1-3 (i.e., Figure 4 the table shown in A of Figure 4 ). In addition, the selection unit 231 may make the selection according to Method 1-4 (i.e.,

[0331] the table shown in B of

[0332] In the case of allocating a context variable (index ctxInc) and performing context decoding, the selection unit 231 provides the encoded data of the bin to the context setting unit 232. In the case of bypass decoding, the selection unit 231 provides the bin to the bypass decoding unit 234.

[0330] The context setting unit 232 performs processing regarding context setting. For example, the context setting unit 232 acquires the encoded data of the bin provided from the selection unit 231. The context setting unit 232 allocates a context variable (index ctxInc) to the bin. The context setting unit 232 performs the allocation according to Method 1 described above in <2-2. Decoding of Adaptive Orthogonal Transform Identifier>. For example, the context setting unit 232 may perform the allocation according to Method 1-1 (i.e., Figure 3 the table shown in A of Figure 3 ). In addition, the context setting unit 232 may perform the allocation according to Method 1-2 (i.e., Figure 4 the table shown in A of Figure 4 ). In addition, the context setting unit 232 may perform the allocation according to Method 1-3 (i.e.,

[0331] the table shown in A of

[0332] In addition, the context setting unit 232 may acquire the decoding result from the context decoding unit 233. The context setting unit 232 may appropriately update the context variable (index ctxInc) using the decoding result. The context setting unit 232 provides the context variable (index ctxInc) derived in this way to the context decoding unit 233.

[0332] The context decoding unit 233 performs processing regarding arithmetic decoding. For example, the context decoding unit 233 acquires a context variable (index ctxInc) provided from the context setting unit 232. Further, the context decoding unit 233 performs arithmetic decoding using the context variable (index ctxInc). That is, context decoding is performed. Further, the context decoding unit 233 supplies the bin to be processed in the bin sequence of the adaptive orthogonal transform identifier, which is the decoding result, to the inverse binarization unit 235.

[0333] The bypass decoding unit 234 performs processing regarding bypass decoding. For example, the bypass decoding unit 234 acquires the encoded data of the bin provided from the selection unit 231. The bypass decoding unit 234 performs bypass decoding (arithmetic decoding) on the encoded data of the bin. The bypass decoding unit 234 supplies the bin to be processed in the bin sequence of the adaptive orthogonal transform identifier, which is the decoding result, to the inverse binarization unit 235.

[0334] The inverse binarization unit 235 performs processing regarding inverse binarization (also referred to as multi-binarization) of the bin sequence of the adaptive orthogonal transform identifier. For example, the inverse binarization unit 235 acquires the bin sequence of the adaptive orthogonal transform identifier provided from the context decoding unit 233 or the bypass decoding unit 234. The inverse binarization unit 235 performs inverse binarization on the acquired bin sequence to derive the adaptive orthogonal transform identifier mts_idx. This inverse binarization is the inverse process of the binarization performed by the binarization unit 131. That is, the inverse binarization unit 235 performs inverse binarization using a truncated unary code (or a truncated Rice code). The inverse binarization unit 235 supplies the derived adaptive orthogonal transform identifier mts_idx to the inverse orthogonal transform unit 214 as Tinfo. The inverse orthogonal transform unit 214 appropriately performs adaptive orthogonal transform based on the adaptive orthogonal transform identifier mts_idx.

[0335] Since each of the processing units (the selection unit 231 to the inverse binarization unit 235) performs the processing as described above, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 1 (for example, any one of Method 1-1 to Method 1-4). Therefore, an increase in the memory usage amount can be suppressed. Further, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0336] <Flow of Image Decoding Processing>

[0337] Next, the flow of each process performed by the image decoding apparatus 200 having the above configuration will be described. First, an example of the flow of the image decoding process will be described with reference to the flowchart in Figure 12 The flow of the image decoding process will be described with reference to the flowchart in

[0338] When the image decoding process starts, in step S201, the accumulation buffer 211 acquires and holds (accumulates) the encoded data (bitstream) provided from outside the image decoding device 200.

[0339] In step S202, the decoding unit 212 decodes the encoded data (bitstream) to obtain the quantized transform coefficient level "level". In addition, the decoding unit 212 parses (analyzes and obtains) various encoding parameters from the encoded data (bitstream) through this decoding.

[0340] In step S203, the inverse quantization unit 213 performs inverse quantization, which is the inverse process of the quantization performed on the encoding side, on the quantized transform coefficient level "level" obtained through the process in step S202 to obtain the transform coefficient Coeff_IQ.

[0341] In step S204, the inverse orthogonal transformation unit 214 performs an inverse orthogonal transformation process, which is the inverse process of the orthogonal transformation process performed on the encoding side, on the transform coefficient Coeff_IQ obtained in step S203 to obtain the prediction residual D'.

[0342] In step S205, the prediction unit 219 performs a prediction process based on the information parsed in step S202 by the prediction method specified on the encoding side, and generates a prediction image P by referring to, for example, the reference image stored in the frame memory 218.

[0343] In step S206, the calculation unit 215 adds the prediction residual D' obtained in step S204 to the prediction image P obtained in step S205 to derive the local decoded image Rlocal.

[0344] In step S207, the loop filter unit 216 performs a loop filter process on the local decoded image Rlocal obtained through the process in step S206.

[0345] In step S208, the rearrangement buffer 217 uses the filtered local decoded image Rlocal obtained through the process in step S207 to derive the decoded image R, and rearranges the decoded image R group from the decoding order to the reproduction order. The decoded image R group rearranged in the reproduction order is output as a moving image to the outside of the image decoding device 200.

[0346] In addition, in step S209, the frame memory 218 stores at least one of the local decoded image Rlocal obtained through the process in step S206 and the filtered local decoded image Rlocal obtained through the process in step S207.

[0347] When the process in step S209 ends, the image decoding process ends.

[0348] <Flow of the decoding process>

[0349] In Figure 12 in the decoding process of step S202, the decoding unit 212 decodes the encoded data of the adaptive orthogonal transform identifier mts_idx. At this time, the decoding unit 212 decodes the encoded data of the adaptive orthogonal transform identifier by applying Method 1 described in <2-2. Decoding of the adaptive orthogonal transform identifier>. An example of the flow of decoding the encoded data of the adaptive orthogonal transform identifier mts_idx will be described with reference to the flowchart in Figure 13 below.

[0350] When the decoding process starts, in step S231, the selection unit 231 of the decoding unit 212 sets the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier mts_idx as the bin to be processed. In this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (i.e., selects context decoding). For example, the selection unit 231 selects context decoding as the decoding method for this bin according to any one of the tables shown in Figure 3 and Figure 4 below (i.e., by applying any one of Methods 1-1 to 1-4).

[0351] In step S232, the context setting unit 232 assigns a predetermined context variable ctx (index ctxInc) to the bin. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0352] In step S233, the selection unit 231 sets the unprocessed bins in the second and subsequent bins in the bin sequence as the bins to be processed. In step S234, the selection unit 231 determines whether to perform bypass decoding on the bin to be processed. For example, the selection unit 231 determines whether to perform bypass decoding on the encoded data of the bin to be processed according to any one of the tables shown in Figure 3 and Figure 4 below (i.e., by applying any one of Methods 1-1 to 1-4).

[0353] In the case where it is determined to perform bypass decoding, the process proceeds to step S235. That is, in this case, the selection unit 231 selects the bypass decoding unit 234 as the supply destination of the encoded data of the bin. In step S235, the bypass decoding unit 234 performs bypass decoding on the encoded data of the bin to be processed. When the process in step S235 is completed, the process proceeds to step S237.

[0354] In addition, in step S234, when it is determined not to perform bypass decoding (perform context decoding), the process proceeds to step S236. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the encoded data of the bin. In step S236, the context setting unit 232 assigns a predetermined context variable ctx (index ctxInc) to the bin. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed. When the process in step S236 is completed, the process proceeds to step S237.

[0355] In step S237, the inverse binarization unit 235 inverse-binarizes the bin sequence using a truncated unary code (or a truncated Rice code) to derive an adaptive orthogonal transform identifier mts_idx.

[0356] In step S238, the decoding unit 212 determines whether to terminate the decoding of the adaptive orthogonal transform identifier mts_idx. When it is determined not to terminate the decoding, the process returns to step S233, and the processes in step S233 and subsequent steps are repeated. In addition, in step S238, when it is determined to terminate the decoding, the decoding process is terminated.

[0357] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 1 (for example, any one of Methods 1-1 to 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0358] <3. Second Embodiment>

[0359] <3-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0360] In this embodiment, as already performed in the table shown in Figure 1 B below, the assignment of context variables to each bin in the bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (Method 0) is performed as follows.

[0361] That is, a context variable based on a parameter regarding the block size is assigned to the first bin in the bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0362] More specifically, the parameter regarding the block size is the difference between the logarithm of the longer side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied. That is, a context variable based on the difference between the logarithm of the longer side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image coding, and context coding (Method 2) is performed on the first bin.

[0363] For example, as shown in the table in Figure 14 A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer one of the logarithm of the transform block size in the horizontal direction (log2W) and the logarithm of the transform block size in the vertical direction (log2H) (max(log2W, log2H)) and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied (log2MinMtsSize) (max(log2W, log2H) - log2MinMtsSize), context coding can be performed on the first bin, and bypass coding (Method 2-1) can be performed on the second to fourth bins in the bin sequence.

[0364] In Figure 14 the case of the example in A, the index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context coding is performed on the first bin, and bypass coding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0365] Furthermore, for example, as shown in the table in Figure 14 B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer one of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W, log2H) - log2MinMtsSize), context coding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, context coding can be performed on the second bin, and bypass coding can be performed on the third and fourth bins in the bin sequence (Method 2-2).

[0366] InFigure 14 In the case of the example in B, an index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. And bypass encoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0367] In addition, for example, as Figure 15 shown in the table in A, a context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin in the bin sequence, and context encoding can be performed on the second bin and the third bin, and bypass encoding can be performed on the fourth bin in the bin sequence (Method 2-3).

[0368] In Figure 15 the case of the example in A, an index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And bypass encoding is performed on the fourth bin (binIdx = 3).

[0369] In addition, for example, as Figure 15As shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize), and context encoding can be performed on the first bin, and different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 2-4).

[0370] In Figure 15 In the case of the example in B, the index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin, the index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin, and the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.

[0371] Note that in Figure 14 and Figure 15 the table, non-overlapping unique values are set in the indices B1, B2, and B3.

[0372] Figure 16 Examples of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are shown in the table in. For example, in the case of Method 0, the number of contexts is 9, the number of context encoding bins is 4, and the number of bypass encoding bins is 0. In contrast, in the case of Method 2-1, the number of contexts is 4, the number of context encoding bins is 1, and the number of bypass encoding bins is 3. In addition, in the case of Method 2-2, the number of contexts is 5, the number of context encoding bins is 2, and the number of bypass encoding bins is 2. In addition, in the case of Method 2-3, the number of contexts is 6, the number of context encoding bins is 3, and the number of bypass encoding bins is 1. In addition, in the case of Method 2-4, the number of contexts is 7, the number of context encoding bins is 4, and the number of bypass encoding bins is 0.

[0373] As described above, in any of the cases of Method 2-1 to Method 2-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 2, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0374] In addition, in any of the cases of Method 2-1 to Method 2-3, the number of context encoding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 2-4, the number of context encoding bins required for encoding is equivalent to the number of context encoding bins required in the case of Method 0. That is, by applying Method 2, bypass encoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context encoding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.

[0375] As described above, by applying Method 2, an increase in the load of the encoding process can be suppressed.

[0376] <3-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0377] Similarly, in the case of decoding, context variables are assigned to each bin in the bin sequence of the binarized adaptive orthogonal transform identifier indicating the mode of inverse adaptive orthogonal transform in image decoding as has been performed in the table shown in Figure 1 B.

[0378] That is, a context variable based on a parameter regarding the block size is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.

[0379] More specifically, the parameter regarding the block size is the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied. That is, a context variable based on the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding (Method 2) is performed on the first bin.

[0380] For example, as Figure 14As shown in the table in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference (max(log2W, log2H) - log2MinMtsSize) between the logarithm (log2W) of the transform block size in the horizontal direction and the logarithm (log2H) of the transform block size in the vertical direction, where max(log2W, log2H) is the longer of the two, and log2MinMtsSize is the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform. Context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 2-1).

[0381] In Figure 14 In the case of the example in A, the index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier. Context decoding is performed on the first bin, and bypass decoding is performed on the second (binIdx = 1) to fourth (binIdx = 3) bins.

[0382] Furthermore, for example, as Figure 14 shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference (max(log2W, log2H) - log2MinMtsSize) between the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction. Context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 2-2).

[0383] In Figure 14 In the case of the example in B, the index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier. Context decoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), context decoding is performed on the second bin, and bypass decoding is performed on the third (binIdx = 2) and fourth (binIdx = 3) bins.

[0384] In addition, for example, as Figure 15 shown in the table in A of Figure 15 , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context decoding can be performed on the second and third bins, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 2-3).

[0385] In Figure 15 the case of the example in A of Figure 15 , the index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin, and bypass decoding (bypass) is performed on the fourth bin (binIdx = 3).

[0386] In addition, for example, as Figure 15 shown in the table in B of Figure 15 , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 2-4).

[0387] In Figure 15In the case of the example in B, an index ctxInc based on the difference (max(log2W, log2H) - log2MinMtsSize) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.

[0388] Note that even in the case of decoding, similar to the case of encoding, Figure 14 and Figure 15 non-overlapping unique values are also set in the indexes B1, B2, and B3 in the table.

[0389] The number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are similar to the case of encoding ( Figure 16 ).

[0390] As described above, in any of the cases of Method 2-1 to Method 2-4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 2, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0391] In addition, in any of the cases of Method 2-1 to Method 2-3, compared with the case of Method 0, the number of context coding bins required for decoding can be reduced. Note that in the case of Method 2-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 2, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context coding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in coding efficiency.

[0392] As described above, by applying Method 2, an increase in the load of the decoding process can be suppressed.

[0393] <3-3. Encoding side>

[0394] <Configuration>

[0395] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side of the first embodiment. That is, the image encoding device 100 in this case has the same as that referred to Figure 6The described configuration is similar to the configuration. Additionally, in this case, the coding unit 115 has a configuration similar to the configuration described with reference to Figure 7 The described configuration.

[0396] <Flow of the coding process>

[0397] Furthermore, in this case, the image coding device 100 performs processing that is substantially similar to the case of the first embodiment. That is, the image coding process performed by the image coding device 100 in this case is executed through a process similar to the process described in the flowchart with reference to Figure 8 The flowchart in.

[0398] Reference will be made to Figure 17 The flowchart in to describe an example of the process of the coding process for encoding the adaptive orthogonal transform identifier performed by the coding unit 115 in this case.

[0399] In this coding process, the processing in steps S301 and S302 is performed similarly to the processing in steps S131 and S132 in Figure 9 The flowchart in. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (i.e., selects context coding). For example, the selection unit 132 selects context coding as the coding method for this bin according to Figure 14 And Figure 15 Any one of the tables shown in the table (i.e., by applying any one of methods 2-1 to 2-4).

[0400] In step S303, the context setting unit 133 assigns the context variable ctx (index ctxInc) to this bin based on the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize). Then, the context coding unit 134 performs arithmetic coding using the context variable. That is, context coding is performed.

[0401] The processing in steps S304 to S308 is also performed similarly to the processing in steps S134 to S138 in Figure 9 The flowchart in. In step S308, when it is determined to terminate the coding, the coding process is terminated.

[0402] By performing each of the processes described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 2 (e.g., any one of Methods 2-1 to 2-4). Accordingly, an increase in the memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0403] <3-4. Decoding side>

[0404] <Configuration>

[0405] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to the configuration of the decoding side of the first embodiment. That is, the image decoding apparatus 200 in this case has a configuration similar to the configuration described with reference to Figure 10 In addition, the decoding unit 212 in this case has a configuration similar to the configuration described with reference to Figure 11

[0406] <Flow of decoding process>

[0407] In addition, the image decoding apparatus 200 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image decoding process performed by the image decoding apparatus 200 in this case is performed through a process similar to the process described in the flowchart with reference to Figure 12

[0408] A flowchart in Figure 18 will be referred to to describe an example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.

[0409] In this decoding process, the process of step S321 is performed in a manner similar to the process of step S231 in Figure 13 . That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (i.e., selects context decoding). For example, the selection unit 231 selects context decoding as the decoding method for this bin according to any one of the tables shown in Figure 14 and Figure 15 (i.e., by applying any one of Methods 2-1 to 2-4).

[0410] ​​In step S322, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin based on the difference between the longer one of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (max(log2W, log2H) - log2MinMtsSize). Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0411] The processing in steps S323 to S328 is also performed Figure 13 similarly to the processing in steps S233 to S238 in

[0412] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 2 (for example, any one of Methods 2-1 to 2-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0413] <4. Third Embodiment>

[0414] <4-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0415] In the present embodiment, the context variable is assigned to each bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding (Method 0) as already performed in the table shown in Figure 1 B of

[0416] That is, the context variable based on the parameter regarding the block size is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0417] More specifically, the parameter regarding the block size is the minimum value between the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold. That is, the context variable based on the smaller value between the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and the predetermined threshold is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin (Method 3).

[0418] Note that the threshold TH is an arbitrary value. By setting the threshold to a value less than the maximum value of the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform, the number of contexts can be reduced compared to the case of Method 2. Therefore, the memory usage can be reduced.

[0419] For example, as shown in the table in A of Figure 19 it is possible to assign the context variable ctx (index ctxInc) to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference (max(log2W, log2H) - log2MinMtsSize) between the longer of the logarithm of the transform block size in the horizontal direction (log2W) and the logarithm of the transform block size in the vertical direction (log2H) (max(log2W, log2H)) and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform (log2MinMtsSize) and a predetermined threshold (TH), and it is possible to perform context encoding on the first bin, and it is possible to perform bypass encoding on the second to fourth bins in the bin sequence (Method 3-1).

[0420] In Figure 19 the case of the example in A of

[0421] In addition, for example, as shown in Figure 19As shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the larger value of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context coding can be performed on the first bin. A predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context coding can be performed on the second bin, and bypass coding can be performed on the third and fourth bins in the bin sequence (Method 3-2).

[0422] In Figure 19 In the case of the example in B, the index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context coding is performed on the second bin, and bypass coding (bypass) is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0423] In addition, for example, as Figure 20 shown in the table in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the larger value of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context coding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context coding can be performed on the second and third bins, and bypass coding can be performed on the fourth bin in the bin sequence (Method 3-3).

[0424] In Figure 20In the case of the example in A, the index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin, and bypass encoding is performed on the fourth bin (binIdx = 3).

[0425] In addition, for example, as Figure 20 shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 3 - 4).

[0426] In Figure 20 the case of the example in B, the index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. The index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.

[0427] Note that in Figure 19 and Figure 20 the table, non - overlapping unique values are set for the indices B1, B2, and B3.

[0428] Figure 21 The table in shows examples of the number of contexts, the number of context - encoded bins, and the number of bypass - encoded bins for each of these methods.Figure 21 The example shown in A of shows an example when the threshold TH = 2. In this example, for instance, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 3-1, the number of contexts is 3, the number of context coding bins is 1, and the number of bypass coding bins is 3. In addition, in the case of method 3-2, the number of contexts is 4, the number of context coding bins is 2, and the number of bypass coding bins is 2. In addition, in the case of method 3-3, the number of contexts is 5, the number of context coding bins is 3, and the number of bypass coding bins is 1. In addition, in the case of method 3-4, the number of contexts is 6, the number of context coding bins is 4, and the number of bypass coding bins is 0.

[0429] In addition, Figure 21 The example shown in B of shows an example when the threshold TH = 1. In this example, in the case of method 3-1, the number of contexts is 2. In addition, in the case of method 3-2, the number of contexts is 3. In addition, in the case of method 3-3, the number of contexts is 4. In addition, in the case of method 3-4, the number of contexts is 5.

[0430] As described above, in any case of methods 3-1 to 3-4, compared with the case of method 0, the number of contexts required for coding can be reduced. That is, by applying method 3, the number of contexts allocated to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0431] In addition, in any one of methods 3-1 to 3-3, compared with the case of method 0, the number of context coding bins required for coding can be reduced. Note that in the case of method 3-4, the number of context coding bins required for coding is equivalent to the number of context coding bins required in the case of method 0. That is, by applying method 3, bypass coding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context coding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in coding efficiency.

[0432] As described above, by applying method 3, an increase in the load of coding processing can be suppressed.

[0433] <4-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0434] Similarly, in the case of decoding, it is performed as follows as in Figure 1Assign context variables to each bin in the bin sequence of the binarized adaptive orthogonal transform identifier indicating the mode of inverse adaptive orthogonal transform in image decoding as has been performed in the table shown in B of

[0435] That is, assign the context variable based on the parameter regarding the block size to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and perform context encoding on the first bin in the bin sequence.

[0436] More specifically, the parameter regarding the block size is the minimum value among the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied and a predetermined threshold. That is, assign the context variable based on the smaller value among the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied and the predetermined threshold to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and perform context decoding on the first bin (Method 3).

[0437] Note that, as in the case of encoding, the threshold TH is an arbitrary value. By setting the threshold to a value less than the maximum value of the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of Method 2. Therefore, the memory usage can be reduced.

[0438] For example, as Figure 19 shown in the table in A of , assign the context variable ctx (index ctxInc) to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value among the difference between the logarithm of the larger value between the logarithm of the transform block size in the horizontal direction (log2W) and the logarithm of the transform block size in the vertical direction (log2H) (max(log2W, log2H)) and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied (log2MinMtsSize) (max(log2W, log2H) - log2MinMtsSize) and a predetermined threshold (TH), and perform context decoding on the first bin, and perform bypass decoding on the second to fourth bins in the bin sequence (Method 3-1).

[0439] In Figure 19In the case of the example in A, an index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context decoding is performed on the first bin, and bypass decoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0440] In addition, for example, as Figure 19 shown in the table in B, a context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin, context decoding can be performed on the second bin, and bypass decoding can be performed on the third bin and the fourth bin in the bin sequence (Method 3-2).

[0441] In Figure 19 the case of the example in B, an index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, context decoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3) (Methods 2 and 3).

[0442] In addition, for example, as Figure 20As shown in the table in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context decoding can be performed on the second and third bins, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 3-3).

[0443] In Figure 20 In the case of the example in A, the index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin, and bypass decoding (bypass) is performed on the fourth bin (binIdx = 3).

[0444] In addition, for example, as Figure 20 As shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 3-4).

[0445] In Figure 20In the case of the example in B, the index ctxInc based on the minimum value (min(max(log2W, log2H) - log2MinMtsSize, TH)) is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier (binIdx = 0), and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.

[0446] Note that even in the case of decoding, similar to the case of encoding, Figure 19 and Figure 20 non - overlapping unique values are also set in the indices B1, B2, and B3 in the table.

[0447] The number of contexts, the number of context - encoding bins, and the number of bypass - encoding bins for each of these methods are similar to these numbers used for encoding ( Figure 21 A in Figure 21 and B in

[0448] As described above, in any of the cases of Method 3 - 1 to Method 3 - 4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 3, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0449] In addition, in any of the cases of Method 3 - 1 to Method 3 - 3, compared with the case of Method 0, the number of context - encoding bins required for decoding can be reduced. Note that in the case of Method 3 - 4, the number of context - encoding bins required for decoding is comparable to the number of context - encoding bins required for decoding in the case of Method 0. That is, by applying Method 3, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context - encoding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.

[0450] As described above, by applying Method 3, an increase in the load of the decoding process can be suppressed.

[0451] <4 - 3. Encoding side>

[0452] <Configuration>

[0453] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side in the first embodiment. That is, the image encoding apparatus 100 in this case has a configuration similar to the configuration described with reference to Figure 6 In addition, the encoding unit 115 in this case has a configuration similar to the configuration described with reference to Figure 7

[0454] <Flow of encoding process>

[0455] In addition, the image encoding apparatus 100 in this case performs processing substantially similar to that in the first embodiment. That is, the image encoding process performed by the image encoding apparatus 100 in this case is executed by a process similar to the process described in the flowchart with reference to Figure 8

[0456] Reference will be made to Figure 22 for an example of the flow of the encoding process for encoding the adaptive orthogonal transform identifier performed by the encoding unit 115 in this case described in the flowchart.

[0457] In this encoding process, the processing in steps S351 and S352 is performed similarly to the processing in steps S131 and S132 in Figure 9 . That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bins (i.e., context encoding is selected). For example, the selection unit 132 selects context encoding as the encoding method for the bins according to any one of the tables shown in Figure 19 and Figure 20 (i.e., by applying any one of methods 3-1 to 3-4).

[0458] In step S353, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bins based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H)-log2MinMtsSize, TH)). Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.

[0459] The processing in steps S354 to S358 is also performed similarly to the processing in steps S134 to S138 in Figure 9 . In step S358, when it is determined to terminate the encoding, the encoding process is terminated.

[0460] ​​By performing each of the processes described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 3 (e.g., any one of Method 3-1 to Method 3-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0461] <4-4. Decoding side>

[0462] <Configuration>

[0463] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to the configuration of the decoding side of the first embodiment. That is, the image decoding apparatus 200 in this case has a configuration similar to the configuration described with reference to Figure 10 In addition, the decoding unit 212 in this case has a configuration similar to the configuration described with reference to Figure 11 described.

[0464] <Flow of decoding process>

[0465] In addition, the image decoding apparatus 200 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image decoding process performed by the image decoding apparatus 200 in this case is performed through a process similar to the process described in the flowchart with reference to Figure 12 described.

[0466] Reference will be made to Figure 23 described in the flowchart an example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.

[0467] In this decoding process, the process of step S371 is performed similarly to the process of step S231 in Figure 13 That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bins (i.e., selects context decoding.). For example, the selection unit 231 selects context decoding as the decoding method for the bins according to any one of the tables shown in Figure 19 and Figure 20 (i.e., by applying any one of Method 3-1 to Method 3-4).

[0468] In step S372, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin based on the smaller value between the difference between the longer one of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)). Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0469] The processes in steps S373 to S378 are also Figure 13 performed similarly to the processes in steps S233 to S238 in

[0470] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 3 (for example, any one of Methods 3-1 to 3-4). Therefore, an increase in the memory usage amount can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0471] <5. Fourth Embodiment>

[0472] <5-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0473] In this embodiment, the context variable is assigned to each bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding (Method 0) as already performed in the table shown in B of Figure 1 That is, the context variable based on the parameter regarding the block size is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0474]

[0475] ​More specifically, the parameter regarding the block size is the result of right-shifting the minimum value between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied by a predetermined threshold. That is, the result of right-shifting the minimum value between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied by a predetermined threshold is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image coding, and context coding (Method 4) is performed on the first bin.

[0476] Note that the threshold TH is an arbitrary value. By setting the threshold to a value less than the maximum value of the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of Method 2. In addition, the value of the scaling parameter shift, which is the amount of right-shifting, is arbitrary. In the case of Method 4, since the minimum value is further right-shifted, the number of contexts can be reduced compared to the case of Method 3. Therefore, the memory usage can be reduced.

[0477] For example, as shown in the table in A of Figure 24 the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the result of right-shifting the smaller value between the difference (max(log2W, log2H) - log2MinMtsSize) between the logarithm of the longer of the logarithm of the transform block size in the horizontal direction (log2W) and the logarithm of the transform block size in the vertical direction (log2H) (max(log2W, log2H)) and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied (log2MinMtsSize) and a predetermined threshold (TH) (min(max(log2W, log2H) - log2MinMtsSize, TH)), context coding can be performed on the first bin, and bypass coding can be performed on the second to fourth bins in the bin sequence (Method 4-1).

[0478] In Figure 24In the case of Example A, the index ctxInc based on the result of a right shift ((min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context encoding can be performed on the first bin, and bypass encoding can be performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0479] In addition, for example, as Figure 24 shown in the table in B of, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), context encoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, context encoding can be performed on the second bin, and bypass encoding can be performed on the third bin and the fourth bin in the bin sequence (Method 4-2).

[0480] In Figure 24 the case of the example in B of, the index ctxInc based on the result of a right shift ((min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift)) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), context encoding is performed on the second bin, and bypass encoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3) (Methods 2 and 3).

[0481] In addition, for example, as Figure 25As shown in the table in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context encoding can be performed on the second and third bins, and bypass encoding (Method 4-3) can be performed on the fourth bin in the bin sequence.

[0482] In Figure 25 In the case of the example in A, the index ctxInc based on the result of right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin, and bypass encoding (bypass) is performed on the fourth bin (binIdx = 3).

[0483] In addition, for example, as Figure 25 As shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 4-4).

[0484] In Figure 25In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.

[0485] Note that in Figure 24 and Figure 25 in the table, non-overlapping unique values are set in the indices B1, B2, and B3.

[0486] Figure 26 Examples of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are shown in the table of Figure 26 The table shown in is an example in the case where the threshold TH = 2 and the scaling parameter (the shift amount of the right shift) shift = 1. In this example, for instance, in the case of method 0, the number of contexts is 9, the number of context encoding bins is 4, and the number of bypass encoding bins is 0. In contrast, in the case of method 4-1, the number of contexts is 2, the number of context encoding bins is 1, and the number of bypass encoding bins is 3. Further, in the case of method 4-2, the number of contexts is 3, the number of context encoding bins is 2, and the number of bypass encoding bins is 2. Additionally, in the case of method 4-3, the number of contexts is 4, the number of context encoding bins is 3, and the number of bypass encoding bins is 1. Moreover, in the case of method 4-4, the number of contexts is 5, the number of context encoding bins is 4, and the number of bypass encoding bins is 0.

[0487] As described above, in any of the cases of method 4-1 to method 4-4, compared with the case of method 0, the number of contexts required for encoding can be reduced. That is, by applying method 4, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0488] In addition, in any of cases of Method 4-1 to Method 4-3, the number of context coding bins required for coding can be reduced compared to the case of Method 0. Note that in the case of Method 4-4, the number of context coding bins required for coding is equivalent to the number of context coding bins required for coding in the case of Method 0. That is, by applying Method 4, bypass coding can be applied to bins corresponding to transform types with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in coding efficiency and suppress an increase in the processing amount (throughput).

[0489] As described above, by applying Method 4, an increase in the load of the coding process can be suppressed.

[0490] <5-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0491] Similarly, in the case of decoding, as has been performed in the table shown in B of Figure 1 context variables are assigned to each bin in the bin sequence of the binarized adaptive orthogonal transform identifier indicating the mode of inverse adaptive orthogonal transform in image decoding.

[0492] That is, a context variable based on a parameter regarding the block size is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context coding is performed on the first bin in the bin sequence.

[0493] More specifically, the parameter regarding the block size is the result of right-shifting the minimum value between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which adaptive orthogonality can be applied and a predetermined threshold. That is, a context variable based on the result of right-shifting the minimum value between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which adaptive orthogonal transform can be applied and a predetermined threshold is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding (Method 4) is performed on the first bin.

[0494] Note that, as in the case of coding, the threshold TH is an arbitrary value. By setting the threshold to a value less than the maximum value of the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of Method 2. In addition, the value of the scaling parameter shift as the amount of right-shifting is arbitrary. In the case of Method 4, since the minimum value is further right-shifted, the number of contexts can be reduced compared to the case of Method 3. Therefore, the memory usage can be reduced.

[0495] For example, as Figure 24As shown in the table in A, based on the smaller value between the difference (max(log2W, log2H) - log2MinMtsSize) between the logarithm (log2W) of the transform block size in the horizontal direction and the logarithm (log2H) of the transform block size in the vertical direction and the logarithm (log2MinMtsSize) of the minimum transform block size applicable to the adaptive orthogonal transform, and a predetermined threshold (TH), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier after right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift). Context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 4-1).

[0496] In Figure 24 In the case of the example in A, the index ctxInc based on the result of right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier. Context decoding is performed on the first bin, and bypass decoding (bypass) is performed on the second (binIdx = 1) to fourth (binIdx = 3) bins.

[0497] In addition, for example, as Figure 24 shown in the table in B, based on the smaller value between the difference between the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform, and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier. Context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 4-2).

[0498] In Figure 24In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. And bypass decoding (Methods 2 and 3) is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0499] In addition, for example, as Figure 25 shown in the table in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin in the bin sequence, and context decoding can be performed on the second bin and the third bin. And bypass decoding (Method 4-3) can be performed on the fourth bin in the bin sequence.

[0500] In Figure 25 the case of the example in A, the index ctxInc based on the result of the right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And bypass decoding (bypass) is performed on the fourth bin (binIdx = 3).

[0501] In addition, for example, as Figure 25As shown in the table in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the smaller value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH)), and context decoding can be performed on the first bin, and different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 4-4).

[0502] In Figure 25 In the case of the example in B, the index ctxInc based on the result of right shift (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift) is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin, the index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin, and the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.

[0503] Note that even in the case of decoding, similar to the case of encoding, non-overlapping unique values are also set in the indexes B1, B2, and B3 in the Figure 24 and Figure 25 table.

[0504] The number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are similar to those used for encoding( Figure 26 ).

[0505] As described above, in any of the cases of Method 4-1 to Method 4-4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 4, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0506] In addition, in any of the cases of Method 4-1 to Method 4-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 4-4, the number of context coding bins required for decoding is equivalent to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 4, bypass decoding can be applied to bins corresponding to transform types with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins and an increase in the processing amount (throughput) while suppressing a decrease in coding efficiency.

[0507] As described above, by applying Method 4, an increase in the load of the decoding process can be suppressed.

[0508] <5-3. Encoding side>

[0509] <Configuration>

[0510] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side of the first embodiment. That is, the image encoding apparatus 100 in this case has a configuration similar to the configuration of the image encoding apparatus 100 described with reference to Figure 6 In addition, the encoding unit 115 in this case has a configuration similar to the configuration described with reference to Figure 7 described.

[0511] <Flow of encoding process>

[0512] In addition, the image encoding apparatus 100 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image encoding process performed by the image encoding apparatus 100 in this case is performed through a process similar to the process described in the flowchart in Figure 8 described.

[0513] Reference will be made to Figure 27 in the flowchart to describe an example of the flow of the encoding process for encoding the adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.

[0514] In this encoding process, the processing in steps S401 and S402 is performed similarly to the processing in steps S131 and S132 in Figure 9 That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (i.e., context coding is selected). For example, the selection unit 132 selects context coding as the coding method for the bin according to any one of the tables shown in Figure 24 and Figure 25 (that is, by applying any one of Method 4-1 to Method 4-4).

[0515] In step S403, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bin based on the result of right-shifting the minimum value between the difference between the longer one of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction by the scaling parameter shift and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H)-log2MinMtsSize, TH)>>shift). Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.

[0516] The processes in steps S404 to S408 are also Figure 9 performed similarly to the processes in steps S134 to S138 in

[0517] By performing each of the above-described processes, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 4 (for example, any one of Methods 4-1 to 4-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0518] <5-4. Decoding side>

[0519] <Configuration>

[0520] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to the configuration of the decoding side of the first embodiment. That is, the image decoding device 200 in this case has a configuration similar to the configuration Figure 10 described with reference to Figure 11 In addition, the decoding unit 212 in this case has a configuration similar to the configuration

[0521] <Flow of decoding process>

[0522] In addition, the image decoding device 200 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image decoding process performed by the image decoding device 200 in this case is performed through a process similar to the process Figure 12 described in the flowchart with reference to

[0523] An example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case will be described with reference to the flowchart Figure 28 in

[0524] In this decoding process, similar to the processing of step S231 of Figure 13 , the processing of step S421 is performed. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (i.e., selects context decoding). For example, the selection unit 231 selects context decoding as the decoding method for the bin according to any one of the tables shown in Figure 24 and Figure 25 (i.e., by applying any one of methods 4-1 to 4-4).

[0525] In step S422, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin based on the result of right-shifting the minimum value between the difference between the longer of the logarithm of the transform block size in the horizontal direction and the logarithm of the transform block size in the vertical direction by the scaling parameter shift and the logarithm of the minimum transform block size applicable to the adaptive orthogonal transform and a predetermined threshold (min(max(log2W, log2H) - log2MinMtsSize, TH) >> shift). Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0526] The processing in steps S423 to S428 is also performed similarly to the processing in steps S233 to S238 of Figure 13 . In step S428, when it is determined to terminate decoding, the decoding process is terminated.

[0527] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 4 (for example, any one of methods 4-1 to 4-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0528] <6. Fifth Embodiment>

[0529] <6-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0530] In this embodiment, the context variable is assigned to each bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding as already performed in the table shown in B of Figure 1 (method 0) as follows.

[0531] That is, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.

[0532] More specifically, the context variable is assigned and context encoding is performed (Method 5) according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold.

[0533] For example, as Figure 29 shown in the table of A in, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than the predetermined threshold (TH) (S < TH?), context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (Method 5-1).

[0534] In Figure 29 the example of A in, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. In addition, bypass encoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0535] Furthermore, for example, as Figure 29 shown in the table of B in, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than the predetermined threshold (S < TH?), context encoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin, context encoding can be performed on the second bin, and bypass encoding can be performed on the third and fourth bins in the bin sequence (Method 5-2).

[0536] In Figure 29In the case of the example in B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. And bypass encoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0537] In addition, for example, as Figure 30 shown in Table A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin, and context encoding can be performed on the second bin and the third bin. And bypass encoding can be performed on the fourth bin in the bin sequence (Method 5-3).

[0538] In Figure 30 the case of the example in A, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And bypass encoding is performed on the fourth bin (binIdx = 3).

[0539] In addition, for example, as Figure 30In the table shown in B, according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold (S < TH?), a context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin, and different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins, and context encoding can be performed on the second to fourth bins (Method 5-4).

[0540] In Figure 30 In the case of the example in B, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.

[0541] Note that in Figure 29 and Figure 30 in the table, non-overlapping unique values are set among the indexes A0, A1, B1, B2, and B3.

[0542] Note that the parameter S can be any parameter as long as the parameter S is related to the block size.

[0543] For example, the parameter S can be the product of the width tbWidth and height tbHeight of the transform block, that is, the area of the transform block.

[0544] S = tbWidth * tbHeight

[0545] In addition, the parameter S can be the sum of the logarithm of the width tbWidth and the logarithm of the height tbHeight of the transform block, that is, the logarithm of the area of the transform block.

[0546] S = log2(tbWidth) + log2(tbHeight)

[0547] In addition, the parameter S can be the maximum value of the width tbWidth and height tbHeight of the transform block, that is, the size of the long side of the transform block.

[0548] S = max(tbWidth, tbHeight)

[0549] In addition, the parameter S can be the minimum value of the width tbWidth and the height tbHeight of the transform block, that is, the size of the shorter side of the transform block.

[0550] S = min(tbWidth, tbHeight)

[0551] In addition, the parameter S can be the maximum value of the logarithm of the width tbWidth and the logarithm of the height tbHeight of the transform block, that is, the logarithm of the size of the longer side of the transform block.

[0552] S = max(log2(tbWidth), log2(tbHeight))

[0553] In addition, the parameter S can be the minimum value of the logarithm of the width tbWidth and the logarithm of the height tbHeight of the transform block, that is, the logarithm of the size of the shorter side of the transform block.

[0554] S = min(log2(tbWidth), log2(tbHeight))

[0555] In addition, the parameter S can be the ratio cbSubDiv of the area of the coding block to the area of the CTU. An example of the ratio cbSubDiv of the area of the coding block to the area of the CTU is as Figure 31 shown in the table.

[0556] S = cbSubDiv

[0557] In addition, the parameter S can be the result of right-shifting the ratio cbSubDiv of the area of the coding block to the area of the CTU by the scaling parameter shift. Note that the scaling parameter shift can be the difference between the logarithm of the block size of the CTU and the logarithm of the maximum transform block size. In addition, the logarithm of the maximum transform block size can be 5.

[0558] S = cbSubDiv >> shift

[0559] shift = (log2CTUSize - log2MaxTsSize)

[0560] log2MaxTsSize = 5

[0561] In addition, the parameter S can be the absolute value of the difference between the logarithm of the width tbWidth and the logarithm of the height tbHeight of the transform block.

[0562] S = abs(log2(tbWidth) - log2(tbHeight))

[0563] Figure 32 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in the table of Figure 32 . In this example, for instance, in the case of Method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of Method 5-1, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 3. Further, in the case of Method 5-2, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 2. Additionally, in the case of Method 5-3, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 1. Moreover, in the case of Method 5-4, the number of contexts is 5, the number of context coding bins is 4, and the number of bypass coding bins is 0.

[0564] As described above, in any of the cases of Method 5-1 to Method 5-4, compared with the case of Method 0, the number of contexts required for coding can be reduced. That is, by applying Method 5, the number of contexts allocated to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0565] Furthermore, in any of the cases of Method 5-1 to Method 5-3, compared with the case of Method 0, the number of context coding bins required for coding can be reduced. Note that in the case of Method 5-4, the number of context coding bins required for coding is comparable to the number of context coding bins required for coding in the case of Method 0. That is, by applying Method 5, bypass coding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context coding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in coding efficiency.

[0566] As described above, by applying Method 5, an increase in the load of the coding process can be suppressed.

[0567] <6-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0568] Similarly, in the case of decoding, as has been performed in the table shown in B of

[0568] , context variables are assigned to each bin in the bin sequence of the binarized adaptive orthogonal transform identifier indicating the mode of the inverse adaptive orthogonal transform in image decoding. Figure 1

[0569] ​That is, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.

[0570] More specifically, a context variable is assigned and context decoding is performed according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold (Method 5).

[0571] For example, as Figure 29 shown in the table of A, a context variable ctx (index ctxInc) can be assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than a predetermined threshold (TH) (S < TH?), context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 5-1).

[0572] In Figure 29 the example of A, when the parameter regarding the block size is less than the predetermined threshold (S < TH), index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter of the block size is equal to or greater than the predetermined threshold (S ≥ TH), index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, bypass decoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0573] Furthermore, for example, as Figure 29 shown in the table of B, a context variable ctx (index ctxInc) can be assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold (S < TH?), context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin, context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 5-2).

[0574] In Figure 29In the example of B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0575] In addition, for example, as Figure 30 shown in the table of A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 5-3).

[0576] In Figure 30 the example of A, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter of the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx = 3).

[0577] In addition, for example, as Figure 30 shown in the table of B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin to the fourth bin, and context decoding can be performed on the second bin to the fourth bin (Method 5-4).

[0578] In Figure 25 the example in B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin of the adaptive orthogonal transform identifier (binIdx = 0), and context decoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. Further, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.

[0579] Note that even in the case of decoding, similar to the case of encoding, non-overlapping unique values are also set in the indices A0, A1, B1, B2, and B3 in the table of Figure 29 and Figure 30 .

[0580] Further, in the case of decoding, the parameter S can be any parameter as long as the parameter S is related to the block size, similar to the case of encoding. For example, the parameter S can be derived by various methods described in <6-1. Encoding of Adaptive Orthogonal Transform Identifier>.

[0581] The number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are similar to those for encoding ( Figure 32 ).

[0582] As described above, in any of the cases of Method 5-1 to Method 5-4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 5, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0583] Further, in any of the cases of Method 5-1 to Method 5-3, compared with the case of Method 0, the number of context encoding bins required for decoding can be reduced. Note that in the case of Method 5-4, the number of context encoding bins required for decoding is comparable to the number of context encoding bins required for decoding in the case of Method 0. That is, by applying Method 5, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context encoding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.

[0584] As described above, by applying Method 5, an increase in the load of the decoding process can be suppressed.

[0585] <6-3. Encoding side>

[0586] <Configuration>

[0587] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side of the first embodiment. That is, the image encoding apparatus 100 in this case has a configuration similar to the configuration described with reference to Figure 6 In addition, the encoding unit 115 in this case has a configuration similar to the configuration described with reference to Figure 7 described.

[0588] <Flow of encoding process>

[0589] In addition, the image encoding apparatus 100 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image encoding process performed by the image encoding apparatus 100 in this case is performed through a process similar to the process described in the flowchart with reference to Figure 8 described.

[0590] Reference will be made to Figure 33 in the flowchart to describe an example of the flow of the encoding process for encoding the adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.

[0591] In this encoding process, the processing in steps S451 and S452 is performed similarly to the processing in steps S131 and S132 in Figure 9 That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (i.e., context encoding is selected). For example, the selection unit 132 selects context encoding as the encoding method for the bin according to any one of the tables shown in Figure 29 and Figure 30 (i.e., by applying any one of Methods 5-1 to 5-4).

[0592] In step S453, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bin according to whether the parameter (S) regarding the block size is equal to or greater than a predetermined threshold (TH) (S < TH?).

[0593] For example, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the context setting unit 133 assigns the index ctxInc = A0 to the bin. In addition, when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the context setting unit 133 assigns the index ctxInc = A1 to the bin.

[0594] Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.

[0595] The processing in steps S454 to S458 is also similar to Figure 9 the processing in steps S134 to S138 and is performed similarly. In step S408, when it is determined to terminate the encoding, the encoding process is terminated.

[0596] By performing each of the above-described processes, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 5 (for example, any one of Methods 5-1 to 5-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0597] <6-4. Decoding side>

[0598] <Configuration>

[0599] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to the configuration of the decoding side of the first embodiment. That is, the image decoding device 200 in this case has a configuration similar to the configuration described with reference to Figure 10 In addition, the decoding unit 212 in this case has a configuration similar to the configuration described with reference to Figure 11 described.

[0600] <Flow of decoding process>

[0601] In addition, the image decoding device 200 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image decoding process performed by the image decoding device 200 in this case is performed by a process similar to the process described in the flowchart with reference to Figure 12 described.

[0602] Reference will be made to [[ID= the flowchart described in the example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.

[0603] In this decoding process, the processing of step S471 is performed similarly to the processing of step S231 of ​ . That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, context decoding is selected). For example, the selection unit 231 selects according to ​ and ​Select context decoding as the decoding method for the bin from any one of the tables shown (i.e., by applying any one of Methods 5-1 to 5-4).

[0604] In step S472, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin according to whether the parameter (S) regarding the block size is equal to or greater than a predetermined threshold (TH) (S < TH?).

[0605] For example, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the context setting unit 232 assigns the index ctxInc = A0 to the bin. Further, when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the context setting unit 232 assigns the index ctxInc = A1 to the bin.

[0606] Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0607] The processing in steps S473 to S478 is also performed similarly to the processing in steps S233 to S238 in ​ In step S478, when it is determined to terminate the decoding, the decoding process is terminated.

[0608] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 5 (for example, any one of Methods 5-1 to 5-4). Therefore, an increase in the memory usage amount can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0609] <7. Sixth Embodiment>

[0610] <7-1. Encoding of Adaptive Orthogonal Transform Identifier>

[0611] In the present embodiment, binarization (Method 0) of the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal transform in image encoding is performed as already performed in the table shown in A of ​ That is, the adaptive orthogonal transform identifier is binarized into a bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types, and is encoded (Method 6).

[0612] For example, as in

[0613] For example, as ​In the table shown in A, when the transform type of the adaptive orthogonal transform identifier is DCT2×DCT2, the adaptive orthogonal transform identifier is binarized into a bin sequence of one bin (=0), and when the transform type of the adaptive orthogonal transform identifier is different from DCT2×DCT2, the adaptive orthogonal transform identifier is binarized into a bin sequence of three bins.

[0614] In ​ the case of the example of A, the adaptive orthogonal transform identifier mts_idx = 0 is binarized into the bin sequence "0" indicating that the transform type is DCT2×DCT2, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx = 1 is binarized into the bin sequence "100" indicating that the transform type is different from DCT2×DCT2 and is DST7×DST7, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx = 2 is binarized into the bin sequence "101" indicating that the transform type is different from DCT2×DCT2 and is DCT8×DST7, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx = 3 is binarized into the bin sequence "110" indicating that the transform type is different from DCT2×DCT2 and is DST7×DCT8, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx = 4 is binarized into the bin sequence "111" indicating that the transform type is different from DCT2×DCT2 and is DCT8×DCT8, and is encoded.

[0615] By binarizing the adaptive orthogonal transform identifier in this way, the length of the bin sequence (bin length) can be at most three bins. In ​ the case of the example of A (method 0), the length of the bin sequence (bin length) is at most 4 bins, so by applying method 6, the bin length can be shortened by one bin.

[0616] Note that, as shown in the table in ​ B, the bin sequence can be divided into a prefix part of one bin indicating whether the transform type is different from DCT2×DCT2 and a suffix part of two bins indicating other transform types, and is binarized.

[0617] In addition, a context variable can be assigned to each bin in the bin sequence of the adaptive orthogonal transform identifier generated by binarization according to method 6 by any one of methods 0 to 5, and encoding can be performed.

[0618] For example, as ​In the table shown in A, a predetermined context variable ctx (index ctxInc), i.e., a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin, and bypass encoding can be performed on the second and third bins in the bin sequence (Method 1-1).

[0619] In ​ In the case of the example in A, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, and bypass encoding is performed on the second bin (binIdx = 1) and the third bin (binIdx = 2).

[0620] In addition, for example, as ​ In the table shown in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first and second bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first and second bins, and bypass encoding can be performed on the third bin in the bin sequence (Method 1-2).

[0621] In ​ In the case of the example in B, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin, and bypass encoding is performed on the third bin (binIdx = 2).

[0622] In addition, for example, as ​ In the table shown in C, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first to third bins (Method 1-3).

[0623] In ​ In the case of the example in C, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin, and the index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin.

[0624] Note that in​ In the table, non-overlapping unique values are set in indexes A0, B1, and B2.

[0625] In ​ the table, examples of the number of contexts, the number of context coding bins, the number of bypass coding bins, and the worst-case bin length for each of these methods are shown. For example, in the case of method 0, the worst-case bin length is 4. For example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of the combination of method 6 and method 1-1 (method 6-1), the worst-case bin length is 3. Accordingly, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 2. Further, in the case of the combination of method 6 and method 1-2 (method 6-2), the worst-case bin length is 3. Accordingly, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 1. Additionally, in the case of the combination of method 6 and method 1-3 (method 6-3), the worst-case bin length is 3. Accordingly, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 0.

[0626] As described above, in any of the cases of method 6-1 to method 6-3, compared with the case of method 0, the number of contexts required for coding can be reduced. That is, by applying method 6, the worst-case bin length can be reduced, and further, by applying method 1, the number of contexts allocated to the first bin (binIdx = 0) can be reduced. Accordingly, an increase in the memory usage can be suppressed.

[0627] Furthermore, in any of the cases of method 6-1 to method 6-3, compared with the case of method 0, the number of context coding bins required for coding can be reduced. That is, by applying method 6, the worst-case bin length is reduced, and further, by applying method 1, bypass coding can be applied to bins corresponding to transform types with relatively low selectivity. Accordingly, an increase in the number of context coding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in the coding efficiency.

[0628] Note that by combining method 6 with another method (for example, any one of methods 2 to 5) in place of method 1, the effects of method 6 and the effects of another method can be combined. Accordingly, in any case, an increase in the memory usage and the processing amount (throughput) can be suppressed.

[0629] As described above, by applying method 6, an increase in the load of the coding process can be suppressed.

[0630] <7-2. Decoding of Adaptive Orthogonal Transform Identifier>

[0631] In this embodiment, the inverse binarization (Method 0) of the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal inverse transform in image decoding is performed as already performed in the table shown in ​ A.

[0632] That is, the inverse binarization is performed on the bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types obtained by decoding, to derive the adaptive orthogonal transform identifier (Method 6).

[0633] For example, in the table shown in ​ A, the inverse binarization is performed on the bin sequence of one bin (=0) to derive the adaptive orthogonal transform identifier with the transform type of DCT2×DCT2, and the inverse binarization is performed on the bin sequence of three bins to derive the adaptive orthogonal transform identifier with the transform type different from DCT2×DCT2.

[0634] In the case of the example in ​ A, the inverse binarization is performed on the bin sequence "0" obtained by decoding the encoded data to derive the adaptive orthogonal transform identifier mts_idx = 0 indicating the transform type of DCT2×DCT2. Further, the inverse binarization is performed on the bin sequence "100" obtained by decoding the encoded data to derive the adaptive orthogonal transform identifier mts_idx = 1 indicating the transform type different from DCT2×DCT2 and being DST7×DST7. Further, the inverse binarization is performed on the bin sequence "101" obtained by decoding the encoded data to derive the adaptive orthogonal transform identifier mts_idx = 2 indicating the transform type different from DCT2×DCT2 and being DCT8×DST7. Further, the inverse binarization is performed on the bin sequence "110" obtained by decoding the encoded data to derive the adaptive orthogonal transform identifier mts_idx = 3 indicating the transform type different from DCT2×DCT2 and being DST7×DCT8. Further, the inverse binarization is performed on the bin sequence "111" obtained by decoding the encoded data to derive the adaptive orthogonal transform identifier mts_idx = 4 indicating the transform type different from DCT2×DCT2 and being DST8×DCT8.

[0635] By performing the inverse binarization on the adaptive orthogonal transform identifier in this way, the length of the bin sequence (bin length) can be at most three bins. In ​In the case of the example of A (Method 0), the length of the bin sequence (bin length) is at most 4 bins. Therefore, by applying Method 6, the bin length can be shortened by one bin.

[0636] Note that, as shown in the table in ​ B, this bin sequence can be divided into a prefix part of one bin indicating whether the transform type is different from DCT2×DCT2 and a suffix part of two bins indicating other transform types, and is inverse-binarized.

[0637] Furthermore, in the case of this inverse-binarization applying Method 6, a context variable can be assigned to each bin in the bin sequence by any one of Methods 0 to 5, and decoding can be performed.

[0638] For example, as shown in the table in ​ A, a predetermined context variable ctx (index ctxInc), i.e., a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first bin, and bypass decoding (Method 1-1) can be performed on the second and third bins in the bin sequence.

[0639] In ​ the case of the example in A, index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding is performed on the second bin (binIdx = 1) to the third bin (binIdx = 2).

[0640] For example, as shown in the table in ​ A, a predetermined context variable ctx (index ctxInc), i.e., a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first bin, and bypass decoding (Method 1-1) can be performed on the second and third bins in the bin sequence.

[0641] In ​ the case of the example in B, index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0642] Furthermore, for example, as shown in ​In the table shown in C, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding (Methods 1-3) can be performed on the first to third bins.

[0643] In ​ In the case of the example in C, the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin.

[0644] Note that ​ in the table, non-overlapping unique values are set in the indices A0, B1, and B2.

[0645] The number of contexts, the number of context-encoded bins, the number of bypass-encoded bins, and the worst-case bin length for each of these methods are similar to the case during encoding ( ​ ).

[0646] As described above, in any of the cases of Methods 6-1 to 6-3, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 6, the worst-case bin length can be reduced. In addition, by applying Method 1, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0647] Furthermore, in any of the cases of Methods 6-1 to 6-3, compared with the case of Method 0, the number of context-encoded bins required for decoding can be reduced. That is, by applying Method 6, the worst-case bin length is reduced. In addition, by applying Method 1, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context-encoded bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.

[0648] Note that by combining Method 6 with another method (for example, any one of Methods 2 to 5) instead of Method 1, the effects of Method 6 and the effects of another method can be combined. Therefore, in any case, an increase in memory usage and processing amount (throughput) can be suppressed.

[0649] As described above, by applying Method 6, an increase in the load of the decoding process can be suppressed.

[0650] <7-3. Encoding side>

[0651] <Configuration>

[0652] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side in the first embodiment. That is, the image encoding apparatus 100 in this case has a configuration similar to the configuration described with reference to ​ described. In addition, the encoding unit 115 in this case has a configuration similar to the configuration described with reference to ​ described.

[0653] <Flow of encoding process>

[0654] In addition, the image encoding apparatus 100 in this case performs processing substantially similar to that in the first embodiment. That is, the image encoding process performed by the image encoding apparatus 100 in this case is performed through a process similar to the process described in the flowchart with reference to ​ described.

[0655] Reference will be made to ​ the flowchart described in the example of the flow of the encoding process for encoding the adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.

[0656] When the encoding process starts, in step S501, the binarization unit 131 of the encoding unit 115 binarizes the adaptive orthogonal transform identifier mts_idx into a bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types.

[0657] The processing in steps S502 to S508 is also performed similarly to the processing in steps S132 to S138 in ​ described.

[0658] By performing each process in this way, the encoding unit 115 can binarize the adaptive orthogonal transform identifier by applying method 6, and can encode the bin sequence by applying method 1 (for example, any one of methods 1-1 to 1-3). Therefore, an increase in the memory usage amount can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0659] <7-4. Decoding side>

[0660] <Configuration>

[0661] Next, the decoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side in the first embodiment. That is, the image decoding apparatus 200 in this case has a configuration similar to the configuration described with reference to ​ In addition, the decoding unit 212 in this case has a configuration similar to the configuration described with reference to ​ The configuration described.

[0662] <Flow of decoding process>

[0663] In addition, the image decoding apparatus 200 in this case performs processing substantially similar to that in the first embodiment. That is, the image decoding process performed by the image decoding apparatus 200 in this case is performed through a process similar to the process described in the flowchart with reference to ​ The flowchart in.

[0664] Reference will be made to ​ The flowchart in describes an example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.

[0665] In this decoding process, the processing in steps S521 to S526 is also performed in a similar manner to the processing in steps S231 to S236 in ​ The processing in.

[0666] In step S527, the inverse binarization unit 235 performs inverse binarization on the bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types to derive the adaptive orthogonal transform identifier mts_idx.

[0667] In step S528, the decoding unit 212 determines whether to terminate the decoding of the adaptive orthogonal transform identifier mts_idx. In the case where it is determined not to terminate the decoding, the process returns to step S523, and the processing in step S523 and subsequent steps is repeated. In addition, in step S528, in the case where it is determined to terminate the decoding, the decoding process is terminated.

[0668] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 1 (for example, any one of methods 1-1 to 1-3) to derive the bin sequence. In addition, the decoding unit 212 can perform inverse binarization on the bin sequence by applying method 6 to derive the adaptive orthogonal transform identifier. Therefore, an increase in the memory usage amount can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0669] <8. Seventh Embodiment>

[0670] <8-1. Encoding of Transform Skip Flag and Adaptive Orthogonal Transform Identifier>

[0671] In the case of the method described in Non-Patent Document 1, the adaptive orthogonal transform identifier tu_mits_idx and the transform skip flag transform_skip_flag indicating whether to apply transform skip are luminance-limited (limited to the 4:2:0 format).

[0672] Therefore, in the case where the chrominance array type is greater than 1 (i.e., in the case of the chrominance format being 4:2:2 or 4:4:4), the transform skip flag and the adaptive orthogonal transform identifier are signaled (encoded / decoded) for each component ID (cIdx) so that transform skip and adaptive orthogonal transform can be applied to the chrominance components (Method 7).

[0673] An example of the syntax of the transform unit in this case is as ​ shown. As ​ shown, the transform mode is signaled for luminance (Y), chrominance (Cb), and chrominance (Cr) as shown in the 17th, 20th, and 23rd lines (gray lines) from the top of this syntax. ​ An example of the syntax of the transform mode in this case is shown. As ​ shown, the transform skip flag transform_skip_flag[x0][y0][cIdx] is signaled in the fifth line (gray line) from the top of this syntax. In addition, the adaptive orthogonal transform identifier tu_mts_idx[x0][y0][cIdx] is signaled in the 8th line (gray line) from the top of this syntax. That is, the transform skip flag and the adaptive orthogonal transform identifier are signaled for each component ID (cIdx).

[0674] By doing so, for chrominance of the chrominance format 4:2:2 or 4:4:4, which has more information than the chrominance format 4:2:0, the application of the adaptive orthogonal transform can be controlled. Therefore, a reduction in coding efficiency can be suppressed.

[0675] Then, in this case, the context variable ctx can be assigned to the first bin (binIdx = 0) in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0, and context encoding can be performed.

[0676] For example, as ​In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context coding can be performed on the first bin, and bypass coding can be performed on the second to fourth bins in the bin sequence (Method 7-1).

[0677] In ​ the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, bypass coding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0678] In addition, for example, as ​ shown in the table of B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context coding can be performed on the first bin, and a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context coding can be performed on the second bin, and bypass coding can be performed on the third and fourth bins in the bin sequence (Method 7-2).

[0679] In ​ the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context coding is performed. In addition, bypass coding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0680] In addition, for example, as ​In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context encoding can be performed on the second and third bins. Bypass encoding (Method 7-3) can be performed on the fourth bin in the bin sequence.

[0681] In ​ In the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. In addition, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context encoding is performed. In addition, in the third bin (binIdx = 2), the index ctxInc = B2 is assigned and context encoding is performed. In addition, bypass encoding (bypass) is performed on the fourth bin (binIdx = 3).

[0682] In addition, for example, as ​ In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 7-4).

[0683] In ​In the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. Further, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context encoding is performed. Further, in the third bin (binIdx = 2), the index ctxInc = B2 is assigned and context encoding is performed. Further, in the fourth bin (binIdx = 3), the index ctxInc = B3 is assigned and context encoding is performed.

[0684] Note that in ​ and ​ 's table, non-overlapping unique values are set in the indices A0, A1, B1, B2, and B3.

[0685] ​ Examples of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each of these methods are shown in the table of

[0686] As described above, in any of the cases of Method 7-1 to Method 7-4, compared with the case of Method 0, the number of contexts required for encoding can be reduced. That is, by applying Method 7, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0687] In addition, in any of the cases of Method 7-1 to Method 7-3, the number of context coding bins required for coding can be reduced compared to the case of Method 0. Note that in the case of Method 7-4, the number of context coding bins required for coding is equivalent to the number of context coding bins required for coding in the case of Method 0. That is, by applying Method 7, bypass coding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins and an increase in the processing amount (throughput) while suppressing a decrease in coding efficiency.

[0688] Then, as described above, for color differences in color difference formats 4:2:2 or 4:4:4, which have more information than the color difference format 4:2:0, the application of the adaptive orthogonal transform can be controlled. Therefore, a decrease in coding efficiency can be suppressed.

[0689] As described above, by applying Method 7, an increase in the load of the coding process can be suppressed.

[0690] Note that the control parameters for transform skip or adaptive orthogonal transform can be signaled (encoded) for each treeType instead of the color component ID (cIdx). That is, [cIdx] for each control parameter can be replaced with [treeType].

[0691] In addition, the above method of assigning context variables to each bin in the bin sequence of the adaptive orthogonal transform identifier mts_idx can also be applied to other syntax elements related to orthogonal transforms and the like. For example, this method can also be applied to the secondary transform identifier st_idx and the transform skip flag ts_flag.

[0692] <8-2. Decoding of Transform Skip Flag and Adaptive Orthogonal Transform Identifier>

[0693] Similarly, in the case of decoding, in the case where the color difference array type is greater than 1 (i.e., in the case of color difference formats 4:2:2 or 4:4:4), the transform skip flag and the adaptive orthogonal transform identifier are signaled (decoded) for each component ID (cIdx) so that transform skip and adaptive orthogonal transform can be applied to the color difference components (Method 7).

[0694] An example of the syntax of the transform unit in this case is ​ as shown. As ​ shown, the transform mode is signaled for the luminance (Y), color difference (Cb), and color difference (Cr) as shown in the 17th, 20th, and 23rd lines (gray lines) from the top of this syntax. ​An example of the syntax of the transform_mode in this case is shown. As ​ shown, in the fifth line (gray line) from the top of this syntax, the transform skip flag transform_skip_flag[x0][y0][cIdx] is signaled. Further, in the eighth line (gray line) from the top of this syntax, the adaptive orthogonal transform identifier tu_mts_idx[x0][y0][cIdx] is signaled. That is, the transform skip flag and the adaptive orthogonal transform identifier are signaled for each component ID (cIdx).

[0695] By doing so, the application of the inverse adaptive orthogonal transform can be controlled for color differences in color difference formats 4:2:2 or 4:4:4, which have more information than the color difference format 4:2:0. Therefore, a reduction in coding efficiency can be suppressed.

[0696] Then, in this case, the context variable ctx can be assigned to the first bin (binIdx = 0) in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0, and context decoding can be performed.

[0697] For example, as ​ shown in the table of A, the context variable ctx can be assigned to the first bin (binIdx = 0) in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 7-1).

[0698] In ​ the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, while when the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. Further, bypass decoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).

[0699] Further, for example, as ​In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context decoding can be performed on the first bin. A predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context decoding can be performed on the second bin. Also, bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 7-2).

[0700] In ​ In the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. Further, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context decoding is performed. Also, bypass decoding (bypass) is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).

[0701] Further, for example, as ​ In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second and third bins in the bin sequence, and context decoding can be performed on the second and third bins. Also, bypass decoding can be performed on the fourth bin in the bin sequence (Method 7-3).

[0702] In ​In the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. Further, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context decoding is performed. Further, in the third bin (binIdx = 2), the index ctxInc = B2 is assigned and context decoding is performed. Further, bypass decoding is performed on the fourth bin (binIdx = 3).

[0703] Further, for example, as ​ shown in the table of B, according to whether the component ID (cIdx) of the transform block is 0 ((cIdx == 0)?), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 7-4).

[0704] In ​ the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx == 0), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the component ID (cIdx) of the transform block is not 0 (cIdx > 0), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. Further, in the second bin (binIdx = 1), the index ctxInc = B1 is assigned and context decoding is performed. Further, in the third bin (binIdx = 2), the index ctxInc = B2 is assigned and context decoding is performed. Further, in the fourth bin (binIdx = 3), the index ctxInc = B3 is assigned and context decoding is performed.

[0705] Note that in ​ and ​ the tables, non-overlapping unique values are set in the indexes A0, A1, B1, B2, and B3.

[0706] The number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are similar to those used for encoding ( ​ ).

[0707] As described above, in any of the cases of Method 7-1 to Method 7-4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 7, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.

[0708] In addition, in any of the cases of Method 7-1 to Method 7-3, compared with the case of Method 0, the number of context coding bins required for decoding can be reduced. Note that in the case of Method 7-4, the number of context coding bins required for decoding is equivalent to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 7, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context coding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in coding efficiency.

[0709] Then, as described above, for color differences in the color difference formats 4:2:2 or 4:4:4, which have more information than the color difference format 4:2:0, the application of the inverse adaptive orthogonal transform can be controlled. Therefore, a decrease in coding efficiency can be suppressed.

[0710] As described above, by applying Method 7, an increase in the load of the decoding process can be suppressed.

[0711] Note that the control parameters for transform skip or adaptive orthogonal transform can be signaled (decoded) for each treeType instead of the color component ID (cIdx). That is, [cIdx] for each control parameter can be replaced with [treeType].[[]END]]

[0712] In addition, the above method of assigning context variables to each bin in the bin sequence of the adaptive orthogonal transform identifier mts_idx can also be applied to other syntax elements related to orthogonal transforms, etc. For example, this method can also be applied to the secondary transform identifier st_idx and the transform skip flag ts_flag.

[0713] <8-3. Encoding side>

[0714] <Configuration>

[0715] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side of the first embodiment. That is, the image encoding apparatus 100 in this case has the same as that referred to ​The described configuration is similar to the configuration. Additionally, in this case, the coding unit 115 has a configuration similar to the one ​ described.

[0716] <Flow of encoding processing>

[0717] Furthermore, in this case, the image coding device 100 performs processing that is substantially similar to that in the case of the first embodiment. That is, the image coding process performed by the image coding device 100 in this case is executed through a process similar to the one described in the flowchart with reference to ​ .

[0718] Reference will be made to ​ for an example of the flowchart describing the process of the coding process for encoding the adaptive orthogonal transform identifier performed by the coding unit 115 in this case.

[0719] In this coding process, the processing in steps S551 and S552 is performed similarly to the processing in ​ in steps S131 and S132. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bins (i.e., context coding is selected). For example, the selection unit 132 selects context coding as the coding method for the bins according to ​ and ​ shown in any one of the tables (i.e., by applying any one of methods 7-1 to 7-4).

[0720] In step S553, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bins according to whether the component is luminance (Y) ((cIdx == 0)?).

[0721] For example, when the component is luminance (Y) (cIdx == 0), the context setting unit 133 assigns the index ctxInc = A0 to the bins. Additionally, when the component is not luminance (Y) (cIdx > 0), the context setting unit 133 assigns the index ctxInc = A1 to the bins.

[0722] Then, the context coding unit 134 performs arithmetic coding using the context variable. That is, context coding is performed.

[0723] The processing in steps S554 to S558 is also performed similarly to the processing in ​ in steps S134 to S138. In step S558, when it is determined to terminate the coding, the coding process is terminated.

[0724] By performing each of the processes described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying Method 7 (e.g., any one of Methods 7-1 to 7-4). Accordingly, an increase in the memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.

[0725] <8-4. Decoding side>

[0726] <Configuration>

[0727] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to the configuration of the decoding side of the first embodiment. That is, the image decoding apparatus 200 in this case has a configuration similar to the configuration described with reference to ​ In addition, the decoding unit 212 in this case has a configuration similar to the configuration described with reference to ​

[0728] <Flow of decoding process>

[0729] In addition, the image decoding apparatus 200 in this case performs processing substantially similar to that in the case of the first embodiment. That is, the image decoding process performed by the image decoding apparatus 200 in this case is performed by a flow similar to the flow described in the flowchart with reference to ​

[0730] Reference will be made to ​ to describe an example of the flow of the decoding process for decoding the encoded data of the adaptive orthogonal transform identifier, which is performed by the decoding unit 212 in this case.

[0731] In this decoding process, the process of step S571 is performed in a manner similar to the process of step S231 in ​ . That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bins (i.e., selects context decoding). For example, the selection unit 231 selects context decoding as the decoding method for the bins according to any one of the tables shown in ​ and Figure 43 (i.e., by applying any one of Methods 7-1 to 7-4).

[0732] In step S572, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bins according to whether the component is luminance (Y) ((cIdx == 0)?).

[0733] ​​For example, when the component is luminance (Y) (cIdx == 0), the context setting unit 232 assigns the index ctxInc = A0 to the bin. In addition, when the component is not luminance (Y) (cIdx > 0), the context setting unit 232 assigns the index ctxInc = A1 to the bin.

[0734] Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.

[0735] The processing in steps S573 to S578 is also similar to Figure 13 the processing in steps S233 to S238 in

[0736] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 7 (for example, any one of Methods 7-1 to 7-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.

[0737] <9. Appendix>

[0738] <Combination>

[0739] As long as there is no contradiction, the present technology described in each of the above embodiments can be combined and applied with the present technology described in any other embodiment.

[0740] <Computer>

[0741] The above series of processes can be executed by hardware or by software. In the case where this series of processes is executed by software, a program for configuring the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, a computer capable of executing various functions by installing various programs, such as a general-purpose personal computer, and so on.

[0742] Figure 47 is a block diagram showing a configuration example of the hardware of a computer that executes the above series of processes by a program.

[0743] In Figure 47 the computer 800 shown, a central processing unit (CPU) 801, a read-only memory (ROM) 802, and a random access memory (RAM) 803 are connected to each other via a bus 804.

[0744] The input / output interface 810 is also connected to the bus 804. The input unit 811, output unit 812, storage unit 813, communication unit 814, and driver 815 are connected to the input / output interface 810.

[0745] The input unit 811 includes, for example, a keyboard, mouse, microphone, touchpad, input terminal, etc. The output unit 812 includes, for example, a display, speaker, output terminal, etc. The storage unit 813 includes, for example, a hard disk, RAM disk, and non-volatile memory, etc. The communication unit 814 includes, for example, a network interface. The driver 815 drives a removable medium 821, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0746] In the computer configured as described above, the CPU 801 loads, for example, a program stored in the storage unit 813 into the RAM 803 via the input / output interface 810 and the bus 804, and executes the program so that the above-described series of processes are performed. In addition, the RAM 803 appropriately stores data and the like required for the CPU 801 to perform various types of processing.

[0747] For example, a program to be executed by the computer can be recorded and applied on a removable medium 821, etc. as a packaged medium, and can be provided. In this case, by attaching the removable medium 821 to the driver 815, the program can be installed in the storage unit 813 via the input / output interface 810.

[0748] In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program can be received by the communication unit 814 and installed in the storage unit 813.

[0749] In addition to the above methods, the program can be pre-installed in the ROM 802 or the storage unit 813.

[0750] <Units of Information and Processing>

[0751] The data units in which the above various types of information are set and the data units to be processed by various types of processing are arbitrary and are not limited to the above examples. For example, these information and processing can be set for each transform unit (TU), transform block (TB), prediction unit (PU), prediction block (PB), coding unit (CU), largest coding unit (LCU), sub-block, block, tile, slice, picture, sequence, or component, or the data in these data units can be used. Of course, the data unit can be set for each information and processing, and it is not necessary to unify the data units for all information and processing. Note that the storage location of this information is arbitrary and can be stored in the headers, parameters, etc. of the above data units. In addition, the information can be stored in multiple locations.

[0752] <Control information>

[0753] The control information of the present technology described in the above embodiments can be sent from the encoding side to the decoding side. For example, the control information for controlling whether to allow (or prohibit) the application of the above present technology (e.g., enabled_flag) can be sent. In addition, for example, the control information indicating the object to which the above present technology is applied (or the object to which the present technology is not applied) (e.g., present_flag) can be sent. For example, the control information for specifying the block size (upper limit, lower limit, or both), frame, component, layer, etc. for applying the present technology (or allowing or prohibiting the application) can be sent.

[0754] <Object to which the present technology is applicable>

[0755] The present technology can be applied to any image encoding / decoding method. That is, the specifications of various types of processing for image encoding / decoding such as transformation (inverse transformation), quantization (inverse quantization), encoding (decoding), and prediction are arbitrary and are not limited to the above examples, as long as there is no contradiction with the above present technology. In addition, as long as there is no contradiction with the above present technology, some parts of the processing can be omitted.

[0756] In addition, the present technology can be applied to a multi-viewpoint image encoding / decoding system that performs encoding / decoding of multi-viewpoint images including multiple viewpoints (views). In this case, the present technology is simply applied to the encoding / decoding of each viewpoint (view).

[0757] In addition, the present technology can be applied to a hierarchical image encoding (scalable encoding) / decoding system that encodes / decodes a hierarchical image of multiple layers (hierarchical) to have a scalable function for a predetermined parameter. In this case, the present technology is simply applied to the encoding / decoding of each layer.

[0758] In addition, in the above description, the image encoding device 100 and the image decoding device 200 have been described as application examples of the present technology, but the present technology can be applied to any configuration.

[0759] The present technology can be applied to, for example, various electronic devices such as transmitters and receivers in satellite broadcasting (such as television receivers and mobile phones), cable broadcasting such as cable television, distribution on the Internet, and distribution to terminals via cellular communication, or devices that record images on media such as optical discs, magnetic disks, and flash memories and reproduce images from these storage media (for example, hard disk recorders and imaging devices).

[0760] In addition, for example, the present technology can be implemented as a configuration of a part of a device, such as a processor (for example, a video processor) such as a system large scale integration (LSI), a module (for example, a video module) using multiple processors, a unit (for example, a video unit) using multiple modules, or a set (for example, a video set) in which other functions are added to the unit (that is, a configuration of a part of a device).

[0761] In addition, for example, the present technology can also be applied to a network system including multiple devices. For example, the present technology can be implemented as cloud computing in which multiple devices share and collaboratively process via a network. For example, the present technology can be implemented as a cloud service that provides services related to images (moving images) to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, or an Internet of Things (IoT) device.

[0762] Note that in this specification, the term "system" means a collection of multiple configuration elements (devices, modules (components), etc.), and it does not matter whether all the configuration elements are in the same housing. Therefore, both multiple devices housed in separate housings and connected via a network and a single device that houses multiple modules in one housing are systems.

[0763] <Fields and Applications to which the Present Technology is Applicable>

[0764] Systems, devices, processing units, etc. to which the present technology is applied can be used in any fields such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, household appliances, weather, and natural monitoring. In addition, the use in any field is also arbitrary.

[0765] For example, the present technology can be applied to systems and devices provided for providing content for appreciation and the like. In addition, for example, the present technology can also be applied to systems and devices for traffic such as traffic condition monitoring and autonomous driving control. In addition, for example, the present technology can also be applied to systems and devices provided for security. In addition, for example, the present technology can be applied to systems and devices provided for automatic control of machines and the like. In addition, for example, the present technology can also be applied to systems and devices provided for agriculture or animal husbandry. In addition, the present technology can also be applied to systems and devices for monitoring natural states such as volcanoes, forests, and oceans, wild animals, and the like. In addition, for example, the present technology can also be applied to systems and devices provided for sports.

[0766] <Other>

[0767] Note that the "flag" in this specification is information for identifying multiple states, and includes not only information for identifying two states of true (1) and false (0), but also information capable of identifying three or more states. Therefore, the values that the "flag" can take can be, for example, binary values 1 / 0 or can be ternary values or more - ary values. That is, the number of bits constituting the "flag" is arbitrary and can be 1 bit or multiple bits. In addition, it is assumed that the identification information (including the flag) is not only in the form of including the identification information in the bit stream, but also in the form of including the difference information between the identification information and specific reference information in the bit stream. Therefore, in this specification, the "flag" and the "identification information" include not only the information itself, but also the difference information with respect to the reference information.

[0768] In addition, various types of information (such as metadata) about the encoded data (bit stream) can be sent or recorded in any form as long as the various types of information are associated with the encoded data. Here, the term "associated" means that, for example, when processing one data, other data can be used (linked). That is, the data associated with each other can be collected as one data or can be separate data. For example, information associated with the encoded data (image) can be sent on a transmission path different from the transmission path of the encoded data (image). In addition, for example, information associated with the encoded data (image) can be recorded on a recording medium different from the encoded data (image) (or another recording area of the same recording medium). Note that this "association" can be a part of the data, rather than the entire data. For example, an image and the information corresponding to the image can be associated with each other in any unit such as multiple frames, one frame, or a part of a frame.

[0769] Note that in this specification, terms such as "combine", "multiplex", "add", "integrate", "include", "store", and "insert" mean putting multiple things into one thing, for example, putting encoded data and metadata into one data, and mean a method of the above - mentioned "association".

[0770] In addition, the embodiments of the present technology are not limited to the above embodiments, and various modifications can be made without departing from the gist of the present technology.

[0771] For example, a configuration described as a single device (or processing unit) can be divided and configured into multiple devices (or processing units). Conversely, configurations described as multiple devices (or processing units) can be configured together as a single device (or processing unit). In addition, configurations other than the above configurations can be added to the configuration of each device (or each processing unit). Further, a part of the configuration of a specific device (or processing unit) can be included in the configuration of another device (or another processing unit), as long as the configuration and operation of the system as a whole are substantially the same.

[0772] In addition, for example, the above program can be executed by any device. In this case, the device only needs to have the necessary functions (function blocks, etc.) and obtain the necessary information.

[0773] In addition, for example, each step of a flowchart can be executed by a single device, or can be shared and executed by multiple devices. In addition, in the case where a step includes multiple processes, the multiple processes can be executed by a single device, or can be shared and executed by multiple devices. In other words, the multiple processes included in a single step can be executed as processes of multiple steps. Conversely, processes described as multiple steps can be executed together as a single step.

[0774] Note that in a program executed by a computer, the processes describing the steps of the program can be executed chronologically in the order described in this specification, or can be executed in parallel or individually at the necessary time, such as when called. That is, as long as there is no contradiction, the processes of each step can be executed in an order different from the above order. In addition, the processes describing the steps of the program can be executed in parallel with the processes of another program, or can be executed in combination with the processes of another program.

[0775] In addition, for example, as long as there is no contradiction, multiple technologies related to the present technology can be independently implemented as a whole. Of course, any number of the present technologies can be implemented together. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. In addition, part or all of any of the above-presented technologies can be implemented in combination with another technology not described above.

[0776] Note that the present technology can also have the following configuration.

[0777] (1) An image processing apparatus, comprising:

[0778] A coding unit configured to assign a predetermined context variable to a first bin in a bin sequence and perform context coding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image coding.

[0779] (2) The image processing apparatus according to (1), wherein

[0780] the coding unit assigns different predetermined context variables to the first to fourth bins in the bin sequence and performs context coding on the first to fourth bins in the bin sequence.

[0781] (3) The image processing apparatus according to (1), wherein

[0782] the coding unit assigns different predetermined context variables to the first to third bins in the bin sequence and performs context coding on the first to third bins in the bin sequence, and performs bypass coding on the fourth bin in the bin sequence.

[0783] (4) The image processing apparatus according to (1), wherein

[0784] the coding unit assigns different predetermined context variables to the first and second bins in the bin sequence and performs context coding on the first and second bins in the bin sequence, and performs bypass coding on the third and fourth bins in the bin sequence.

[0785] (5) The image processing apparatus according to (1), wherein

[0786] the coding unit assigns a predetermined context variable to the first bin in the bin sequence and performs context coding on the first bin in the bin sequence, and performs bypass coding on the second to fourth bins in the bin sequence.

[0787] (6) The image processing apparatus according to any one of (1) to (5), wherein

[0788] the coding unit binarizes the adaptive orthogonal transform identifier into a bin sequence and encodes the bin sequence, the bin sequence being configured by one bit indicating whether the transform type is different from the transform type DCT2×DCT2 and two bits indicating other transform types.

[0789] (7) An image processing method, comprising:

[0790] Assign a predetermined context variable to the first bin in the bin sequence, and perform context encoding on the first bin in the bin sequence, where the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0791] (8) An image processing apparatus, comprising:

[0792] The encoding unit is configured to assign a context variable based on a parameter regarding a block size to the first bin in the bin sequence, and perform context encoding on the first bin in the bin sequence, where the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0793] (9) The image processing apparatus according to (8), wherein,

[0794] The parameter regarding the block size is the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied.

[0795] (10) The image processing apparatus according to (8), wherein,

[0796] The parameter regarding the block size is the minimum value among the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied and a predetermined threshold.

[0797] (11) The image processing apparatus according to (8), wherein,

[0798] The parameter regarding the block size is the result of right-shifting the minimum value among the difference between the logarithm of the long side of the transform block and the logarithm of the minimum transform block size to which the adaptive orthogonal transform can be applied and a predetermined threshold.

[0799] (12) The image processing apparatus according to (8), wherein,

[0800] The encoding unit assigns a context variable and performs context encoding according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold.

[0801] (13) The image processing apparatus according to any one of (8) to (12), wherein,

[0802] The encoding unit binarizes the adaptive orthogonal transform identifier into a bin sequence, and encodes the bin sequence, where the bin sequence is configured by one bit indicating whether the transform type is different from the transform type DCT2×DCT2 and two bits indicating other transform types.

[0803] (14) The image processing apparatus according to (8), wherein,

[0804] The encoding unit binarizes the adaptive orthogonal transform identifier for each component and performs encoding.

[0805] (15) An image processing method, comprising:

[0806] Assigning a context variable based on a parameter regarding a block size to a first bin in a bin sequence, and performing context encoding on the first bin in the bin sequence, the bin sequence being obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.

[0807] List of reference numerals

[0808] 100 Image encoding apparatus

[0809] 115 Encoding unit

[0810] 131 Binarization unit

[0811] 132 Selection unit

[0812] 133 Context setting unit

[0813] 134 Context encoding unit

[0814] 135 Bypass encoding unit

[0815] 200 Image decoding apparatus

[0816] 212 Decoding unit

[0817] 231 Selection unit

[0818] 232 Context setting unit

[0819] 233 Context decoding unit

[0820] 234 Bypass decoding unit

[0821] 235 Inverse binarization unit

Claims

1. An image processing apparatus, comprising: A circuit system configured to: Assign a context variable to a first bin in a bin sequence based on a parameter regarding a tree type, the bin sequence being obtained by binarizing a second-order orthogonal transform identifier indicating a mode of a second-order orthogonal transform in image coding, and Perform context coding on the first bin in the bin sequence.

2. The image processing apparatus according to claim 1, wherein, The circuit system assigns different predetermined context variables to the first bin to the last bin in the bin sequence, and performs the context coding on the first bin to the last bin in the bin sequence.

3. The image processing apparatus according to claim 1, wherein The circuit system performs bypass coding on at least a second bin among the remaining bins in the bin sequence.

4. The image processing apparatus according to claim 2, wherein, The predetermined context variables are non-overlapping.

5. The image processing apparatus according to claim 1, wherein The circuit system binarizes the second-order orthogonal transform identifier into a bin sequence configured by one bit indicating whether a transform type is different from a transform type DCT2×DCT2 and two bits indicating an additional transform type, and codes the bin sequence.

6. The image processing apparatus according to claim 1, wherein, The parameter regarding the tree type indicates one of a single tree or a double tree.

7. An image processing method performed by a circuit system, the image processing method comprising: Assigning a context variable to a first bin in a bin sequence based on a parameter regarding a tree type, the bin sequence being obtained by binarizing a second-order orthogonal transform identifier indicating a mode of a second-order orthogonal transform in image coding, and Performing context coding on the first bin in the bin sequence.

8. The image processing method according to claim 7 further includes: Assigning different predetermined context variables to the first bin to the last bin in the bin sequence, and performing the context coding on the first bin to the last bin in the bin sequence.

9. An image processing apparatus, comprising: A circuit system configured to: Perform decoding of a first bin in a bin sequence encoded according to a mode of a second-order orthogonal transform in image coding, the first bin in the bin sequence being assigned a context variable based on a parameter regarding a tree type, the bin sequence being obtained by binarizing a second-order orthogonal transform identifier indicating the mode of the second-order orthogonal transform in the image coding.

10. The image processing apparatus according to claim 9, wherein, Different predetermined context variables are assigned to the first bin to the last bin in the bin sequence, and context coding is performed on the first bin to the last bin in the bin sequence.

11. The image processing apparatus according to claim 9, wherein, The circuit system decodes at least a second bin in the bin sequence that has been encoded via bypass coding.

12. The image processing apparatus according to claim 10, wherein, The predetermined context variables are non-overlapping.

13. The image processing apparatus according to claim 9, wherein, The parameter regarding the tree type indicates one of a single tree or a double tree.

14. An image processing method performed by a circuit system, the image processing method comprising: Performing decoding of a first bin in a bin sequence encoded according to a mode of a second-order orthogonal transform in image coding, the first bin in the bin sequence being assigned a context variable based on a parameter regarding a tree type, the bin sequence being obtained by binarizing a second-order orthogonal transform identifier indicating the mode of the second-order orthogonal transform in the image coding.

15. The image processing method according to claim 14, wherein, Respectively different predetermined context variables are assigned to the first bin to the last bin in the bin sequence, and context encoding is performed for the first bin to the last bin in the bin sequence.