Image processing device and method
By assigning the context variable to the first bin in the bin sequence of the adaptive orthogonal transform identifier in image encoding and performing context encoding, the problem of increasing the context variable during image encoding and decoding is solved, and the effect of reducing memory usage and processing load is achieved.
Patent Information
- Application Number
- CN202080042943.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-19
- Filing Date
- 2020-05-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-05-08
AI Technical Summary
In the prior art, unnecessary increase in context variables during image encoding and decoding results in an increase in memory usage and processing load.
By assigning a predetermined context variable to the first bin in the bin sequence of adaptive orthogonal transform identifiers in the image encoding, and performing context encoding on the bin, an unnecessary increase of the context variable is suppressed.
It effectively suppresses the increase in the load of encoding and decoding processing, reduces the memory usage, and improves the processing efficiency.
Smart Images

Figure CN113940072B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing apparatus and method, and more particularly to an image processing apparatus and method capable of suppressing an increase in load. Background Art
[0002] Conventionally, in image encoding, an adaptive orthogonal transform identifier mts_idx has been signaled (encoded / decoded) as mode information about adaptive orthogonal transform (multiple transform selection (MTS)). For encoding of the adaptive orthogonal transform identifier mts_idx, context encoding is applied in which the adaptive orthogonal transform identifier mts_idx is binarized and a context variable ctx is assigned to each bin in a bin sequence bins to perform arithmetic encoding. In addition, context decoding corresponding to the context encoding is applied to decoding of the encoded data of the adaptive orthogonal transform identifier mts_idx.
[0003] Reference List
[0004] Non-patent literature
[0005] Non-Patent Literature 1: Benjamin Bross, Jianle Chen, Shan Liu, “Versatile Video Coding (Draft 5)”, JVET-N1001v8, 14th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019. Summary of the invention
[0006] Problem that the invention aims to solve
[0007] However, in the case of such a method, there is a possibility that context variables increase unnecessarily and memory usage increases unnecessarily. That is, there is a possibility that the load of encoding processing and decoding processing increases.
[0008] The present disclosure has been made in view of the above-mentioned problems, and an object is to suppress an increase in the load of encoding processing and decoding processing.
[0009] Solutions to technical problems
[0010] An image processing device according to one aspect of the present technology is an image processing device comprising an encoding unit configured to assign a predetermined context variable to a first bin in a bin sequence and to perform context encoding on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0011] An image processing method according to one aspect of the present technology is an image processing method comprising assigning a predetermined context variable to a first bin in a bin sequence, and performing context encoding on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0012] An image processing device according to another aspect of the present technology is an image processing device comprising an encoding unit configured to assign a context variable based on a parameter regarding a block size to a first bin in a bin sequence, and to perform context encoding on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0013] An image processing method according to another aspect of the present technology is an image processing method comprising assigning a context variable based on a parameter regarding a block size to a first bin in a bin sequence, and performing context encoding on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0014] In an image processing device and method according to one aspect of the present technology, a predetermined context variable is assigned to a first bin in a bin sequence, and context encoding is performed on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0015] In an image processing device and method according to another aspect of the present technology, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence, and context encoding is performed on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a diagram for describing an example of a state in which an adaptive orthogonal transform identifier is encoded.
[0017] Figure 2 is a diagram for describing an example of a state in which an adaptive orthogonal transform identifier is encoded.
[0018] Figure 3 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0019] Figure 4 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0020] Figure 5 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0021] Figure 6 is a block diagram showing a main configuration example of an image encoding device.
[0022] Figure 7 is a block diagram showing a main configuration example of an encoding unit.
[0023] Figure 8 is a flowchart for describing an example of the flow of image encoding processing.
[0024] Fig. 9 is a flowchart for describing an example of the flow of image encoding processing.
[0025] Fig.10 is a block diagram showing a main configuration example of an image decoding device.
[0026] Fig.11 is a block diagram showing a main configuration example of a decoding unit.
[0027] Fig.12 is a flowchart for describing an example of the flow of image decoding processing.
[0028] Fig.13 is a flowchart for describing an example of the flow of a decoding process.
[0029] Fig.14 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0030] Fig.15 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0031] Fig.16 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0032] Fig.17 is a flowchart for describing an example of the flow of encoding processing.
[0033] Fig.18 is a flowchart for describing an example of the flow of a decoding process.
[0034] Fig.19 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0035] Fig. 20 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0036] Fig.21 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0037] Fig. 22 is a flowchart for describing an example of the flow of encoding processing.
[0038] Fig.23 is a flowchart for describing an example of the flow of a decoding process.
[0039] Fig.24 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0040] Fig.25 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0041] Fig.26 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0042] Fig. 27 is a flowchart for describing an example of the flow of encoding processing.
[0043] Fig.28 is a flowchart for describing an example of the flow of a decoding process.
[0044] Fig.29 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0045] Fig.30 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0046] Fig.31 is a diagram showing an example of a ratio of an area of a coding block to an area of a CTU.
[0047] Fig.32 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0048] Fig.33 is a flowchart for describing an example of the flow of encoding processing.
[0049] Fig.34 is a flowchart for describing an example of the flow of a decoding process.
[0050] Fig.35 is a diagram showing an example of a binarized state of an adaptive orthogonal transform identifier.
[0051] Fig.36 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0052] Fig.37 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0053] Fig.38 is a flowchart for describing an example of the flow of encoding processing.
[0054] Fig.39 is a flowchart for describing an example of the flow of a decoding process.
[0055] Fig.40 is a diagram showing an example of syntax regarding a transform unit.
[0056] Fig.41 is a diagram showing an example of syntax regarding orthogonal transform modes.
[0057] Fig.42 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0058] Fig.43 is a diagram showing an example of assigning a context variable to each bin in a bin sequence of an adaptive orthogonal transform identifier.
[0059] Fig.44 is a diagram showing a comparative example of the number of contexts, the number of context encoding bins, and the number of bypass encoding bins for each method.
[0060] Fig.45 is a flowchart for describing an example of the flow of encoding processing.
[0061] Fig.46 is a flowchart for describing an example of the flow of a decoding process.
[0062] Fig.47 is a block diagram showing a main configuration example of a computer. DETAILED DESCRIPTION
[0063] Hereinafter, a mode for implementing the present disclosure (hereinafter, referred to as an embodiment) will be described. Note that the description will be given in the following order.
[0064] 1. Encoding of Adaptive Orthogonal Transform Identifier
[0065] 2. First Implementation
[0066] 3. Second Implementation
[0067] 4. Third Implementation
[0068] 5. Fourth embodiment
[0069] 6. Fifth Implementation
[0070] 7. Sixth Implementation
[0071] 8. Seventh Implementation Method
[0072] 9. Appendix
[0073] <1. Encoding of Adaptive Orthogonal Transform Identifier>
[0074] <1-1. Documents supporting technical content and technical terms, etc.>
[0075] The scope disclosed in the present technology includes not only the contents described in the embodiments but also the contents described in the following non-patent documents and the like known at the time of filing the application and the contents of other documents referred to in the following non-patent documents.
[0076] Non-Patent Document 1: (described above).
[0077] Non-patent document 2: ITU-T Recommendation H.264 (04 / 2017) “Advanced video coding for generic audiovisual services”, April 2017.
[0078] Non-patent document 3: ITU-T Recommendation H.265 (12 / 2016) “High efficiency video coding”, April 2016.
[0079] Non-Patent Literature 4: J. Chen, E. Alshina, G. J. Sullivan, J. R. Ohm, and J. Boyce, “Algorithm Description of Joint Exploration Test Model (JEM7)”, JVET-G1001, Seventh Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Turin, Italy, July 13, 2017 to July 21, 2017.
[0080] Non-Patent Literature 5: Bross, J. Chen and S. Liu, “Versatile Video Coding (Draft 3)”, JVET-L1001, 12th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Macau, China, October 3, 2018 to October 12, 2018.
[0081] Non-Patent Literature 6: J. Chen, Y. Ye, and S. Kim, “Algorithm description for Versatile Video Coding and Test Model 3 (VTM 3)”, JVET-L1002, 12th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG16WP 3 and ISO / IEC JTC1 / SC 29 / WG 11: Macau, China, October 3, 2018 to October 12, 2018.
[0082] Non-Patent Literature 7: Jianle Chen, Yan Ye, and Seung Hwan Kim: “Algorithm descriptionfor Versatile Video Coding and Test Model 5 (VTM 5)”, JVET-N1002-v2, 14th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.
[0083] Non-Patent Literature 8: Moonmo Koo, Jaehyun Lim, Mehdi Salehifar, and Seung Hwan Kim, “CE6: Reduced Secondary Transform (RST) (CE6-3.1)”, JVET-N0193, 14th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.
[0084] Non-Patent Document 9: Mischa Siekmann, Martin Winken, Heiko Schwarz, and Detlev Marpe, “CE6-related: Simplification of the Reduced Secondary Transform,” JVET-N0555-v3, 14th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.
[0085] Non-Patent Literature 10: C. Rosewarne and J. Gan, “CE6-related: RST binarization,” JVET-N0105-v2, 14th Meeting of the Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11: Geneva, Switzerland, March 19, 2019 to March 27, 2019.
[0086] That is, the contents described in the above non-patent literature are also used as the basis for determining the support requirements. For example, even if the quadtree block structure and the quadtree plus binary tree (QTBT) block structure described in the above non-patent literature are not directly described in the examples, these contents fall within the disclosure scope of the present technology and meet the support requirements of the claims. In addition, for example, even if technical terms such as parsing, grammar, and semantics are not directly described in the examples, these technical terms similarly fall within the disclosure scope of the present technology and meet the support requirements of the claims.
[0087] In addition, in the present specification, unless otherwise specified, a "block" (not a block indicating a processing unit) used for description as a partial area or a processing unit of an image (picture) indicates an arbitrary partial area in a picture, and the size, shape, characteristics, etc. of the block are not limited. For example, a "block" includes an arbitrary partial area (unit of processing) such as a transform block (TB), a transform unit (TU), a prediction block (PB), a prediction unit (PU), a minimum coding unit (SCU), a coding unit (CU), a maximum coding unit (LCU), a coding tree block (CTB), a coding tree unit (CTU), a transform block, a subblock, a macroblock, a tile, or a slice described in the above non-patent literature.
[0088] In addition, when specifying the size of such a block, the block size can be specified not only directly but also indirectly. For example, the block size can be specified using identification information for identifying the size. In addition, for example, the block size can be specified by a ratio or difference with the size of a reference block (e.g., LCU, SCU, etc.). For example, in the case where information for specifying the block size is sent as a syntax element, etc., the information for indirectly specifying the size as described above can be used as the information. Through this configuration, the amount of information can be reduced, and in some cases the coding efficiency can be improved. In addition, the specification of the block size also includes the specification of a range of block sizes (e.g., the specification of a range of allowable block sizes, etc.).
[0089] Furthermore, in the present specification, encoding includes not only the entire process of transforming an image into a bit stream, but also a part of the process. For example, encoding includes not only a process including prediction processing, orthogonal transform, quantization, arithmetic coding, etc., but also a process collectively referred to as quantization and arithmetic coding, including a prediction processing, quantization, and arithmetic coding process, and the like. Similarly, decoding includes not only the entire process of transforming a bit stream into an image, but also a part of the process. For example, decoding includes not only a process including inverse arithmetic decoding, inverse quantization, inverse orthogonal transform, prediction processing, etc., but also a process including inverse arithmetic decoding and inverse quantization, a process including inverse arithmetic decoding, inverse quantization, and prediction processing, and the like.
[0090] <1-2. Context encoding / context decoding of adaptive orthogonal transform identifier>
[0091] Conventionally, in image encoding and decoding, an adaptive orthogonal transform identifier mts_idx has been signaled (encoded / decoded) as mode information on adaptive orthogonal transform (multiple transform selection (MTS)). Context encoding using the following context has been applied to encoding of the adaptive orthogonal transform identifier mts_idx.
[0092] First, if Figure 1In the table shown in A of , the adaptive orthogonal transform identifier mts_idx is binarized by a truncated unary (TU) code to obtain a bin sequence bins. Note that the TU code is equivalent to a truncated Rice (TR) code with a Rice parameter cRiceParam=0.
[0093] Next, arithmetic coding is performed with reference to a context variable ctx corresponding to each binIdx (index indicating a bin number) in the bin sequence bins obtained by TR. The index for identifying the context variable ctx is referred to as ctxInc (or ctxIdx).
[0094] Specifically, Figure 1 In the table shown in B of FIG. 1 , the context variable ctx corresponding to the value of the CQT partition depth cqtDepth is assigned to the first bin (binIdx=0) in the bin sequence. The CQT partition depth cqtDepth represents the depth at which the CTU is partitioned by the quadtree of the CU. Figure 1 In the example of B, the smaller one of the CQT partition depth cqtDepth and 5 is set to the context index ctxInc corresponding to the first bin (binIdx=0) in the bin sequence (ctxInc=min(cqtDepth,5)). That is, since the output (frequency) of the 0th order 1 in the adaptive orthogonal transform is likely to change according to the partition depth, in order to improve efficiency, the context is also correspondingly variable.
[0095] In addition, the context variable ctx (in Figure 1 In the example in B, ctxInc=6 to 8) are assigned to the second bin (binIdx=1) to the fourth bin (binIdx=3) in the bin sequence bins.
[0096] Note that Figure 2 In the table shown in A of , each bin in the bin sequence bins of the adaptive orthogonal transform identifier mts_idx can be interpreted as a flag corresponding to a transform type. In this example, the value of the first bin (binIdx=0) corresponds to a flag indicating whether the transform type is DCT2×DCT2 (0 indicates "yes" and 1 indicates "no"), the value of the second bin (binIdx=1) corresponds to a flag indicating whether the transform type is DST7×DST7 (0 indicates "yes" and 1 indicates "no"), the value of the third bin (binIdx=2) corresponds to a flag indicating whether the transform type is DCT8×DST7 (0 indicates "yes" and 1 indicates "no"), and the value of the fourth bin (binIdx=3) corresponds to a flag indicating whether the transform type is DST7×DCT8 (0 indicates "yes" and 1 indicates "no").
[0097] The coded data of the adaptive orthogonal transform identifier mts_idx has been decoded by a method corresponding to such coding. That is, context decoding using a context has been applied.
[0098] However, in the case of such context encoding and context decoding, there is a possibility that the processing load increases.
[0099] For example, the adaptive orthogonal transform identifier mts_idx does not appear in the case of a specific value of the CQT division depth cqtDepth. Therefore, there is a completely unused context variable ctx, and there is a possibility that the memory usage is unnecessarily increased due to the context variable ctx (there is a possibility that the memory capacity required for processing increases).
[0100] For example, in the case where all CU partitions are performed with a quadtree, when the CTU size = 128×128, the CU size corresponding to each CQT partition depth cqtDepth is as follows: Figure 2 As shown in the table in B. Since the adaptive orthogonal transform is not applicable to blocks larger than 32×32, the adaptive orthogonal transform identifier mts_idx does not appear for CUs of 128×128 and 64×64. Therefore, in the context variable ctx corresponding to the first bin in the bin sequence bins of the adaptive orthogonal transform identifier mts_idx, ctxInc=0 and 1 are not used at all. That is, due to these context variables ctx, there is a possibility that the memory usage is unnecessarily increased.
[0101] Therefore, a predetermined context variable is assigned to the first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0102] By doing so, an increase in the number of contexts of the bin sequence assigned to the adaptive orthogonal transform identifier can be suppressed, and thus an increase in memory usage can be suppressed and an increase in the encoding process and the load of the encoding process can be suppressed.
[0103] Furthermore, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0104] By doing so, an increase in the number of contexts of the bin sequence assigned to the adaptive orthogonal transform identifier can be suppressed, and thus an increase in memory usage can be suppressed and an increase in the encoding process and the load of the encoding process can be suppressed.
[0105] <1-3. Bypass Coding>
[0106] In addition, the selectivity of the transform type is discrete cosine transform (DCT) 2×DCT2, discrete sine transform (DST) 7×DST7, DCT8×DST7, DST7×DCT8 or DCT8×DCT8 in descending order. That is, the selectivity of each transform type is not uniform. Therefore, it is inefficient to similarly assign the context variable ctx to all transform types, which may unnecessarily increase the total number of context coding bins. As described above, each bin in the bin sequence bins of the adaptive orthogonal transform identifier mts_idx can be interpreted as a flag corresponding to the transform type. That is, it is inefficient to similarly assign the context variable ctx to each bin in the bin sequence bins, which may unnecessarily increase the total number of context coding bins.
[0107] As the total number of context coding bins increases in this way, there is a possibility that the processing amount (throughput) of context-based adaptive binary arithmetic coding (CABAC) increases.
[0108] Therefore, bypass coding is applied to bins corresponding to transform types with relatively low selectivity. By doing so, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in coding efficiency and suppress an increase in the amount of processing (throughput) of CABAC. That is, it is possible to suppress an increase in the load of encoding processing and decoding processing.
[0109] <2. First Embodiment>
[0110] <2-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0111] In this embodiment, the following is performed as in Figure 1 As has been performed in the table shown in B of , a context variable is assigned to each bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (method 0).
[0112] That is, a predetermined context variable (fixed (one-to-one correspondence) context variable) ctx is assigned to the first bin in the bin sequence, and context encoding is performed on the first bin (method 1).
[0113] For example, Figure 3 In the table shown in A, a predetermined context variable (index ctxInc for identifying the context variable ctx) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (method 1-1).
[0114] exist Figure 3 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, and bypass coding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0115] In addition, for example, Figure 3 In the table shown in B, predetermined context variables ctx (index ctxInc) different from each other can be assigned to the first bin and the second bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin and the second bin, and bypass encoding can be performed on the third bin and the fourth bin in the bin sequence (method 1-2).
[0116] exist Figure 3 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and bypass coding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0117] In addition, for example, Figure 4 In the table shown in A, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first to third bins, and bypass encoding can be performed on the fourth bin in the bin sequence (method 1-3).
[0118] exist Figure 4 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and bypass coding is performed on the fourth bin (binIdx=3).
[0119] In addition, for example, Figure 4In the table shown in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first to fourth bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first to fourth bins (method 1-4).
[0120] exist Figure 4 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context coding is performed on the fourth bin.
[0121] Note that Figure 3 and Figure 4 In the table, non-overlapping unique values are set in indexes A0, B1, B2 and B3.
[0122] Figure 5 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in the table. For example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 1-1, the number of contexts is 1, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 1-2, the number of contexts is 2, the number of context coding bins is 2, and the number of bypass coding bins is 2. Furthermore, in the case of method 1-3, the number of contexts is 3, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 1-4, the number of contexts is 4, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0123] As described above, in any case of Method 1-1 to Method 1-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 1, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0124] Furthermore, in any of the cases of Method 1-1 to Method 1-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 1-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 1, bypass coding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0125] As described above, by applying method 1, an increase in the load of encoding processing can be suppressed.
[0126] <2-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0127] Similarly, in the case of decoding, the following is performed as in Figure 1 As has been performed in the table shown in B, a context variable is assigned to each bin in a bin sequence of a binarized adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding.
[0128] That is, a predetermined context variable (fixed (one-to-one correspondence) context variable) ctx is assigned to the first bin in the bin sequence, and context decoding is performed on the first bin (method 1).
[0129] For example, Figure 3 In the table shown in A, a predetermined context variable (index ctxInc for identifying the context variable ctx) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (method 1-1).
[0130] exist Figure 3 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0131] In addition, for example, Figure 3 In the table shown in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first bin and the second bin in the bin sequence of the bin-valued adaptive orthogonal transform identifier, and context decoding can be performed on the first bin and the second bin, and bypass decoding can be performed on the third bin and the fourth bin in the bin sequence (method 1-2).
[0132] exist Figure 3 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0133] In addition, for example, Figure 4 In the table shown in A, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first to third bins, and bypass decoding can be performed on the fourth bin in the bin sequence (method 1-3).
[0134] exist Figure 4 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx=3).
[0135] In addition, for example, Figure 4 In the table shown in B, different predetermined context variables ctx (index ctxInc) can be assigned to the first to fourth bins in the bin sequence of the binned adaptive orthogonal transform identifier, and context decoding can be performed on the first to fourth bins (method 1-4).
[0136] exist Figure 4 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context decoding is performed on the fourth bin.
[0137] Note that even in the case of decoding, similar to the case of encoding, Figure 3 and Figure 4 Non-overlapping unique values are also set in indexes A0, B1, B2, and B3 in the table.
[0138] The number of contexts, the number of context bins, and the number of bypass bins for each of these methods are similar to those for encoding ( Figure 5 ).
[0139] As described above, in any case of Method 1-1 to Method 1-4, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 1, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0140] Furthermore, in any of the cases of Method 1-1 to Method 1-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 1-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 1, bypass decoding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0141] As described above, by applying method 1, an increase in the load of the decoding process can be suppressed.
[0142] <2-3. Encoding side>
[0143] <Image Coding Device>
[0144] Next, the encoding side will be described. Figure 6 : is a block diagram showing an example of the configuration of an image encoding device as one mode of an image processing device to which the present technology is applied. Figure 6 The image encoding device 100 shown in FIG. 1 is a device that encodes image data of a moving image. For example, the image encoding device 100 encodes image data of a moving image by an encoding method described in any one of non-patent documents 1 to 10.
[0145] Notice, Figure 6 shows the main processing units (blocks), data flows, etc., and Figure 6 That is, in the image encoding device 100, there may be Figure 6 The processing units shown as blocks in the diagram may exist without Figure 6 A process or data flow is shown as arrows or the like.
[0146] like Figure 6 As shown, the image encoding device 100 includes a control unit 101, a reordering buffer 111, a computing unit 112, an orthogonal transform unit 113, a quantization unit 114, an encoding unit 115, an accumulation buffer 116, an inverse quantization unit 117, an inverse orthogonal transform unit 118, a computing unit 119, a loop filtering unit 120, a frame memory 121, a prediction unit 122 and a rate control unit 123.
[0147] <Control Unit>
[0148] The control unit 101 divides the moving image data held by the reordering buffer 111 into blocks of processing units (CU, PU, transform block, etc.) based on the block size of the processing unit specified externally or in advance. In addition, the control unit 101 determines encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc.) to be provided to each block based on, for example, rate-distortion optimization (RDO).
[0149] The details of these encoding parameters will be described below. After determining the above encoding parameters, the control unit 101 provides the encoding parameters to each block. Specifically, the encoding parameters are as follows.
[0150] Header information Hinfo is provided for each block.
[0151] The prediction mode information Pinfo is provided to the encoding unit 115 and the prediction unit 122 .
[0152] The transformation information Tinfo is provided to the encoding unit 115 , the orthogonal transformation unit 113 , the quantization unit 114 , the inverse quantization unit 117 , and the inverse orthogonal transformation unit 118 .
[0153] The filter information Finfo is provided to the loop filtering unit 120 .
[0154] <Reorder Buffer>
[0155] Each field (input image) of the moving image data is input to the image encoding device 100 in a reproduction order (display order). The reordering buffer 111 acquires and holds (stores) each input image in its reproduction order (display order). The reordering buffer 111 reorders the input image in a coding order (decoding order) or divides the input image into blocks of processing units based on the control of the control unit 101. The reordering buffer 111 supplies the processed input image to the calculation unit 112. In addition, the reordering buffer 111 also supplies the input image (original image) to the prediction unit 122 and the loop filtering unit 120.
[0156] <Computational Unit>
[0157] The calculation unit 112 receives as input an image I corresponding to a block of the processing unit and a predicted image P provided from the prediction unit 122, subtracts the predicted image P from the image I to derive a predicted residual D as shown in the following expression, and provides the predicted residual D to the orthogonal transformation unit 113.
[0158] D=IP.
[0159] <Orthogonal Transformation Unit>
[0160] The orthogonal transform unit 113 uses the prediction residual D provided from the calculation unit 112 and the transform information Tinfo provided from the control unit 101 as input, and performs an orthogonal transform on the prediction residual D based on the transform information Tinfo to derive a transform coefficient Coeff. Note that the orthogonal transform unit 113 may perform an adaptive orthogonal transform for adaptively selecting the type (transform coefficient) of the orthogonal transform. The orthogonal transform unit 113 provides the obtained transform coefficient Coeff to the quantization unit 114.
[0161] <Quantization Unit>
[0162] The quantization unit 114 uses the transform coefficient Coeff provided from the orthogonal transform unit 113 and the transform information Tinfo provided from the control unit 101 as inputs, and scales (quantizes) the transform coefficient Coeff based on the transform information Tinfo. Note that the rate of this quantization is controlled by the rate control unit 123. The quantization unit 114 provides the quantized transform coefficient obtained by quantization (i.e., the quantized transform coefficient level "level") to the encoding unit 115 and the inverse quantization unit 117.
[0163] <Coding unit>
[0164] The encoding unit 115 uses as input the quantized transform coefficient level “level” supplied from the quantization unit 114, various encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, etc.) supplied from the control unit 101, information about the filter (e.g., filter coefficient) supplied from the loop filtering unit 120, and information about the optimal prediction mode supplied from the prediction unit 122. The encoding unit 115 performs variable length encoding (e.g., arithmetic encoding) on the quantized transform coefficient level “level” to generate a bit string (encoded data).
[0165] Furthermore, the encoding unit 115 derives residual information Rinfo from the quantized transform coefficient level "level", and encodes the residual information Rinfo to generate a bit string.
[0166] In addition, the encoding unit 115 includes information about the filter provided from the loop filtering unit 120 into the filtering information Finfo, and includes information about the optimal prediction mode provided from the prediction unit 122 into the prediction mode information Pinfo. Then, the encoding unit 115 encodes the above-mentioned various encoding parameters (header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, filtering information Finfo, etc.) to generate a bit string.
[0167] Furthermore, the encoding unit 115 multiplexes the bit strings of the various types of information generated as described above to generate encoded data. The encoding unit 115 supplies the encoded data to the accumulation buffer 116 .
[0168] <Accumulation Buffer>
[0169] The accumulation buffer 116 temporarily stores the encoded data obtained by the encoding unit 115. The accumulation buffer 116 outputs the stored encoded data as a bit stream or the like to the outside of the image encoding device 100 at a predetermined timing. For example, the encoded data is transmitted to the decoding side via an arbitrary recording medium, an arbitrary transmission medium, an arbitrary information processing device, etc. That is, the accumulation buffer 116 is also a transmission unit that transmits the encoded data (bit stream).
[0170] <Inverse Quantization Unit>
[0171] The inverse quantization unit 117 performs processing related to inverse quantization. For example, the inverse quantization unit 117 uses the quantized transform coefficient level "level" provided from the quantization unit 114 and the transform information Tinfo provided from the control unit 101 as inputs, and scales (inversely quantizes) the value of the quantized transform coefficient level "level" based on the transform information Tinfo. Note that inverse quantization is an inverse process of quantization performed in the quantization unit 114.
[0172] The inverse quantization unit 117 supplies the transform coefficient Coeff_IQ obtained by the inverse quantization to the inverse orthogonal transform unit 118. Note that since the inverse orthogonal transform unit 118 is similar to the inverse orthogonal transform unit on the decoding side (to be described below), the description given for the decoding side (to be described below) can be applied to the inverse quantization unit 117.
[0173] <Inverse Orthogonal Transformation Unit>
[0174] The inverse orthogonal transform unit 118 performs processing related to inverse orthogonal transform. For example, the inverse orthogonal transform unit 118 uses the transform coefficient Coeff_IQ provided from the inverse quantization unit 117 and the transform information Tinfo provided from the control unit 101 as input, and performs inverse orthogonal transform on the transform coefficient Coeff_IQ based on the transform information Tinfo to derive the prediction residual D'. Note that the inverse orthogonal transform is an inverse process of the orthogonal transform performed in the orthogonal transform unit 113. That is, the inverse orthogonal transform unit 118 can perform an adaptive inverse orthogonal transform for adaptively selecting the type (transform coefficient) of the inverse orthogonal transform.
[0175] The inverse orthogonal transform unit 118 supplies the prediction residual D' obtained by the inverse orthogonal transform to the calculation unit 119. Note that since the inverse orthogonal transform unit 118 is similar to the inverse orthogonal transform unit on the decoding side (to be described below), the description given for the decoding side (to be described below) can be applied to the inverse orthogonal transform unit 118.
[0176] <Computational Unit>
[0177] The calculation unit 119 uses the prediction residual D' provided from the inverse orthogonal transform unit 118 and the prediction image P provided from the prediction unit 122 as input. The calculation unit 119 adds the prediction residual D' and the prediction image P corresponding to the prediction residual D' to derive a local decoded image Rlocal. The calculation unit 119 provides the derived local decoded image Rlocal to the loop filtering unit 120 and the frame memory 121.
[0178] <Loop filter unit>
[0179] The loop filtering unit 120 performs processing related to the loop filtering process. For example, the loop filtering unit 120 uses the local decoded image Rlocal provided from the calculation unit 119, the filter information Finfo provided from the control unit 101, and the input image (original image) provided from the reordering buffer 111 as input. Note that the information input to the loop filtering unit 120 is arbitrary and may include information other than the information mentioned above. For example, information such as prediction mode, motion information, code amount target value, quantization parameter QP, picture type, block (CU, CTU, etc.) may be input to the loop filtering unit 120 as needed.
[0180] The loop filtering unit 120 appropriately performs a filtering process on the local decoded image Rlocal based on the filtering information Finfo. The loop filtering unit 120 also uses an input image (original image) and other input information for filtering processing as necessary.
[0181] For example, the loop filtering unit 120 applies the four loop filters in the order of a bilateral filter, a deblocking filter (DBF), an adaptive offset filter (sampling adaptive offset (SAO)), and an adaptive loop filter (adaptive loop filter (ALF)). Note that which filter is applied and in what order the filters are applied are arbitrary and can be appropriately selected.
[0182] Of course, the filtering process performed by the loop filtering unit 120 is arbitrary and is not limited to the above example. For example, the loop filtering unit 120 may apply a Wiener filter or the like.
[0183] The loop filtering unit 120 supplies the filtered local decoded image Rlocal to the frame memory 121. Note that the loop filtering unit 120 supplies the information about the filter to the encoding unit 115 in the case where information about the filter (for example, filter coefficient) is transmitted to the decoding side.
[0184] <Frame Memory>
[0185] The frame memory 121 performs processing related to storage of data related to an image. For example, the frame memory 121 uses the local decoded image Rlocal supplied from the calculation unit 119 and the filtered local decoded image Rlocal supplied from the loop filtering unit 120 as inputs, and holds (stores) the inputs. In addition, the frame memory 121 reconstructs and holds the decoded image R for each picture unit using the local decoded image Rlocal (storing the decoded image R in a buffer in the frame memory 121). The frame memory 121 supplies the decoded image R (or part thereof) to the prediction unit 122 in response to a request from the prediction unit 122.
[0186] <Prediction Unit>
[0187] The prediction unit 122 performs processing on the generation of a predicted image. For example, the prediction unit 122 uses the prediction mode information Pinfo provided from the control unit 101, the input image (original image) provided from the reordering buffer 111, and the decoded image R (or part thereof) read from the frame memory 121 as input. The prediction unit 122 performs prediction processing such as inter-frame prediction, intra-frame prediction, etc. using the prediction mode information Pinfo and the input image (original image), performs prediction using the decoded image R as a reference image, performs motion compensation processing based on the prediction result, and generates a predicted image P. The prediction unit 122 provides the generated predicted image P to the calculation unit 112 and the calculation unit 119. In addition, the prediction unit 122 provides the prediction mode selected by the above-mentioned processing (that is, information on the optimal prediction mode) to the encoding unit 115 as needed.
[0188] <Rate Control Unit>
[0189] The rate control unit 123 performs processing related to rate control. For example, the rate control unit 123 controls the rate of the quantization operation of the quantization unit 114 based on the code amount of the encoded data accumulated in the accumulation buffer 116 so that overflow or underflow does not occur.
[0190] Note that these processing units (control unit 101 and reorder buffer 111 to rate control unit 123) have arbitrary configurations. For example, each processing unit can be configured by a logic circuit that implements the above-mentioned processing. In addition, each processing unit may include, for example, a central processing unit (CPU), a read-only memory (ROM), a random access memory (RAM), etc., and implements the above-mentioned processing by executing a program using the above resources. Of course, each processing unit can have these two configurations, and implement part of the above-mentioned processing by a logic circuit, and implement another part of the processing by executing a program. The configurations of the processing units can be independent of each other. For example, some of the processing units can implement part of the above-mentioned processing by a logic circuit, other of the processing units can implement the above-mentioned processing by executing a program, and still other of the processing units can implement the above-mentioned processing by both the logic circuit and the execution of the program.
[0191] <Coding unit>
[0192] Figure 7 It is shown Figure 6 1 is a block diagram of an example of a main configuration of the encoding unit 115 in FIG. Figure 7 As shown, the encoding unit 115 includes a binarization unit 131, a selection unit 132, a context setting unit 133, a context encoding unit 134 and a bypass encoding unit 135.
[0193] Note that although the encoding of the adaptive orthogonal transform identifier is described here, the encoding unit 115 also encodes other encoding parameters, residual information Rinfo, etc. as described above. The encoding unit 115 encodes the adaptive orthogonal transform identifier by applying method 1 described in <2-1. Encoding of adaptive orthogonal transform identifier>.
[0194] The binarization unit 131 performs binarization processing on the adaptive orthogonal transform identifier. For example, the binarization unit 131 obtains the adaptive orthogonal transform identifier mts_idx provided from the control unit 101. In addition, the binarization unit 131 binarizes the adaptive orthogonal transform identifier using a truncated unary code (or a truncated Rice code) to generate a bin sequence. In addition, the binarization unit 131 provides the generated bin sequence to the selection unit 132.
[0195] The selection unit 132 performs processing regarding selection of a supply destination for each bin in the bin sequence of the adaptive orthogonal transform identifier. For example, the selection unit 132 acquires the bin sequence of the adaptive orthogonal transform identifier supplied from the binarization unit 131.
[0196] Furthermore, the selection unit 132 selects whether to set the supply destination to the context setting unit 133 or the bypass encoding unit 135 for each bin in the bin sequence of the acquired adaptive orthogonal transform identifier. The selection unit 132 performs the selection according to the method 1 described above in <2-1. Encoding of adaptive orthogonal transform identifier>. For example, the selection unit 132 may select according to the method 1-1 (i.e., Figure 3 In addition, the selection unit 132 may perform selection according to method 1-2 (ie, Figure 3 In addition, the selection unit 132 may perform selection according to methods 1-3 (ie, Figure 4 In addition, the selection unit 132 may perform selection according to methods 1-4 (i.e., Figure 4 The table shown in B) performs the selection.
[0197] In the case of allocating a context variable (index ctxInc) and performing context encoding, the selection unit 132 supplies the bin to the context setting unit 133. In the case of bypass encoding, the selection unit 132 supplies the bin to the bypass encoding unit 135.
[0198] The context setting unit 133 performs processing on context setting. For example, the context setting unit 133 obtains the bin provided from the selection unit 132. The context setting unit 133 assigns the context variable (index ctxInc) to the bin. The context setting unit 133 performs assignment according to method 1 described above in <2-1. Encoding of adaptive orthogonal transform identifier>. For example, the context setting unit 133 may perform assignment according to method 1-1 (i.e., Figure 3 In addition, the context setting unit 133 may perform allocation according to method 1-2 (ie, Figure 3 In addition, the context setting unit 133 may perform allocation according to methods 1-3 (ie, Figure 4 In addition, the context setting unit 133 may perform allocation according to methods 1-4 (ie, Figure 4 The allocation is performed according to the table shown in B).
[0199] In addition, the context setting unit 133 can acquire the encoding result from the context encoding unit 134. The context setting unit 133 can appropriately update the context variable (index ctxInc) using the encoding result. The context setting unit 133 supplies the context variable (index ctxInc) derived in this way to the context encoding unit 134.
[0200] The context encoding unit 134 performs processing related to arithmetic coding. For example, the context encoding unit 134 obtains the context variable (index ctxInc) provided from the context setting unit 133. In addition, the context encoding unit 134 performs arithmetic coding using the context variable (index ctxInc). That is, context coding is performed. In addition, the context encoding unit 134 provides the encoding result as encoded data to the accumulation buffer 116.
[0201] The bypass encoding unit 135 performs processing related to bypass encoding. For example, the bypass encoding unit 135 acquires the bin supplied from the selection unit 132. The bypass encoding unit 135 performs bypass encoding (arithmetic encoding) on the bin. The bypass encoding unit 135 supplies the encoding result to the accumulation buffer 116 as encoded data.
[0202] Since each of the processing units (binarization unit 131 to bypass encoding unit 135) performs the processing as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 1 (for example, any one of methods 1-1 to 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0203] <Flow of Image Coding Process>
[0204] Next, we will refer to Figure 8 The flowchart describes an example of the flow of an image encoding process performed by the image encoding device 100 having the above configuration.
[0205] When the image encoding process starts, in step S101 , the reordering buffer 111 is controlled by the control unit 101 , and frames of input moving image data are reordered from a display order to an encoding order.
[0206] In step S102 , the control unit 101 sets a processing unit (performs block division) on the input image held by the rearrangement buffer 111 .
[0207] In step S103 , the control unit 101 determines (sets) encoding parameters of the input image held by the reordering buffer 111 .
[0208] In step S104, the prediction unit 122 performs prediction processing, and generates a predicted image of the optimal prediction mode, etc. For example, in the prediction processing, the prediction unit 122 performs intra prediction to generate a predicted image of the optimal intra prediction mode, performs inter prediction to generate a predicted image of the optimal inter prediction mode, and selects the optimal prediction mode from the predicted images based on a cost function value, etc.
[0209] In step S105, the calculation unit 112 calculates the difference between the input image and the predicted image of the optimal mode selected by the prediction process in step S104. That is, the calculation unit 112 generates a prediction residual D between the input image and the predicted image. The prediction residual D obtained in this way reduces the amount of data compared to the original image data. Therefore, compared with the case where the image is encoded as it is, the amount of data can be compressed.
[0210] In step S106 , the orthogonal transform unit 113 performs an orthogonal transform process on the prediction residual D generated by the process in step S105 to derive a transform coefficient Coeff.
[0211] In step S107 , the quantization unit 114 quantizes the transform coefficient Coeff obtained by the process in step S106 by using the quantization parameter calculated by the control unit 101 or the like to derive a quantized transform coefficient level “level”.
[0212] In step S108 , the inverse quantization unit 117 inversely quantizes the quantized transform coefficient level “level” generated by the process in step S107 using characteristics corresponding to the characteristics of the quantization in step S107 to derive a transform coefficient Coeff_IQ.
[0213] In step S109, the inverse orthogonal transform unit 118 performs inverse orthogonal transform on the transform coefficient Coeff_IQ obtained by the process in step S108 by a method corresponding to the orthogonal transform process in step S106 to derive the prediction residual D'. Note that since the inverse orthogonal transform process is similar to the inverse orthogonal transform process (to be described below) performed on the decoding side, the description of the decoding side (to be given below) can be applied to the inverse orthogonal transform process in step S109.
[0214] In step S110 , the calculation unit 119 adds the predicted image obtained by the prediction process in step S104 and the prediction residual D′ derived by the process in step S109 to generate a local decoded image.
[0215] In step S111 , the loop filtering unit 120 performs a loop filtering process on the local decoded image derived by the process in step S110 .
[0216] In step S112 , the frame memory 121 stores the local decoded image derived by the process in step S110 and the local decoded image filtered in step S111 .
[0217] In step S113, the encoding unit 115 encodes the quantized transform coefficient level "level" obtained by the process in step S107. For example, the encoding unit 115 encodes the quantized transform coefficient level "level" as information about the image by arithmetic coding or the like to generate encoded data. In addition, at this time, the encoding unit 115 encodes various encoding parameters (header information Hinfo, prediction mode information Pinfo, and transform information Tinfo). In addition, the encoding unit 115 derives residual information RInfo from the quantized transform coefficient level "level" and encodes the residual information RInfo.
[0218] In step S114, the accumulation buffer 116 accumulates the encoded data thus obtained, and outputs the encoded data as a bit stream, for example, to the outside of the image encoding device 100. For example, the bit stream is sent to the decoding side via a transmission path or a recording medium. In addition, the rate control unit 123 performs rate control as needed.
[0219] When the processing in step S114 ends, the image encoding processing ends.
[0220] <Encoding Process Flow>
[0221] exist Figure 8 In the encoding process of step S113 in the above, the encoding unit 115 encodes the adaptive orthogonal transform identifier mts_idx. At this time, the encoding unit 115 encodes the adaptive orthogonal transform identifier by applying the method 1 described in <2-1. Encoding of the adaptive orthogonal transform identifier>. Fig. 9 The flowchart in describes an example of a process of encoding an adaptive orthogonal transform identifier mts_idx.
[0222] When the encoding process starts, in step S131, the binarization unit 131 of the encoding unit 115 binarizes the adaptive orthogonal transform identifier mts_idx by truncated unary code (or truncated Rice code) to generate a bin sequence.
[0223] In step S132, the selection unit 132 sets the first bin (binIdx=0) in the bin sequence as the bin to be processed. In this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (i.e., selects context encoding). For example, the selection unit 132 selects Figure 3 and Figure 4Any one of the tables shown in (i.e., by applying any one of Methods 1-1 to 1-4) selects context coding as the encoding method for that bin.
[0224] In step S133, the context setting unit 133 assigns a predetermined context variable ctx (index ctxInc) determined in advance to the bin. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0225] In step S134, the selection unit 132 sets the unprocessed bins in the second bin and the subsequent bins in the bin sequence as the bins to be processed. In step S135, the selection unit 132 determines whether to perform bypass encoding on the bin to be processed. For example, the selection unit 132 determines whether to perform bypass encoding on the bin to be processed. Figure 3 and Figure 4 Any of the tables shown in determines (by applying any of Methods 1-1 to 1-4) whether to bypass encode the bin to be processed.
[0226] In the case where it is determined to perform bypass coding, the process proceeds to step S136. That is, in this case, the selection unit 132 selects the bypass coding unit 135 as the bin supply destination. In step S136, the bypass coding unit 135 performs bypass coding (arithmetic coding) on the bin to be processed. When the process in step S136 is completed, the process proceeds to step S138.
[0227] In addition, in step S135, in the case where it is determined that bypass encoding is not performed (context encoding is performed), the processing proceeds to step S137. That is, in this case, the selection unit 132 selects the context setting unit 133 as the bin supply destination. In step S137, the context setting unit 133 assigns a predetermined context variable ctx (index ctxInc) determined in advance to the bin. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed. When the processing in step S137 is completed, the processing proceeds to step S138.
[0228] In step S138, the encoding unit 115 determines whether to terminate the encoding of the adaptive orthogonal transform identifier mts_idx. In the case where it is determined that the encoding is not terminated, the process returns to step S134, and the processes in step S134 and subsequent steps are repeated. In addition, in step S138, in the case where it is determined that the encoding is terminated, the encoding process is terminated.
[0229] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 1 (for example, any one of method 1-1 to method 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0230] <2-4. Decoding side>
[0231] <Image decoding device>
[0232] Next, the decoding side will be described. Fig.10 : is a block diagram showing an example of the configuration of an image decoding device as one mode of an image processing device to which the present technology is applied. Fig.10 The image decoding device 200 shown in the figure is a device for encoding the encoded data of the moving image. For example, the image decoding device 200 decodes the encoded data by the decoding method described in any one of the non-patent documents 1 to 10 to generate the moving image data. For example, the image decoding device 200 decodes the encoded data (bit stream) generated by the above-mentioned image encoding device 100 to generate the moving image data.
[0233] Notice, Fig.10 shows the main processing units, data flows, etc., and Fig.10 That is, in the image decoding device 200, there may be Fig.10 The processing units are shown as blocks, or there are no Fig.10 A process or data flow is shown as arrows or the like.
[0234] exist Fig.10 , the image decoding device 200 includes an accumulation buffer 211, a decoding unit 212, an inverse quantization unit 213, an inverse orthogonal transform unit 214, a calculation unit 215, a loop filtering unit 216, a reordering buffer 217, a frame memory 218, and a prediction unit 219. Note that the prediction unit 219 includes an intra-frame prediction unit and an inter-frame prediction unit (not shown).
[0235] <Accumulation Buffer>
[0236] The accumulation buffer 211 acquires a bit stream input to the image decoding device 200 and holds (stores) the bit stream. For example, the accumulation buffer 211 supplies the accumulated bit stream to the decoding unit 212 at a predetermined timing or when a predetermined condition is satisfied.
[0237] <Decoding Unit>
[0238] The decoding unit 212 performs processing related to image decoding. For example, the decoding unit 212 uses the bit stream supplied from the accumulation buffer 211 as input, and performs variable length decoding on the syntax value of each syntax element from the bit string according to the definition of the syntax table to derive parameters.
[0239] Parameters derived from syntax elements and syntax values of syntax elements include, for example, information such as header information Hinfo, prediction mode information Pinfo, transform information Tinfo, residual information Rinfo, and filter information Finfo. That is, the decoding unit 212 parses (analyzes and obtains) such information from the bitstream. This information will be described below.
[0240] <Header Information Hinfo>
[0241] The header information Hinfo includes, for example, header information such as video parameter set (VPS) / sequence parameter set (SPS) / picture parameter set (PPS) / slice header (SH). The header information Hinfo includes, for example, information defining the following items: image size (width PicWidth and height PicHeight), bit depth (luminance bitDepthY and color difference bitDepthC), color difference array type ChromaArrayType, CU size maximum value MaxCUSize / minimum value MinCUSize, maximum depth MaxQTDepth / minimum depth MinQTDepth of quadtree partitioning, maximum depth MaxBTDepth / minimum depth MinBTDepth of binary tree partitioning, maximum value MaxTSSize of transform skip block (also called maximum transform skip block size), on / off flag of each encoding tool (also called enable flag), etc.
[0242] For example, an example of the on / off flag of the coding tool included in the header information Hinfo includes an on / off flag related to the following transform and quantization processing. Note that the on / off flag of the coding tool can also be interpreted as a flag indicating whether the syntax related to the coding tool is present in the coded data. In addition, when the value of the on / off flag is 1 (true), the value indicates that the coding tool is available. When the value of the on / off flag is 0 (false), the value indicates that the coding tool is not available. Note that the interpretation of the flag value can be swapped.
[0243] The inter-component prediction enable flag (ccp_enabled_flag) is flag information indicating whether inter-component prediction (cross-component prediction (CCP), also referred to as CC prediction) is available. For example, when the flag information is "1" (true), the flag information indicates that inter-component prediction is available. When the flag information is "0" (false), the flag information indicates that inter-component prediction is not available.
[0244] Note that this CCP is also called inter-component linear prediction (CCLM or CCLMP).
[0245] <Prediction mode information Pinfo>
[0246] The prediction mode information Pinfo includes, for example, information such as size information PBSize (prediction block size) of a prediction block (PB) to be processed, intra prediction mode information IPinfo, and motion prediction information MVinfo.
[0247] The intra prediction mode information IPinfo includes, for example, prev_intra_luma_pred_flag, mpm_idx, and rem_intra_pred_mode in JCTVC-W1005, 7.3.8.5 coding unit syntax, luma intra prediction mode IntraPredModeY derived from the syntax, and the like.
[0248] In addition, the intra-frame prediction mode information IPinfo includes, for example, an inter-component prediction flag (ccp_flag(cclmp_flag)), a multi-class linear prediction mode flag (mclm_flag), a chroma sample location type identifier (chroma_sample_loc_type_idx), a chroma MPM identifier (chroma_mpm_idx), a luma intra-frame prediction mode (IntraPredModeC) derived from these syntaxes, and the like.
[0249] The inter-component prediction flag (ccp_flag (cclmp_flag)) is flag information indicating whether inter-component linear prediction is applied. For example, ccp_flag==1 indicates that inter-component prediction is applied, and ccp_flag==0 indicates that inter-component prediction is not applied.
[0250] The multi-class linear prediction mode flag (mclm_flag) is information about the linear prediction mode (linear prediction mode information). More specifically, the multi-class linear prediction mode flag (mclm_flag) is flag information indicating whether the multi-class linear prediction mode is set. For example, "0" indicates a one-class mode (single-class mode) (e.g., CCLMP), and "1" indicates a two-class mode (multi-class mode) (e.g., MCLMP).
[0251] The chroma sample position type identifier (chroma_sample_loc_type_idx) is an identifier for identifying the type of the pixel position of the chroma component (also referred to as the chroma sample position type). For example, in the case where the chroma array type (ChromaArrayType) as information on the color format indicates a 420 format, the chroma sample position type identifier is assigned as in the following expression.
[0252] chroma_sample_loc_type_idx==0: Type 2
[0253] chroma_sample_loc_type_idx == 1: Type 3
[0254] chroma_sample_loc_type_idx == 2: Type 0
[0255] chroma_sample_loc_type_idx == 3: Type 1
[0256] Note that the color difference sample position type identifier (chroma_sample_loc_type_idx) is transmitted as the information (chroma_sample_loc_info()) on the pixel position of the color difference component (chroma_sample_loc_info()) by being stored in the information.
[0257] The chroma MPM identifier (chroma_mpm_idx) is an identifier indicating which prediction mode candidate in the chroma intra prediction mode candidate list (intraPredModeCandListC) is to be designated as the chroma intra prediction mode.
[0258] The motion prediction information MVinfo includes, for example, information such as merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X={0,1}, mvd, etc. (for example, see JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax).
[0259] Of course, the information included in the prediction mode information Pinfo is arbitrary and may include information other than the above information.
[0260] <Transformation information Tinfo>
[0261] The transformation information Tinfo includes, for example, the following information. Of course, the information included in the transformation information Tinfo is arbitrary, and may include information other than the above information:
[0262] The width TBWSize and height TBHSize of the transform block to be processed (or the logarithmic values of TBWSize and TBHSize with base 2, log2TBWSize and log2TBHSize);
[0263] Transform skip flag (transform_ts_flag): A flag indicating whether to skip (inverse) primary transform and (inverse) secondary transform.
[0264] Scan identifier (scanIdx);
[0265] Secondary transformation identifier (st_idx);
[0266] Adaptive orthogonal transform identifier (mts_idx);
[0267] quantization parameter (qp); and
[0268] Quantization matrix (scaling_matrix (eg, JCTVC-W1005, 7.3.4 Scaling List Data Syntax)).
[0269] <Residual information Rinfo>
[0270] The residual information Rinfo (for example, see 7.3.8.11 Residual Coding Syntax of JCTVC-W1005) includes, for example, the following syntax:
[0271] cbf (coded_block_flag): residual data presence / absence flag;
[0272] last_sig_coeff_x_pos: the last non-zero transform coefficient X coordinate;
[0273] last_sig_coeff_y_pos: the last non-zero transform coefficient Y coordinate;
[0274] coded_sub_block_flag: flag indicating the presence / absence of sub-block non-zero transform coefficients;
[0275] sig_coeff_flag: non-zero transform coefficient presence / absence flag;
[0276] gr1_flag: A flag indicating whether the level of the non-zero transform coefficients is greater than 1 (also called the GR1 flag);
[0277] gr2_flag: A flag indicating whether the level of non-zero transform coefficients is greater than 2 (also known as the GR2 flag);
[0278] sign_flag: a code indicating the sign of a non-zero transform coefficient (also called a sign code);
[0279] coeff_abs_level_remaining: residual level of non-zero transform coefficients (also known as non-zero transform coefficient residual level);
[0280] etc.
[0281] Of course, the information included in the residual information Rinfo is arbitrary and may include information other than the above information.
[0282] <Filter information Finfo>
[0283] The filtering information Finfo includes, for example, control information about the following filtering processes:
[0284] Control information about the deblocking filter (DBF)
[0285] Control information about pixel adaptive offset (SAO)
[0286] Control information about the adaptive loop filter (ALF)
[0287] Control information about other linear / non-linear filters
[0288] More specifically, the filter information Finfo includes, for example, a picture to which each filter is applied, information for specifying an area in the picture, filter on / off control information for each CU, filter on / off control information for slice and tile boundaries, etc. Of course, the information included in the filter information Finfo is arbitrary and may include information other than the above information.
[0289] Return to the description of the decoding unit 212 . The decoding unit 212 refers to the residual information Rinfo and derives the quantized transform coefficient level “level” at each coefficient position in each transform block. The decoding unit 212 supplies the quantized transform coefficient level “level” to the inverse quantization unit 213 .
[0290] In addition, the decoding unit 212 provides the parsed header information Hinfo, prediction mode information Pinfo, quantized transform coefficient level "level", transform information Tinfo, and filter information Finfo to each block. A detailed description is given below.
[0291] The header information Hinfo is supplied to the inverse quantization unit 213 , the inverse orthogonal transform unit 214 , the prediction unit 219 , and the loop filtering unit 216 .
[0292] The prediction mode information Pinfo is supplied to the inverse quantization unit 213 and the prediction unit 219 .
[0293] The transform information Tinfo is supplied to the inverse quantization unit 213 and the inverse orthogonal transform unit 214 .
[0294] The filter information Finfo is provided to the loop filtering unit 216 .
[0295] Of course, the above example is an example, and the present embodiment is not limited to this example. For example, each encoding parameter can be provided to any processing unit. In addition, other information can be provided to any processing unit.
[0296] <Inverse Quantization Unit>
[0297] The inverse quantization unit 213 has at least a configuration required to perform processing related to inverse quantization. For example, the inverse quantization unit 213 uses the transform information Tinfo and the quantized transform coefficient level "level" provided from the decoding unit 212 as inputs, and scales (inverse quantizes) the value of the quantized transform coefficient level "level" based on the transform information Tinfo to derive the transform coefficient Coeff_IQ after inverse quantization.
[0298] Note that this inverse quantization is performed as an inverse process of the quantization performed by the quantization unit 114 of the image encoding device 100. In addition, the inverse quantization is a process similar to the inverse quantization performed by the inverse quantization unit 117 of the image encoding device 100. That is, the inverse quantization unit 117 of the image encoding device 100 performs a process (inverse quantization) similar to that of the inverse quantization unit 213.
[0299] The inverse quantization unit 213 supplies the derived transform coefficient Coeff_IQ to the inverse orthogonal transform unit 214 .
[0300] <Inverse Orthogonal Transformation Unit>
[0301] The inverse orthogonal transform unit 214 performs inverse orthogonal transform processing. For example, the inverse orthogonal transform unit 214 uses the transform coefficient Coeff_IQ provided from the inverse quantization unit 213 and the transform information Tinfo provided from the decoding unit 212 as input, and performs inverse orthogonal transform processing on the transform coefficient Coeff_IQ based on the transform information Tinfo to derive the prediction residual D'.
[0302] Note that this inverse orthogonal transform is performed as an inverse process of the orthogonal transform performed by the orthogonal transform unit 113 of the image encoding device 100. In addition, the inverse orthogonal transform is a process similar to the inverse orthogonal transform performed by the inverse orthogonal transform unit 118 of the image encoding device 100. That is, the inverse orthogonal transform unit 118 of the image encoding device 100 performs a process similar to that of the inverse orthogonal transform unit 214 (inverse orthogonal transform).
[0303] The inverse orthogonal transform unit 214 supplies the derived prediction residual D′ to the calculation unit 215 .
[0304] <Computational Unit>
[0305] The calculation unit 215 performs processing related to the addition of information about the image. For example, the calculation unit 215 uses the prediction residual D' provided from the inverse orthogonal transform unit 214 and the predicted image P provided from the prediction unit 219 as input. The calculation unit 215 adds the prediction residual D' and the predicted image P (prediction signal) corresponding to the prediction residual D' to derive a local decoded image Rlocal as shown in the following expression.
[0306] Rlocal=D'+P
[0307] The calculation unit 215 provides the derived local decoded image Rlocal to the loop filtering unit 216 and the frame memory 218 .
[0308] <Loop filter unit>
[0309] The loop filtering unit 216 performs processing related to the loop filtering process. For example, the loop filtering unit 216 uses the local decoded image Rlocal provided from the calculation unit 215 and the filter information Finfo provided from the decoding unit 212 as input. Note that the information input to the loop filtering unit 216 is arbitrary and may include information other than the above-mentioned information.
[0310] The loop filtering unit 216 appropriately performs filtering processing on the local decoded image Rlocal based on the filter information Finfo.
[0311] For example, the loop filtering unit 216 applies the four loop filters in the order of a bilateral filter, a deblocking filter (DBF), an adaptive offset filter (sample adaptive offset (SAO)), and an adaptive loop filter (adaptive loop filter (ALF)). Note that which filter is applied and in what order the filters are applied are arbitrary and can be appropriately selected.
[0312] The loop filtering unit 216 performs a filtering process corresponding to the filtering process performed on the encoding side (e.g., by the loop filtering unit 120 of the image encoding device 100). Of course, the filtering process performed by the loop filtering unit 216 is arbitrary and is not limited to the above example. For example, the loop filtering unit 216 can apply a Wiener filter or the like.
[0313] The loop filtering unit 216 supplies the filtered local decoded image Rlocal to the reordering buffer 217 and the frame memory 218 .
[0314] <Reorder Buffer>
[0315] The reordering buffer 217 uses the local decoded image Rlocal supplied from the loop filtering unit 216 as an input, and holds (stores) the local decoded image Rlocal. The reordering buffer 217 reconstructs the decoded image R for each picture unit using the local decoded image Rlocal, and holds (stores) the decoded image R (in the buffer). The reordering buffer 217 reorders the obtained decoded image R from the decoding order to the reproduction order. The reordering buffer 217 outputs the reordered decoded image R group as motion image data to the outside of the image decoding device 200.
[0316] <Frame Memory>
[0317] The frame memory 218 performs processing related to the storage of data related to the image. For example, the frame memory 218 uses the local decoded image Rlocal provided from the calculation unit 215 as input, reconstructs the decoded image R for each picture unit, and stores the decoded image R in a buffer in the frame memory 218.
[0318] Furthermore, the frame memory 218 reconstructs a decoded image R for each picture unit using the loop-filtered local decoded image Rlocal supplied from the loop filtering unit 216 as an input, and stores the decoded image R in a buffer in the frame memory 218. The frame memory 218 appropriately supplies the stored decoded image R (or a portion thereof) to the prediction unit 219 as a reference image.
[0319] Note that the frame memory 218 may store header information Hinfo, prediction mode information Pinfo, transform information Tinfo, filter information Finfo, and the like related to generation of a decoded image.
[0320] <Prediction Unit>
[0321] The prediction unit 219 performs processing related to the generation of the predicted image. For example, the prediction unit 219 uses the prediction mode information Pinfo provided from the decoding unit 212 as an input, and performs prediction by the prediction method specified by the prediction mode information Pinfo to derive the predicted image P. When deriving, the prediction unit 219 uses the decoded image R (or a part thereof) before or after filtering stored in the frame memory 218 as a reference image, and the decoded image R is specified by the prediction mode information Pinfo. The prediction unit 219 provides the derived predicted image P to the calculation unit 215.
[0322] Note that these processing units (accumulation buffer 211 to prediction unit 219) have an arbitrary configuration. For example, each processing unit can be configured by a logic circuit that implements the above-mentioned processing. In addition, each processing unit may include, for example, a CPU, ROM, RAM, etc., and implement the above-mentioned processing by executing a program using the above resources. Of course, each processing unit can have these two configurations, and implement a part of the above-mentioned processing by a logic circuit, and implement another part of the processing by executing a program. The configurations of the processing units can be independent of each other. For example, some of the processing units can implement a part of the above-mentioned processing by a logic circuit, other of the processing units can implement the above-mentioned processing by executing a program, and still other of the processing units can implement the above-mentioned processing by both the logic circuit and the execution of the program.
[0323] <Decoding Unit>
[0324] Fig.11 It is shown Fig.10 212 in the decoding unit 212 of the main configuration example of the block diagram. Fig.11 As shown, the decoding unit 212 includes a selection unit 231, a context setting unit 232, a context decoding unit 233, a bypass decoding unit 234 and a debinarization unit 235.
[0325] Note that although the decoding of the coded data of the adaptive orthogonal transform identifier is described here, the decoding unit 212 also decodes the coded data of other coding parameters, residual information Rinfo, etc. as described above. The decoding unit 212 decodes the coded data of the adaptive orthogonal transform identifier by applying method 1 described in <2-2. Decoding of adaptive orthogonal transform identifier>.
[0326] The selection unit 231 performs processing for selecting a supply destination of the encoded data of each bin in the bin sequence of the bin-valued adaptive orthogonal transform identifier. For example, the selection unit 231 acquires the encoded data of the bin sequence of the adaptive orthogonal transform identifier supplied from the accumulation buffer 211.
[0327] In addition, the selection unit 231 selects whether to set the supply destination to the context setting unit 232 or the bypass decoding unit 234 for the encoded data of each bin in the bin sequence of the acquired adaptive orthogonal transform identifier. The selection unit 231 performs the selection according to the method 1 described above in <2-2. Decoding of adaptive orthogonal transform identifier>. For example, the selection unit 231 may select according to the method 1-1 (i.e., Figure 3 In addition, the selection unit 231 can select according to method 1-2 (ie, Figure 3 In addition, the selection unit 231 can select according to methods 1-3 (ie, Figure 4 In addition, the selection unit 231 can select according to methods 1-4 (ie, Figure 4 Select from the table shown in B).
[0328] In the case of allocating a context variable (index ctxInc) and performing context decoding, the selection unit 231 supplies the encoded data of the bin to the context setting unit 232. In the case of bypass decoding, the selection unit 231 supplies the bin to the bypass decoding unit 234.
[0329] The context setting unit 232 performs processing related to context setting. For example, the context setting unit 232 acquires the encoded data of the bin provided from the selection unit 231. The context setting unit 232 allocates the context variable (index ctxInc) to the bin. The context setting unit 232 performs allocation according to method 1 described above in <2-2. Decoding of adaptive orthogonal transform identifier>. For example, the context setting unit 232 may perform allocation according to method 1-1 (i.e., Figure 3 In addition, the context setting unit 232 may perform allocation according to method 1-2 (ie, Figure 3 In addition, the context setting unit 232 may perform allocation according to methods 1-3 (ie, Figure 4 In addition, the context setting unit 232 may perform allocation according to methods 1-4 (ie, Figure 4 The allocation is performed according to the table shown in B).
[0330] In addition, the context setting unit 232 can acquire the decoding result from the context decoding unit 233. The context setting unit 232 can appropriately update the context variable (index ctxInc) using the decoding result. The context setting unit 232 provides the context variable (index ctxInc) derived in this way to the context decoding unit 233.
[0331] The context decoding unit 233 performs processing related to arithmetic decoding. For example, the context decoding unit 233 obtains the context variable (index ctxInc) provided from the context setting unit 232. In addition, the context decoding unit 233 performs arithmetic decoding using the context variable (index ctxInc). That is, context decoding is performed. In addition, the context decoding unit 233 provides the bin to be processed in the bin sequence of the adaptive orthogonal transform identifier, which is the decoding result, to the debinarization unit 235.
[0332] The bypass decoding unit 234 performs processing related to bypass decoding. For example, the bypass decoding unit 234 obtains the encoded data of the bin provided from the selection unit 231. The bypass decoding unit 234 performs bypass decoding (arithmetic decoding) on the encoded data of the bin. The bypass decoding unit 234 provides the decoding result, that is, the bin to be processed in the bin sequence of the adaptive orthogonal transform identifier, to the debinarization unit 235.
[0333] The debinarization unit 235 performs a process of debinarization (also referred to as multi-binarization) of the bin sequence of the adaptive orthogonal transform identifier. For example, the debinarization unit 235 obtains the bin sequence of the adaptive orthogonal transform identifier provided from the context decoding unit 233 or the bypass decoding unit 234. The debinarization unit 235 debinarizes the obtained bin sequence to derive the adaptive orthogonal transform identifier mts_idx. This debinarization is the inverse process of the binarization performed by the binarization unit 131. That is, the debinarization unit 235 performs debinarization using a truncated unary code (or a truncated Rice code). The debinarization unit 235 provides the derived adaptive orthogonal transform identifier mts_idx as Tinfo to the inverse orthogonal transform unit 214. The inverse orthogonal transform unit 214 appropriately performs adaptive orthogonal transform based on the adaptive orthogonal transform identifier mts_idx.
[0334] Since each of the processing units (selection unit 231 to inverse binarization unit 235) performs the processing as described above, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 1 (for example, any one of methods 1-1 to 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0335] <Flow of Image Decoding Process>
[0336] Next, the flow of each process performed by the image decoding device 200 having the above configuration will be described. Fig.12 The flowchart in describes an example of the flow of image decoding processing.
[0337] When the image decoding process starts, in step S201 , the accumulation buffer 211 acquires and holds (accumulates) encoded data (bit stream) supplied from the outside of the image decoding device 200 .
[0338] In step S202, the decoding unit 212 decodes the encoded data (bit stream) to obtain the quantized transform coefficient level “level.” Furthermore, the decoding unit 212 parses (analyzes and obtains) various encoding parameters from the encoded data (bit stream) through the decoding.
[0339] In step S203 , the inverse quantization unit 213 performs inverse quantization, which is an inverse process of quantization performed on the encoding side, on the quantized transform coefficient level “level” obtained by the process in step S202 to obtain a transform coefficient Coeff_IQ.
[0340] In step S204 , the inverse orthogonal transform unit 214 performs an inverse orthogonal transform process, which is an inverse process of the orthogonal transform process performed on the encoding side, on the transform coefficient Coeff_IQ obtained in step S203 to obtain a prediction residual D′.
[0341] In step S205 , the prediction unit 219 performs prediction processing by a prediction method specified on the encoding side based on the information parsed in step S202 , and generates a predicted image P by referring to a reference image stored in the frame memory 218 , for example.
[0342] In step S206 , the calculation unit 215 adds the prediction residual D′ obtained in step S204 to the predicted image P obtained in step S205 to derive a local decoded image Rlocal.
[0343] In step S207 , the loop filtering unit 216 performs a loop filtering process on the local decoded image Rlocal obtained by the process in step S206 .
[0344] In step S208, the rearrangement buffer 217 uses the filtered local decoded image Rlocal obtained by the process in step S207 to derive a decoded image R, and rearranges the decoded image R group from the decoding order to the reproduction order. The decoded image R group rearranged in the reproduction order is output as a moving image to the outside of the image decoding device 200.
[0345] Furthermore, in step S209 , the frame memory 218 stores at least one of the local decoded image Rlocal obtained by the process in step S206 and the local decoded image Rlocal after the filtering process obtained by the process in step S207 .
[0346] When the processing in step S209 ends, the image decoding processing ends.
[0347] <Decoding Process Flow>
[0348] exist Fig.12 In the decoding process in step S202 of , the decoding unit 212 decodes the coded data of the adaptive orthogonal transform identifier mts_idx. At this time, the decoding unit 212 decodes the coded data of the adaptive orthogonal transform identifier by applying the method 1 described in <2-2. Decoding of the adaptive orthogonal transform identifier>. Fig.13 The flowchart in describes an example of the flow of decoding of encoded data of the adaptive orthogonal transform identifier mts_idx.
[0349] When the decoding process starts, in step S231, the selection unit 231 of the decoding unit 212 sets the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier mts_idx as the bin to be processed. In this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects context decoding). For example, the selection unit 231 selects Figure 3 and Figure 4 Any one of the tables shown in (ie, by applying any one of Methods 1-1 to 1-4) selects context decoding as the decoding method for that bin.
[0350] In step S232, the context setting unit 232 assigns a predetermined context variable ctx (index ctxInc) determined in advance to the bin. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.
[0351] In step S233, the selection unit 231 sets the unprocessed bins in the second bin and the subsequent bins in the bin sequence as the bins to be processed. In step S234, the selection unit 231 determines whether to perform bypass decoding on the bin to be processed. For example, the selection unit 231 determines whether to perform bypass decoding on the bin to be processed. Figure 3 and Figure 4 Any of the tables shown in (ie, by applying any of methods 1-1 to 1-4) determines whether to bypass decode the encoded data of the bin to be processed.
[0352] In the case where it is determined that bypass decoding is performed, the process proceeds to step S235. That is, in this case, the selection unit 231 selects the bypass decoding unit 234 as the supply destination of the encoded data of the bin. In step S235, the bypass decoding unit 234 bypass-decodes the encoded data of the bin to be processed. When the process in step S235 is completed, the process proceeds to step S237.
[0353] In addition, in step S234, in the case where it is determined that bypass decoding is not performed (context decoding is performed), the processing proceeds to step S236. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the encoded data of the bin. In step S236, the context setting unit 232 assigns a predetermined context variable ctx (index ctxInc) determined in advance to the bin. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed. When the processing in step S236 is completed, the processing proceeds to step S237.
[0354] In step S237, the debinarization unit 235 debinarizes the bin sequence using a truncated unary code (or a truncated Rice code) to derive an adaptive orthogonal transform identifier mts_idx.
[0355] In step S238, the decoding unit 212 determines whether to terminate the decoding of the adaptive orthogonal transform identifier mts_idx. In the case where it is determined that the decoding is not terminated, the process returns to step S233, and the processes in step S233 and subsequent steps are repeated. In addition, in step S238, in the case where it is determined that the decoding is terminated, the decoding process is terminated.
[0356] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 1 (for example, any one of methods 1-1 to 1-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0357] <3. Second Embodiment>
[0358] <3-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0359] In this embodiment, if Figure 1 Assigning a context variable to each bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (method 0) as has been performed in the table shown in B is performed as follows.
[0360] That is, a context variable based on a parameter about a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0361] More specifically, the parameter regarding the block size is the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied. That is, a context variable based on the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied is assigned to the first bin in the bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of the adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin (method 2).
[0362] For example, Fig.14 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value of the transform block size in the horizontal direction (log2W) and the logarithmic value of the transform block size in the vertical direction (log2H), and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, and context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (Method 2-1).
[0363] exist Fig.14 In the case of the example in A, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, and bypass encoding (binIdx=1) to the fourth bin (binIdx=3) is performed.
[0364] In addition, for example, Fig.14 As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context encoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context encoding can be performed on the second bin, and bypass encoding can be performed on the third and fourth bins in the bin sequence (Method 2-2).
[0365] exist Fig.14 In the example case of B, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and bypass coding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0366] In addition, for example, Fig.15 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context encoding can be performed on the first bin, predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context encoding can be performed on the second bin and the third bin, and bypass encoding can be performed on the fourth bin in the bin sequence (Method 2-3).
[0367] exist Fig.15 In the example case of A, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and bypass coding is performed on the fourth bin (binIdx=3).
[0368] In addition, for example, Fig.15As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context encoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 2-4).
[0369] exist Fig.15 In the example case of B, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context coding is performed on the fourth bin.
[0370] Note that Fig.14 and Fig.15 In the table, non-overlapping unique values are set in indexes B1, B2, and B3.
[0371] Fig.16 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in the table. For example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 2-1, the number of contexts is 4, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 2-2, the number of contexts is 5, the number of context coding bins is 2, and the number of bypass coding bins is 2. Furthermore, in the case of method 2-3, the number of contexts is 6, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 2-4, the number of contexts is 7, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0372] As described above, in any case of Method 2-1 to Method 2-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 2, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0373] Furthermore, in any of the cases of Method 2-1 to Method 2-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 2-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 2, bypass coding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0374] As described above, by applying method 2, an increase in the load of encoding processing can be suppressed.
[0375] <3-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0376] Similarly, in the case of decoding, the following is performed as in Figure 1 As has been performed in the table shown in B, a context variable is assigned to each bin in a bin sequence of a binarized adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding.
[0377] That is, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.
[0378] More specifically, the parameter about the block size is the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied. That is, a context variable based on the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding is performed on the first bin (method 2).
[0379] For example, Fig.14As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value of the transform block size in the horizontal direction (log2W) and the logarithmic value of the transform block size in the vertical direction (log2H), and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 2-1).
[0380] exist Fig.14 In the case of the example in A, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0381] In addition, for example, Fig.14 As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 2-2).
[0382] exist Fig.14 In the example case of B, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0383] In addition, for example, Fig.15 As shown in Table A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the bin of the bin identifier of the binarized adaptive orthogonal transform based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context decoding can be performed on the first bin, predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 2-3).
[0384] exist Fig.15 In the example case of A, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx=3).
[0385] In addition, for example, Fig.15 As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the bin of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied (max(log2W,log2H)-log2MinMtsSize), and context decoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 2-4).
[0386] exist Fig.15In the example case of B, the index ctxInc based on the difference (max(log2W,log2H)-log2MinMtsSize) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context decoding is performed on the fourth bin.
[0387] Note that even in the case of decoding, similar to the case of encoding, Fig.14 and Fig.15 Non-overlapping unique values are also set in indexes B1, B2, and B3 in the table.
[0388] The number of contexts, the number of context bins, and the number of bypass bins for each of these methods are similar to those for encoding ( Fig.16 ).
[0389] As described above, in any case of Method 2-1 to Method 2-4, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 2, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0390] Furthermore, in any of the cases of Method 2-1 to Method 2-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 2-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 2, bypass decoding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0391] As described above, by applying method 2, an increase in the load of the decoding process can be suppressed.
[0392] <3-3. Encoding side>
[0393] <Configuration>
[0394] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to that of the encoding side in the first embodiment. That is, the image encoding device 100 in this case has the same Figure 6In addition, the encoding unit 115 in this case has the same Figure 7 The configuration described is similar to the configuration.
[0395] <Encoding Process Flow>
[0396] In addition, the image encoding device 100 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image encoding processing performed by the image encoding device 100 in this case is performed by comparing the image encoding processing with the reference image encoding processing. Figure 8 The process described in the flowchart is similar to that described in the flowchart.
[0397] Will refer to Fig.17 The flowchart in describes an example of the flow of encoding processing for encoding an adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0398] In this encoding process, Fig. 9 The processing in step S131 and step S132 in 1 is similar to the processing in step S301 and step S302. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (that is, selects the context encoding). For example, the selection unit 132 selects Fig.14 and Fig.15 Any one of the tables shown in (i.e., by applying any one of Methods 2-1 to 2-4) selects context coding as the encoding method for that bin.
[0399] In step S303, the context setting unit 133 allocates the context variable ctx (index ctxInc) to the bin based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0400] The processing in steps S304 to S308 is also the same as Fig. 9 The processing in step S134 to step S138 in is similarly performed. In step S308, in the case where it is determined to terminate the encoding, the encoding process is terminated.
[0401] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 2 (for example, any one of method 2-1 to method 2-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0402] <3-4. Decoding side>
[0403] <Configuration>
[0404] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to that of the decoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0405] <Decoding Process Flow>
[0406] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0407] Will refer to Fig.18 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0408] In this decoding process, Fig.13 The processing of step S231 is performed similarly to the processing of step S321. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects the context decoding). For example, the selection unit 231 selects the context setting unit 232 according to Fig.14 and Fig.15 Any one of the tables shown in (i.e., by applying any one of Methods 2-1 to 2-4) selects context decoding as the decoding method for that bin.
[0409] In step S322, the context setting unit 232 allocates the context variable ctx (index ctxInc) to the bin based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.
[0410] The processing in steps S323 to S328 is also the same as Fig.13 The processing in step S233 to step S238 in is similarly performed. In step S328, in the case where it is determined to terminate the decoding, the decoding processing is terminated.
[0411] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 2 (for example, any one of method 2-1 to method 2-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0412] <4. Third Embodiment>
[0413] <4-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0414] In this embodiment, the following is performed as in Figure 1 As has been performed in the table shown in B of , a context variable is assigned to each bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (method 0).
[0415] That is, a context variable based on a parameter about a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0416] More specifically, the parameter regarding the block size is the minimum value between the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and a predetermined threshold. That is, a context variable based on the smaller value between the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the predetermined threshold is assigned to the first bin in the bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of the adaptive orthogonal transform in the image encoding, and context encoding is performed on the first bin (method 3).
[0417] Note that the threshold TH is an arbitrary value. By setting the threshold to a value smaller than the maximum value of the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of method 2. Therefore, the memory usage can be reduced.
[0418] For example, Fig.19 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value of the transform block size in the horizontal direction (log2W) and the logarithmic value of the transform block size in the vertical direction (log2H) (max(log2W,log2H)) and the logarithmic value of the minimum transform block size (log2MinMtsSize) to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) of the predetermined threshold (TH, and context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (Method 3-1).
[0419] exist Fig.19 In the example of A, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, and bypass encoding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0420] In addition, for example, Fig.19As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, the predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context encoding can be performed on the second bin, and bypass encoding can be performed on the third and fourth bins in the bin sequence (Method 3-2).
[0421] exist Fig.19 In the example case of B, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and bypass coding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0422] In addition, for example, Fig. 20 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context encoding can be performed on the second bin and the third bin, and bypass encoding can be performed on the fourth bin in the bin sequence (Method 3-3).
[0423] exist Fig. 20In the example case of A, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and bypass coding is performed on the fourth bin (binIdx=3).
[0424] In addition, for example, Fig. 20 As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 3-4).
[0425] exist Fig. 20 In the example case of B, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context coding is performed on the fourth bin.
[0426] Note that Fig.19 and Fig. 20 In the table, non-overlapping unique values are set in indexes B1, B2, and B3.
[0427] Fig.21 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in Table . Fig.21 The table shown in A shows an example in the case where the threshold value TH=2. In this example, for example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 3-1, the number of contexts is 3, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 3-2, the number of contexts is 4, the number of context coding bins is 2, and the number of bypass coding bins is 2. Furthermore, in the case of method 3-3, the number of contexts is 5, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 3-4, the number of contexts is 6, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0428] also, Fig.21 Table B shows an example in the case where the threshold value TH=1. In this example, in the case of method 3-1, the number of contexts is 2. Furthermore, in the case of method 3-2, the number of contexts is 3. Furthermore, in the case of method 3-3, the number of contexts is 4. Furthermore, in the case of method 3-4, the number of contexts is 5.
[0429] As described above, in any case of Method 3-1 to Method 3-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 3, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0430] Furthermore, in any of the cases of Method 3-1 to Method 3-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 3-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 3, bypass coding can be applied to bins corresponding to a transform type having relatively low selectivity. Therefore, an increase in the number of context coding bins can be suppressed while suppressing a decrease in encoding efficiency and an increase in processing volume (throughput) can be suppressed.
[0431] As described above, by applying method 3, an increase in the load of encoding processing can be suppressed.
[0432] <4-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0433] Similarly, in the case of decoding, the following is performed as in Figure 1As has been performed in the table shown in B, a context variable is assigned to each bin in a bin sequence of a binarized adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding.
[0434] That is, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.
[0435] More specifically, the parameter about the block size is the minimum value between the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and a predetermined threshold. That is, a context variable based on the smaller value between the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the predetermined threshold is assigned to the first bin in the bin sequence of the bin of the binarized adaptive orthogonal transform identifier, and context decoding is performed on the first bin (method 3).
[0436] Note that the threshold value TH is an arbitrary value as in the case of encoding. By setting the threshold value to a value smaller than the maximum value of the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of method 2. Therefore, the memory usage can be reduced.
[0437] For example, Fig.19 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value of the transform block size in the horizontal direction (log2W) and the logarithmic value of the transform block size in the vertical direction (log2H) (max(log2W,log2H)) and the logarithmic value of the minimum transform block size (log2MinMtsSize) to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) of the predetermined threshold (TH, and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 3-1).
[0438] exist Fig.19In the example in A, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0439] In addition, for example, Fig.19 As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, and the predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 3-2).
[0440] exist Fig.19 In the case of the example in B, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3) (Methods 2 and 3).
[0441] In addition, for example, Fig. 20As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the bin of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 3-3).
[0442] exist Fig. 20 In the example case of A, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx=3).
[0443] In addition, for example, Fig. 20 As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 3-4).
[0444] exist Fig. 20In the example case of B, the index ctxInc based on the minimum value (min(max(log2W,log2H)-log2MinMtsSize,TH)) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context decoding is performed on the fourth bin.
[0445] Note that even in the case of decoding, similar to the case of encoding, Fig.19 and Fig. 20 Non-overlapping unique values are also set in indexes B1, B2, and B3 in the table.
[0446] The number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are similar to those used for encoding ( Fig.21 A and Fig.21 B).
[0447] As described above, in any case of Method 3-1 to Method 3-4, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 3, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0448] Furthermore, in any of the cases of Method 3-1 to Method 3-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 3-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 3, bypass decoding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0449] As described above, by applying method 3, an increase in the load of the decoding process can be suppressed.
[0450] <4-3. Encoding side>
[0451] <Configuration>
[0452] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to that of the encoding side in the first embodiment. That is, the image encoding device 100 in this case has the same Figure 6 In addition, the encoding unit 115 in this case has the same Figure 7 The configuration described is similar to the configuration.
[0453] <Encoding Process Flow>
[0454] In addition, the image encoding device 100 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image encoding processing performed by the image encoding device 100 in this case is performed by comparing the image encoding processing with the reference image encoding processing. Figure 8 The process described in the flowchart is similar to that described in the flowchart.
[0455] Will refer to Fig. 22 The flowchart in describes an example of the flow of encoding processing for encoding an adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0456] In this encoding process, Fig. 9 The processing in steps S131 and S132 in 1 is similar to the processing in steps S351 and S352. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (that is, the context encoding is selected.). For example, the selection unit 132 selects Fig.19 and Fig. 20 Any one of the tables shown (ie, by applying any one of Methods 3-1 to 3-4) selects context coding as the encoding method for the bin.
[0457] In step S353, the context setting unit 133 allocates the context variable ctx (index ctxInc) to the bin based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) among the predetermined thresholds. Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0458] The processing in steps S354 to S358 is also the same as Fig. 9 The processing in steps S134 to S138 in is similarly performed. In step S358, in the case where it is determined to terminate the encoding, the encoding processing is terminated.
[0459] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 3 (for example, any one of method 3-1 to method 3-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0460] <4-4. Decoding side>
[0461] <Configuration>
[0462] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to that of the decoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0463] <Decoding Process Flow>
[0464] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0465] Will refer to Fig.23 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0466] In this decoding process, Fig.13 The processing of step S231 is similarly performed to the processing of step S371. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects the context decoding.). For example, the selection unit 231 selects the context setting unit 232 according to Fig.19 and Fig. 20 Any one of the tables shown (ie, by applying any one of Methods 3-1 to 3-4) selects context decoding as the decoding method for the bin.
[0467] In step S372, the context setting unit 232 allocates the context variable ctx (index ctxInc) to the bin based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) among the predetermined thresholds. Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.
[0468] The processing in steps S373 to S378 is also the same as Fig.13 The processing in steps S233 to S238 in is similarly performed. In step S378, in the case where it is determined to terminate the decoding, the decoding processing is terminated.
[0469] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 3 (for example, any one of method 3-1 to method 3-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0470] <5. Fourth embodiment>
[0471] <5-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0472] In this embodiment, the following is performed as in Figure 1 As has been performed in the table shown in B of , a context variable is assigned to each bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (method 0).
[0473] That is, a context variable based on a parameter about a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0474] More specifically, the parameter about the block size is a result of right-shifting the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which adaptive orthogonal transform can be applied and the minimum value among predetermined thresholds. That is, the result of right-shifting the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which adaptive orthogonal transform can be applied and the minimum value among predetermined thresholds is assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier indicating the mode of adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin (method 4).
[0475] Note that the threshold value TH is an arbitrary value. By setting the threshold value to a value smaller than the maximum value of the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of method 2. In addition, the value of the scaling parameter shift, which is the shift amount of the right shift, is arbitrary. In the case of method 4, since the minimum value is further right-shifted, the number of contexts can be reduced compared to the case of method 3. Therefore, the memory usage can be reduced.
[0476] For example, Fig.24 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the right shift result (min(max(log2W,log2H)-log2MinMtsSize)) between the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value (log2W) of the transform block size in the horizontal direction and the logarithmic value (log2H) of the transform block size in the vertical direction and the logarithmic value (log2MinMtsSize) of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) of the predetermined threshold (TH), and context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (Method 4-1).
[0477] exist Fig.24In the case of example A, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin, and bypass encoding (bypass) can be performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0478] In addition, for example, Fig.24 As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, the predetermined context variable (indexctxInc) can be assigned to the second bin in the bin sequence, and context encoding can be performed on the second bin, and bypass encoding can be performed on the third and fourth bins in the bin sequence (Method 4-2).
[0479] exist Fig.24 In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and bypass coding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3) (Methods 2 and 3).
[0480] In addition, for example, Fig.25As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, predetermined context variables (indexctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context encoding can be performed on the second bin and the third bin, and bypass encoding can be performed on the fourth bin in the bin sequence (Method 4-3).
[0481] exist Fig.25 In the case of the example in A, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and bypass coding is performed on the fourth bin (binIdx=3).
[0482] In addition, for example, Fig.25 As shown in Table B of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context encoding can be performed on the first bin, and predetermined context variables (indexctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 4-4).
[0483] exist Fig.25In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context coding is performed on the fourth bin.
[0484] Note that Fig.24 and Fig.25 In the table in , non-overlapping unique values are set in indexes B1, B2, and B3.
[0485] Fig.26 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in Table . Fig.26 The table shown in is an example where the threshold TH=2 and the scaling parameter (shift amount of the right shift) shift=1. In this example, for example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 4-1, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 4-2, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 2. Furthermore, in the case of method 4-3, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 4-4, the number of contexts is 5, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0486] As described above, in any case of Method 4-1 to Method 4-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 4, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0487] In addition, in any case of Method 4-1 to Method 4-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 4-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 4, bypass coding can be applied to bins corresponding to a transform type with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0488] As described above, by applying method 4, an increase in the load of encoding processing can be suppressed.
[0489] <5-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0490] Similarly, in the case of decoding, the following is performed as in Figure 1 As has been performed in the table shown in B, a context variable is assigned to each bin in a bin sequence of a binarized adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding.
[0491] That is, a context variable based on a parameter regarding a block size is assigned to a first bin in a bin sequence of a binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.
[0492] More specifically, the parameter about the block size is a result of right-shifting the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which adaptive orthogonal transform can be applied and the minimum value among the predetermined thresholds. That is, a context variable based on the result of right-shifting the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which adaptive orthogonal transform can be applied and the minimum value among the predetermined thresholds is assigned to the first bin in the bin sequence of the bin of the bin-valued adaptive orthogonal transform identifier, and context decoding is performed on the first bin (method 4).
[0493] Note that, as in the case of encoding, the threshold value TH is an arbitrary value. By setting the threshold value to a value smaller than the maximum value of the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied, the number of contexts can be reduced compared to the case of method 2. In addition, the value of the scaling parameter shift, which is the shift amount of the right shift, is arbitrary. In the case of method 4, since the minimum value is further right-shifted, the number of contexts can be reduced compared to the case of method 3. Therefore, the memory usage can be reduced.
[0494] For example, Fig.24As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the right shift result (min(max(log2W,log2H)-log2MinMtsSize)) between the difference (max(log2W,log2H)-log2MinMtsSize) between the longer one of the logarithmic value (log2W) of the transform block size in the horizontal direction and the logarithmic value (log2H) of the transform block size in the vertical direction and the logarithmic value (log2MinMtsSize) of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value (min(max(log2W,log2H)-log2MinMtsSize,TH)) of the predetermined threshold (TH), and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 4-1).
[0495] exist Fig.24 In the case of the example in A, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding (bypass) is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0496] In addition, for example, Fig.24 As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, the predetermined context variable (indexctxInc) can be assigned to the second bin in the bin sequence, and context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 4-2).
[0497] exist Fig.24In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3) (Methods 2 and 3).
[0498] In addition, for example, Fig.25 As shown in Table A of , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier through binarization based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, predetermined context variables (indexctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 4-3).
[0499] exist Fig.25 In the case of the example in A, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx=3).
[0500] In addition, for example, Fig.25As shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier based on the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the smaller value among the predetermined thresholds (min(max(log2W,log2H)-log2MinMtsSize,TH)), and context decoding can be performed on the first bin, and predetermined context variables (indexctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 4-4).
[0501] exist Fig.25 In the case of the example in B, the index ctxInc based on the result of the right shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift) is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin, and the index ctxInc=B3 is assigned to the fourth bin (binIdx=3), and context decoding is performed on the fourth bin.
[0502] Note that even in the case of decoding, similar to the case of encoding, Fig.24 and Fig.25 Non-overlapping unique values are also set in indexes B1, B2, and B3 in the table.
[0503] The number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are similar to those used for encoding ( Fig.26 ).
[0504] As described above, in any case of Method 4-1 to Method 4-4, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 4, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0505] Furthermore, in any of the cases of Method 4-1 to Method 4-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 4-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 4, bypass decoding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in coding efficiency and suppress an increase in processing volume (throughput).
[0506] As described above, by applying method 4, an increase in the load of the decoding process can be suppressed.
[0507] <5-3. Encoding side>
[0508] <Configuration>
[0509] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to that of the encoding side in the first embodiment. That is, the image encoding device 100 in this case has the same Figure 6 The configuration of the image encoding device 100 described in the previous embodiment is similar to that of the image encoding device 100 described in the previous embodiment. Figure 7 The configuration described is similar to the configuration.
[0510] <Encoding Process Flow>
[0511] In addition, the image encoding device 100 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image encoding processing performed by the image encoding device 100 in this case is performed by comparing the image encoding processing with the reference image encoding processing. Figure 8 The process described in the flowchart is similar to that described in the flowchart.
[0512] Will refer to Fig. 27 The flowchart in describes an example of the flow of encoding processing for encoding an adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0513] In this encoding process, Fig. 9 The processing in steps S131 and S132 in 1 is similar to the processing in steps S401 and S402. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (that is, the context encoding is selected). For example, the selection unit 132 selects Fig.24 and Fig.25 Any one of the tables shown (ie, by applying any one of Methods 4-1 to 4-4) selects context coding as the encoding method for the bin.
[0514] In step S403, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bin based on the result of right-shifting the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the minimum value among the predetermined thresholds by the scaling parameter shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift). Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0515] The processing in steps S404 to S408 is also the same as Fig. 9 The processing in steps S134 to S138 in is similarly performed. In step S408, in the case where it is determined to terminate the encoding, the encoding processing is terminated.
[0516] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 4 (for example, any one of method 4-1 to method 4-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0517] <5-4. Decoding side>
[0518] <Configuration>
[0519] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to that of the decoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0520] <Decoding Process Flow>
[0521] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0522] Will refer to Fig.28 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0523] In this decoding process, Fig.13 The processing of step S231 is similarly performed in step S421. That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects the context decoding.). For example, the selection unit 231 selects the context setting unit 232 according to Fig.24 and Fig.25 Any one of the tables shown (ie, by applying any one of Method 4-1 to Method 4-4) selects context decoding as the decoding method for the bin.
[0524] In step S422, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin based on the result of right-shifting the difference between the longer of the logarithmic value of the transform block size in the horizontal direction and the logarithmic value of the transform block size in the vertical direction and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform can be applied and the minimum value among the predetermined thresholds by the scaling parameter shift (min(max(log2W,log2H)-log2MinMtsSize,TH)>>shift). Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.
[0525] The processing in steps S423 to S428 is also the same as Fig.13 The processing in steps S233 to S238 in is similarly performed. In step S428, in the case where it is determined to terminate the decoding, the decoding processing is terminated.
[0526] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 4 (for example, any one of method 4-1 to method 4-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0527] <6. Fifth embodiment>
[0528] <6-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0529] In this embodiment, the following is performed as in Figure 1 As has been performed in the table shown in B of , a context variable is assigned to each bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding (method 0).
[0530] That is, a context variable based on a parameter regarding the block size is assigned to the first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image coding, and context coding is performed on the first bin in the bin sequence.
[0531] More specifically, the context variable is assigned and context coding is performed (Method 5) according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold.
[0532] For example, as Fig.29 shown in Table A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than the predetermined threshold (TH) (S < TH?), context coding can be performed on the first bin, and bypass coding can be performed on the second to fourth bins in the bin sequence (Method 5-1).
[0533] In Fig.29 the example in A, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) of the adaptive orthogonal transform identifier, and context coding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, bypass coding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).
[0534] Furthermore, for example, as Fig.29 shown in Table B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than the predetermined threshold (S < TH?), context coding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin, context coding can be performed on the second bin, and bypass coding can be performed on the third and fourth bins in the bin sequence (Method 5-2).
[0535] In Fig.29In the case of the example in B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. And bypass encoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).
[0536] In addition, for example, as Fig.30 shown in Table A, depending on whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin, and context encoding can be performed on the second bin and the third bin. And bypass encoding can be performed on the fourth bin in the bin sequence (Method 5-3).
[0537] In Fig.30 the case of the example in A, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin. While when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin. The index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin. And bypass encoding is performed on the fourth bin (binIdx = 3).
[0538] In addition, for example, as Fig.30In the table shown in B, a context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing an adaptive orthogonal transform identifier according to whether a parameter regarding a block size is equal to or greater than a predetermined threshold (S < TH?), context encoding can be performed on the first bin, predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins, and context encoding can be performed on the second to fourth bins (Method 5-4).
[0539] In Fig.30 In the case of the example in B, when the parameter regarding the block size is less than the predetermined threshold (S < TH), index ctxInc = A0 is assigned to the first bin (binIdx = 0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, while when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), index ctxInc = A1 is assigned to the first bin, and context encoding is performed on the first bin, index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context encoding is performed on the second bin, index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context encoding is performed on the third bin, and index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context encoding is performed on the fourth bin.
[0540] Note that in Fig.29 and Fig.30 the table, non-overlapping unique values are set among indices A0, A1, B1, B2, and B3.
[0541] Note that parameter S can be any parameter as long as parameter S is related to the block size.
[0542] For example, parameter S can be the product of the width tbWidth and height tbHeight of the transform block, i.e., the area of the transform block.
[0543] S = tbWidth * tbHeight
[0544] In addition, parameter S can be the sum of the logarithm of the width tbWidth and the logarithm of the height tbHeight of the transform block, i.e., the logarithm of the area of the transform block.
[0545] S = log2(tbWidth) + log2(tbHeight)
[0546] In addition, parameter S can be the maximum value of the width tbWidth and height tbHeight of the transform block, i.e., the size of the long side of the transform block.
[0547] S = max(tbWidth, tbHeight)
[0548] In addition, the parameter S may be the minimum value of the width tbWidth and the height tbHeight of the transform block, that is, the size of the short side of the transform block.
[0549] S = min(tbWidth, tbHeight)
[0550] In addition, the parameter S may be a maximum value of a logarithmic value of a width tbWidth and a logarithmic value of a height tbHeight of the transform block, that is, a logarithmic value of a size of a long side of the transform block.
[0551] S=max(log2(tbWidth),log2(tbHeight))
[0552] In addition, the parameter S may be a minimum value of a logarithmic value of a width tbWidth and a logarithmic value of a height tbHeight of the transform block, that is, a logarithmic value of a size of a short side of the transform block.
[0553] S=min(log2(tbWidth),log2(tbHeight))
[0554] In addition, the parameter S may be a ratio cbSubDiv of the area of the coding block to the area of the CTU. An example of a ratio cbSubDiv of the area of the coding block to the area of the CTU is as follows: Fig.31 As shown in the table.
[0555] S=cbSubDiv
[0556] In addition, the parameter S may be the result of right-shifting the ratio cbSubDiv of the area of the coding block to the area of the CTU by the scaling parameter shift. Note that the scaling parameter shift may be the difference between the logarithmic value of the block size of the CTU and the logarithmic value of the maximum transform block size. In addition, the logarithmic value of the maximum transform block size may be 5.
[0557] S=cbSubDiv>>shift
[0558] shift=(log2CTUSize-log2MaxTsSize)
[0559] log2MaxTsSize=5
[0560] Also, the parameter S may be an absolute value of a difference between a logarithmic value of a width tbWidth and a logarithmic value of a height tbHeight of the transform block.
[0561] S=abs(log2(tbWidth)-log2(tbHeight))
[0562] Fig.32 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in the table. In this example, for example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 5-1, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 5-2, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 2. In addition, in the case of method 5-3, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 5-4, the number of contexts is 5, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0563] As described above, in any case of Method 5-1 to Method 5-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 5, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0564] In addition, in any case of Method 5-1 to Method 5-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 5-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 5, bypass coding can be applied to bins corresponding to a transform type with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins and suppress an increase in processing volume (throughput) while suppressing a decrease in encoding efficiency.
[0565] As described above, by applying method 5, an increase in the load of encoding processing can be suppressed.
[0566] <6-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0567] Similarly, in the case of decoding, the following is performed as in Figure 1 As has been performed in the table shown in B, a context variable is assigned to each bin in a bin sequence of a binarized adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding.
[0568] That is, a context variable based on a parameter regarding the block size is assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context encoding is performed on the first bin in the bin sequence.
[0569] More specifically, the context variable is assigned and context decoding is performed (Method 5) according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold.
[0570] For example, as Fig.29 shown in the table of A in Fig.29 , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the parameter (S) regarding the block size is equal to or greater than the predetermined threshold (TH) (S < TH?), context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 5-1).
[0571] In Fig.29 the example of A in Fig.29 , when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter of the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, bypass decoding is performed on the second bin (binIdx = 1) to the fourth bin (binIdx = 3).
[0572] Furthermore, for example, as Fig.29 shown in the table of B in Fig.29 , the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin, context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 5-2).
[0573] In Fig.29In the example of B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx = 2) and the fourth bin (binIdx = 3).
[0574] In addition, for example, as Fig.30 shown in the table of A, according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin and the third bin, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 5-3).
[0575] In Fig.30 the example of A, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter of the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin, and bypass decoding is performed on the fourth bin (binIdx = 3).
[0576] In addition, for example, as Fig.30 shown in the table of B, according to whether the parameter regarding the block size is equal to or greater than the predetermined threshold (S < TH?), the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence of the binarized adaptive orthogonal transform identifier, and context decoding can be performed on the first bin. Different predetermined context variables (index ctxInc) can be assigned to the second bin to the fourth bin, and context decoding can be performed on the second bin to the fourth bin (Method 5-4).
[0577] In Fig.25 the example in B, when the parameter regarding the block size is less than a predetermined threshold (S < TH), the index ctxInc = A0 is assigned to the first bin (binIdx = 0) of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin. When the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the index ctxInc = A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, the index ctxInc = B1 is assigned to the second bin (binIdx = 1), and context decoding is performed on the second bin. The index ctxInc = B2 is assigned to the third bin (binIdx = 2), and context decoding is performed on the third bin. And the index ctxInc = B3 is assigned to the fourth bin (binIdx = 3), and context decoding is performed on the fourth bin.
[0578] Note that even in the case of decoding, similar to the case of encoding, non - overlapping unique values are also set in the indexes A0, A1, B1, B2, and B3 in the table of Fig.29 and Fig.30 .
[0579] In addition, in the case of decoding, the parameter S can be any parameter as long as the parameter S is related to the block size, similar to the case of encoding. For example, the parameter S can be derived by various methods described in <6 - 1. Encoding of Adaptive Orthogonal Transform Identifiers>.
[0580] The number of contexts, the number of context - encoding bins, and the number of bypass - encoding bins for each of these methods are similar to those for encoding ( Fig.32 ).
[0581] As described above, in any of the cases of Method 5 - 1 to Method 5 - 4, compared with the case of Method 0, the number of contexts required for decoding can be reduced. That is, by applying Method 5, the number of contexts assigned to the first bin (binIdx = 0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0582] In addition, in any of the cases of Method 5 - 1 to Method 5 - 3, compared with the case of Method 0, the number of context - encoding bins required for decoding can be reduced. Note that in the case of Method 5 - 4, the number of context - encoding bins required for decoding is equivalent to the number of context - encoding bins required for decoding in the case of Method 0. That is, by applying Method 5, bypass decoding can be applied to the bins corresponding to the transform types with relatively low selectivity. Therefore, an increase in the number of context - encoding bins and an increase in the processing amount (throughput) can be suppressed while suppressing a decrease in encoding efficiency.
[0583] As described above, by applying Method 5, an increase in the load of the decoding process can be suppressed.
[0584] <6-3. Encoding side>
[0585] <Configuration>
[0586] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side of the first embodiment. That is, the image encoding apparatus 100 in this case has a configuration similar to the configuration described with reference to Figure 6 In addition, the encoding unit 115 in this case has a configuration similar to the configuration described with reference to Figure 7 described.
[0587] <Flow of encoding process>
[0588] In addition, the image encoding apparatus 100 in this case performs processing substantially similar to that in the first embodiment. That is, the image encoding process performed by the image encoding apparatus 100 in this case is performed through a process similar to the process described in the flowchart with reference to Figure 8 described.
[0589] Reference will be made to Fig.33 in the flowchart to describe an example of the flow of the encoding process for encoding the adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0590] In this encoding process, the processing in steps S451 and S452 is performed similarly to the processing in steps S131 and S132 in Fig. 9 That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bins (i.e., context encoding is selected). For example, the selection unit 132 selects context encoding as the encoding method for the bins according to any one of the tables shown in Fig.29 and Fig.30 (i.e., by applying any one of Methods 5-1 to 5-4).
[0591] In step S453, the context setting unit 133 assigns the context variable ctx (index ctxInc) to the bins according to whether the parameter (S) regarding the block size is equal to or greater than a predetermined threshold (TH) (S < TH?).
[0592] For example, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the context setting unit 133 assigns the index ctxInc = A0 to the bins. In addition, when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the context setting unit 133 assigns the index ctxInc = A1 to the bins.
[0593] Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0594] The processing in steps S454 to S458 is also the same as Fig. 9 The processing in steps S134 to S138 in is similarly performed. In step S408, in the case where it is determined to terminate the encoding, the encoding processing is terminated.
[0595] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 5 (for example, any one of method 5-1 to method 5-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in processing volume (throughput) can be suppressed. That is, an increase in the load of encoding processing can be suppressed.
[0596] <6-4. Decoding side>
[0597] <Configuration>
[0598] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to that of the decoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0599] <Decoding Process Flow>
[0600] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0601] Will refer to Fig.34 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0602] In this decoding process, Fig.13 The processing of step S471 is performed similarly to the processing of step S231 of . That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects the context decoding.). For example, the selection unit 231 selects Fig.29 and Fig.30Select context decoding as the decoding method for the bin from any one of the tables shown (i.e., by applying any one of Methods 5-1 to 5-4).
[0603] In step S472, the context setting unit 232 assigns the context variable ctx (index ctxInc) to the bin according to whether the parameter (S) regarding the block size is equal to or greater than a predetermined threshold (TH) (S < TH?).
[0604] For example, when the parameter regarding the block size is less than the predetermined threshold (S < TH), the context setting unit 232 assigns the index ctxInc = A0 to the bin. In addition, when the parameter regarding the block size is equal to or greater than the predetermined threshold (S ≥ TH), the context setting unit 232 assigns the index ctxInc = A1 to the bin.
[0605] Then, the context decoding unit 233 performs arithmetic decoding using the context variable. That is, context decoding is performed.
[0606] The processing in steps S473 to S478 is also performed similarly to the processing in steps S233 to S238 in Fig.13 In step S478, when it is determined to terminate the decoding, the decoding process is terminated.
[0607] By performing each process in this way, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying Method 5 (for example, any one of Methods 5-1 to 5-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the processing amount (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0608] <7. Sixth Embodiment>
[0609] <7-1. Encoding of Adaptive Orthogonal Transform Identifier>
[0610] In this embodiment, binarization (Method 0) of the adaptive orthogonal transform identifier indicating the mode of adaptive orthogonal transform in image encoding is performed as already performed in the table shown in A of Figure 2 That is, the adaptive orthogonal transform identifier is binarized into a bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types, and is encoded (Method 6).
[0611] For example, as in
[0612] For example, as Fig.35In the table shown in A, when the transform type of the adaptive orthogonal transform identifier is DCT2×DCT2, the adaptive orthogonal transform identifier is binarized into a bin sequence of one bin (=0), and when the transform type of the adaptive orthogonal transform identifier is different from DCT2×DCT2, the adaptive orthogonal transform identifier is binarized into a bin sequence of three bins.
[0613] exist Fig.35 In the case of example A, the adaptive orthogonal transform identifier mts_idx=0 is binarized as a bin sequence "0" indicating that the transform type is DCT2×DCT2, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx=1 is binarized as a bin sequence "100" indicating that the transform type is different from DCT2×DCT2 and is DST7×DST7, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx=2 is binarized as a bin sequence "101" indicating that the transform type is different from DCT2×DCT2 and is DCT8×DST7, and is encoded. In addition, the adaptive orthogonal transform identifier mts_idx=3 is binarized as a bin sequence "110" indicating that the transform type is different from DCT2×DCT2 and is DST7×DCT8, and is encoded. Furthermore, the adaptive orthogonal transform identifier mts_idx=4 is binarized into a bin sequence "111" indicating that the transform type is different from DCT2×DCT2 and is DCT8×DCT8, and is encoded.
[0614] By binarizing the adaptive orthogonal transform identifier in this way, the length of the bin sequence (bin length) can be at most three bins. Figure 2 In the case of example A (method 0), the length of the bin sequence (bin length) is at most 4 bins, so by applying method 6, the bin length can be shortened by one bin.
[0615] Note that, as in Fig.35 In the table shown in B of , the bin sequence may be divided into a prefix part of one bin indicating whether the transform type is different from DCT2×DCT2 and a suffix part of two bins indicating other transform types, and may be binarized.
[0616] Furthermore, a context variable may be assigned to each bin in a bin sequence of an adaptive orthogonal transform identifier generated by binarization according to Method 6 by any one of Methods 0 to 5, and encoding may be performed.
[0617] For example, Fig.36In the table shown in A, a predetermined context variable ctx (index ctxInc), that is, a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin, and bypass encoding can be performed on the second and third bins in the bin sequence (method 1-1).
[0618] exist Fig.36 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context encoding is performed on the first bin, and bypass encoding is performed on the second bin (binIdx=1) and the third bin (binIdx=2).
[0619] In addition, for example, Fig.36 In the table shown in B, predetermined context variables ctx (index ctxInc) different from each other can be assigned to the first bin and the second bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first bin and the second bin, and bypass encoding can be performed on the third bin in the bin sequence (method 1-2).
[0620] exist Fig.36 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and bypass coding is performed on the third bin (binIdx=2).
[0621] In addition, for example, Fig.36 In the table shown in C, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context encoding can be performed on the first to third bins (method 1-3).
[0622] exist Fig.36 In the example case of C, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context coding is performed on the second bin, and the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context coding is performed on the third bin.
[0623] Note that Fig.36 In the table, non-overlapping unique values are set in indexes A0, B1 and B2.
[0624] exist Fig.37 An example of the number of contexts, the number of context coding bins, the number of bypass coding bins, and the worst-case bin length for each of these methods is shown in the table. For example, in the case of method 0, the worst-case bin length is 4. For example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of a combination of method 6 and method 1-1 (method 6-1), the worst-case bin length is 3. Therefore, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 2. In addition, in the case of a combination of method 6 and method 1-2 (method 6-2), the worst-case bin length is 3. Therefore, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 1. In addition, in the case of a combination of method 6 and method 1-3 (method 6-3), the worst-case bin length is 3. Therefore, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 0.
[0625] As described above, in any case of Method 6-1 to Method 6-3, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 6, the bin length in the worst case can be reduced, and further, by applying Method 1, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0626] Furthermore, in any case of Method 6-1 to Method 6-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 6, the bin length in the worst case is reduced, and further, by applying Method 1, bypass coding can be applied to bins corresponding to a transform type having relatively low selectivity. Therefore, an increase in the number of context coding bins can be suppressed while suppressing a decrease in encoding efficiency and an increase in processing volume (throughput) can be suppressed.
[0627] Note that the effect of method 6 can be combined with the effect of another method by combining another method (for example, any one of methods 2 to 5) instead of method 1 with method 6. Therefore, in either case, increases in memory usage and processing volume (throughput) can be suppressed.
[0628] As described above, by applying method 6, an increase in the load of encoding processing can be suppressed.
[0629] <7-2. Decoding of Adaptive Orthogonal Transform Identifier>
[0630] In this embodiment, the following is performed as in Figure 2 Inverse binarization of the adaptive orthogonal transform identifier indicating the mode of the adaptive orthogonal inverse transform in image decoding has been performed as shown in the table shown in A (method 0).
[0631] That is, the bin sequence obtained by decoding, which is composed of one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types, is debinarized to derive an adaptive orthogonal transform identifier (Method 6).
[0632] For example, Fig.35 In the table shown in A, the bin sequence of one bin (=0) is debinarized to derive an adaptive orthogonal transform identifier whose transform type is DCT2×DCT2, and the bin sequence of three bins is debinarized to derive an adaptive orthogonal transform identifier whose transform type is different from DCT2×DCT2.
[0633] exist Fig.35 In the case of the example in A of , the bin sequence "0" obtained by decoding the coded data is debinarized to derive an adaptive orthogonal transform identifier mts_idx=0 indicating that the transform type is DCT2×DCT2. In addition, the bin sequence "100" obtained by decoding the coded data is debinarized to derive an adaptive orthogonal transform identifier mts_idx=1 indicating that the transform type is different from DCT2×DCT2 and is DST7×DST7. In addition, the bin sequence "101" obtained by decoding the coded data is debinarized to derive an adaptive orthogonal transform identifier mts_idx=2 indicating that the transform type is different from DCT2×DCT2 and is DCT8×DST7. In addition, the bin sequence "110" obtained by decoding the coded data is debinarized to derive an adaptive orthogonal transform identifier mts_idx=3 indicating that the transform type is different from DCT2×DCT2 and is DST7×DCT8. Furthermore, the bin sequence "111" obtained by decoding the encoded data is debinarized to derive an adaptive orthogonal transform identifier mts_idx=4 indicating that the transform type is different from DCT2×DCT2 and is DST8×DCT8.
[0634] By debinarizing the adaptive orthogonal transform identifier in this way, the length of the bin sequence (bin length) can be made at most three bins. Figure 2In the case of example A (method 0), the length of the bin sequence (bin length) is at most 4 bins, so by applying method 6, the bin length can be shortened by one bin.
[0635] Note that, as in Fig.35 In the table shown in B of , the bin sequence may be divided into a prefix part of one bin indicating whether the transform type is different from DCT2×DCT2 and a suffix part of two bins indicating other transform types, and is inverse-binarized.
[0636] Furthermore, in the case of applying such debinarization of method 6, a context variable may be assigned to each bin in a bin sequence by any one of methods 0 to 5, and decoding may be performed.
[0637] For example, Fig.36 In the table shown in A, a predetermined context variable ctx (index ctxInc), that is, a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first bin, and bypass decoding can be performed on the second and third bins in the bin sequence (method 1-1).
[0638] exist Fig.36 In the example case of A, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and bypass decoding is performed on the second bin (binIdx=1) to the third bin (binIdx=2).
[0639] For example, Fig.36 In the table shown in A, a predetermined context variable ctx (index ctxInc), that is, a fixed (one-to-one) context variable ctx, can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first bin, and bypass decoding can be performed on the second and third bins in the bin sequence (method 1-1).
[0640] exist Fig.36 In the example case of B, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, and bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0641] In addition, for example, Fig.36In the table shown in C, different predetermined context variables ctx (index ctxInc) can be assigned to the first to third bins in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier, and context decoding can be performed on the first to third bins (method 1-3).
[0642] exist Fig.36 In the example case in C, the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, the index ctxInc=B1 is assigned to the second bin (binIdx=1), and context decoding is performed on the second bin, the index ctxInc=B2 is assigned to the third bin (binIdx=2), and context decoding is performed on the third bin.
[0643] Note that Fig.36 In the table, non-overlapping unique values are set in indexes A0, B1 and B2.
[0644] The number of contexts, number of context coding bins, number of bypass coding bins, and worst-case bin length for each of these methods are compared to the case when encoding ( Fig.37 )similar.
[0645] As described above, in any case of Method 6-1 to Method 6-3, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 6, the bin length in the worst case can be reduced, and further, by applying Method 1, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0646] Furthermore, in any case of Method 6-1 to Method 6-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 6, the bin length in the worst case is reduced, and further, by applying Method 1, bypass decoding can be applied to bins corresponding to a transform type having relatively low selectivity. Therefore, an increase in the number of context coding bins can be suppressed while suppressing a decrease in encoding efficiency and an increase in processing volume (throughput) can be suppressed.
[0647] Note that the effect of method 6 can be combined with the effect of another method by combining another method (for example, any one of methods 2 to 5) with method 6 instead of method 1. Therefore, in either case, the increase in memory usage and processing volume (throughput) can be suppressed.
[0648] As described above, by applying method 6, an increase in the load of the decoding process can be suppressed.
[0649] <7-3. Encoding side>
[0650] <Configuration>
[0651] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to that of the encoding side in the first embodiment. That is, the image encoding device 100 in this case has the same Figure 6 In addition, the encoding unit 115 in this case has the same Figure 7 The configuration described is similar to the configuration.
[0652] <Encoding Process Flow>
[0653] In addition, the image encoding device 100 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image encoding processing performed by the image encoding device 100 in this case is performed by comparing the image encoding processing with the reference image encoding processing. Figure 8 The process described in the flowchart is similar to that described in the flowchart.
[0654] Will refer to Fig.38 The flowchart in describes an example of the flow of encoding processing for encoding an adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0655] When the encoding process starts, in step S501, the binarization unit 131 of the encoding unit 115 binarizes the adaptive orthogonal transform identifier mts_idx into a bin sequence consisting of one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types.
[0656] The processing in steps S502 to S508 is also the same as Fig. 9 The processing in steps S132 to S138 in is performed similarly.
[0657] By performing each process in this manner, the encoding unit 115 can binarize the adaptive orthogonal transform identifier by applying method 6, and can encode the bin sequence by applying method 1 (for example, any one of methods 1-1 to 1-3). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.
[0658] <7-4. Decoding side>
[0659] <Configuration>
[0660] Next, the decoding side will be described. The configuration of the encoding side in this case is similar to the configuration of the encoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0661] <Decoding Process Flow>
[0662] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0663] Will refer to Fig.39 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0664] In this decoding process, Fig.13 The processing in steps S231 to S236 in is similarly performed as the processing in steps S521 to S526.
[0665] In step S527, the debinarization unit 235 debinarizes the bin sequence configured by one bin (one bit) indicating whether the transform type is different from the transform type DCT2×DCT2 and two bins (two bits) indicating other transform types to derive the adaptive orthogonal transform identifier mts_idx.
[0666] In step S528, the decoding unit 212 determines whether to terminate the decoding of the adaptive orthogonal transform identifier mts_idx. In the case where it is determined that the decoding is not terminated, the process returns to step S523, and the processes in step S523 and subsequent steps are repeated. In addition, in step S528, in the case where it is determined that the decoding is terminated, the decoding process is terminated.
[0667] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 1 (e.g., any one of methods 1-1 to 1-3) to derive the bin sequence. In addition, the decoding unit 212 can debinarize the bin sequence by applying method 6 to derive the adaptive orthogonal transform identifier. Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0668] <8. Seventh Implementation Method>
[0669] <8-1. Encoding of Transform Skip Flag and Adaptive Orthogonal Transform Identifier>
[0670] In the case of the method described in Non-Patent Literature 1, an adaptive orthogonal transform identifier tu_mits_idx and a transform skip flag transform_skip_flag indicating whether transform skipping is applied are luma-limited (4:2:0 format-limited).
[0671] Therefore, when the color difference array type is greater than 1 (i.e., when the color difference format is 4:2:2 or 4:4:4), the (encoding / decoding) transform skip flag and adaptive orthogonal transform identifier are signaled for each component ID (cIdx) so that transform skip and adaptive orthogonal transform can be applied to the color difference component (Method 7).
[0672] An example of the syntax of the transform_unit in this case is Fig.40 As shown. Fig.40 As shown, the transform mode (transform_mode) is signaled for luma (Y), chroma (Cb), and chroma (Cr) as shown in the 17th, 20th, and 23rd lines (gray lines) from the top of the syntax. Fig.41 An example of the syntax of the transform mode (transform_mode) in this case is shown. Fig.41 As shown, the transform skip flag transform_skip_flag[x0][y0][cIdx] is signaled in the fifth line (gray line) from the top of the syntax. In addition, the adaptive orthogonal transform identifier tu_mts_idx[x0][y0][cIdx] is signaled in the eighth line (gray line) from the top of the syntax. That is, the transform skip flag and the adaptive orthogonal transform identifier are signaled for each component ID (cIdx).
[0673] By doing so, it is possible to control the application of adaptive orthogonal transform to the color difference format 4:2:2 or 4:4:4 having a larger amount of information than the color difference format 4:2:0. Therefore, it is possible to suppress a decrease in encoding efficiency.
[0674] Then, in this case, the context variable ctx can be assigned to the first bin (binIdx=0) in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0, and context encoding can be performed.
[0675] For example, Fig.42In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context encoding can be performed on the first bin, and bypass encoding can be performed on the second to fourth bins in the bin sequence (Method 7-1).
[0676] exist Fig.42 In the case of the example in A of , when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, bypass coding is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0677] In addition, for example, Fig.42 In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context encoding can be performed on the first bin, and a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context encoding can be performed on the second bin, and bypass encoding can be performed on the third and fourth bins in the bin sequence (Method 7-2).
[0678] exist Fig.42 In the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context coding is performed. In addition, bypass coding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0679] In addition, for example, Fig.43In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context coding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context coding can be performed on the second bin and the third bin, and bypass coding can be performed on the fourth bin in the bin sequence (Method 7-3).
[0680] exist Fig.43 In the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context coding is performed. In addition, in the third bin (binIdx=2), the index ctxInc=B2 is assigned and context coding is performed. In addition, bypass coding is performed on the fourth bin (binIdx=3).
[0681] In addition, for example, Fig.43 In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context encoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context encoding can be performed on the second to fourth bins (Method 7-4).
[0682] exist Fig.43In the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context coding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context coding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context coding is performed. In addition, in the third bin (binIdx=2), the index ctxInc=B2 is assigned and context coding is performed. In addition, in the fourth bin (binIdx=3), the index ctxInc=B3 is assigned and context coding is performed.
[0683] Note that Fig.42 and Fig.43 In the table, non-overlapping unique values are set in indexes A0, A1, B1, B2 and B3.
[0684] Fig.44 Examples of the number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are shown in the table. For example, in the case of method 0, the number of contexts is 9, the number of context coding bins is 4, and the number of bypass coding bins is 0. In contrast, in the case of method 7-1, the number of contexts is 2, the number of context coding bins is 1, and the number of bypass coding bins is 3. Furthermore, in the case of method 7-2, the number of contexts is 3, the number of context coding bins is 2, and the number of bypass coding bins is 2. In addition, in the case of method 7-3, the number of contexts is 4, the number of context coding bins is 3, and the number of bypass coding bins is 1. Furthermore, in the case of method 7-4, the number of contexts is 5, the number of context coding bins is 4, and the number of bypass coding bins is 0.
[0685] As described above, in any case of Method 7-1 to Method 7-4, the number of contexts required for encoding can be reduced compared to the case of Method 0. That is, by applying Method 7, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0686] In addition, in any case of Method 7-1 to Method 7-3, the number of context coding bins required for encoding can be reduced compared to the case of Method 0. Note that in the case of Method 7-4, the number of context coding bins required for encoding is comparable to the number of context coding bins required for encoding in the case of Method 0. That is, by applying Method 7, bypass coding can be applied to bins corresponding to a transform type with relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in encoding efficiency and suppress an increase in processing volume (throughput).
[0687] Then, as described above, for the color difference format 4:2:2 or 4:4:4 having a larger amount of information than the color difference format 4:2:0, the application of adaptive orthogonal transform can be controlled. Therefore, the reduction in encoding efficiency can be suppressed.
[0688] As described above, by applying method 7, an increase in the load of encoding processing can be suppressed.
[0689] Note that the control parameters of transform skip or adaptive orthogonal transform may be signaled (coded) for each treeType instead of color component ID (cIdx). That is, [cIdx] of each control parameter may be replaced with [treeType].
[0690] In addition, the above method of assigning a context variable to each bin in the bin sequence of the adaptive orthogonal transform identifier mts_idx can also be applied to other syntax elements related to orthogonal transform, etc. For example, the method can also be applied to the secondary transform identifier st_idx and the transform skip flag ts_flag.
[0691] <8-2. Decoding of Transform Skip Flag and Adaptive Orthogonal Transform Identifier>
[0692] Similarly, in the case of decoding, when the color difference array type is greater than 1 (i.e., when the color difference format is 4:2:2 or 4:4:4), the transform skip flag and adaptive orthogonal transform identifier are signaled (decoded) for each component ID (cIdx) so that transform skip and adaptive orthogonal transform can be applied to the color difference component (Method 7).
[0693] An example of the syntax of a transform_unit in this case is Fig.40 As shown. Fig.40 As shown, the transform mode (transform_mode) is signaled for luma (Y), chroma (Cb), and chroma (Cr) as shown in the 17th, 20th, and 23rd lines (gray lines) from the top of the syntax. Fig.41An example of the syntax of the transform mode (transform_mode) in this case is shown. Fig.41 As shown, the transform skip flag transform_skip_flag[x0][y0][cIdx] is signaled in the fifth line (gray line) from the top of the syntax. In addition, the adaptive orthogonal transform identifier tu_mts_idx[x0][y0][cIdx] is signaled in the eighth line (gray line) from the top of the syntax. That is, the transform skip flag and the adaptive orthogonal transform identifier are signaled for each component ID (cIdx).
[0694] By doing so, it is possible to control the application of inverse adaptive orthogonal transform for the color difference format 4:2:2 or 4:4:4 having a larger amount of information than the color difference format 4:2:0. Therefore, it is possible to suppress a decrease in encoding efficiency.
[0695] Then, in this case, the context variable ctx may be assigned to the first bin (binIdx=0) in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0, and context decoding may be performed.
[0696] For example, Fig.42 In the table shown in A, the context variable ctx can be assigned to the first bin (binIdx=0) in the bin sequence of the binarized adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context decoding can be performed on the first bin, and bypass decoding can be performed on the second to fourth bins in the bin sequence (Method 7-1).
[0697] exist Fig.42 In the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and the first bin performs context decoding, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, bypass decoding is performed on the second bin (binIdx=1) to the fourth bin (binIdx=3).
[0698] In addition, for example, Fig.42In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context decoding can be performed on the first bin, a predetermined context variable (index ctxInc) can be assigned to the second bin in the bin sequence, and context decoding can be performed on the second bin, and bypass decoding can be performed on the third and fourth bins in the bin sequence (Method 7-2).
[0699] exist Fig.42 In the case of the example in B, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context decoding is performed. In addition, bypass decoding is performed on the third bin (binIdx=2) and the fourth bin (binIdx=3).
[0700] In addition, for example, Fig.43 In the table shown in A, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context decoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second bin and the third bin in the bin sequence, and context decoding can be performed on the second bin and the third bin, and bypass decoding can be performed on the fourth bin in the bin sequence (Method 7-3).
[0701] exist Fig.43In the case of the example in A, when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context decoding is performed. In addition, in the third bin (binIdx=2), the index ctxInc=B2 is assigned and context decoding is performed. In addition, bypass decoding is performed on the fourth bin (binIdx=3).
[0702] In addition, for example, Fig.43 In the table shown in B, the context variable ctx (index ctxInc) can be assigned to the first bin in the bin sequence obtained by binarizing the adaptive orthogonal transform identifier according to whether the component ID (cIdx) of the transform block is 0 ((cIdx==0)?), and context decoding can be performed on the first bin, and predetermined context variables (index ctxInc) different from each other can be assigned to the second to fourth bins in the bin sequence, and context decoding can be performed on the second to fourth bins (Method 7-4).
[0703] exist Fig.43 In the case of the example in B in , when the component ID (cIdx) of the transform block is 0 (cIdx==0), the index ctxInc=A0 is assigned to the first bin (binIdx=0) in the bin sequence of the adaptive orthogonal transform identifier, and context decoding is performed on the first bin, and when the component ID (cIdx) of the transform block is not 0 (cIdx>0), the index ctxInc=A1 is assigned to the first bin, and context decoding is performed on the first bin. In addition, in the second bin (binIdx=1), the index ctxInc=B1 is assigned and context decoding is performed. In addition, in the third bin (binIdx=2), the index ctxInc=B2 is assigned and context decoding is performed. In addition, in the fourth bin (binIdx=3), the index ctxInc=B3 is assigned and context decoding is performed.
[0704] Note that Fig.42 and Fig.43 In the table, non-overlapping unique values are set in indexes A0, A1, B1, B2 and B3.
[0705] The number of contexts, the number of context coding bins, and the number of bypass coding bins for each of these methods are similar to those used for encoding ( Fig.44 ).
[0706] As described above, in any case of Method 7-1 to Method 7-4, the number of contexts required for decoding can be reduced compared to the case of Method 0. That is, by applying Method 7, the number of contexts allocated to the first bin (binIdx=0) can be reduced. Therefore, an increase in memory usage can be suppressed.
[0707] Furthermore, in any of the cases of Method 7-1 to Method 7-3, the number of context coding bins required for decoding can be reduced compared to the case of Method 0. Note that in the case of Method 7-4, the number of context coding bins required for decoding is comparable to the number of context coding bins required for decoding in the case of Method 0. That is, by applying Method 7, bypass decoding can be applied to bins corresponding to transform types having relatively low selectivity. Therefore, it is possible to suppress an increase in the number of context coding bins while suppressing a decrease in coding efficiency and suppress an increase in processing volume (throughput).
[0708] Then, as described above, for the color difference format 4:2:2 or 4:4:4 having a larger amount of information than the color difference format 4:2:0, the application of the inverse adaptive orthogonal transform can be controlled. Therefore, the reduction in encoding efficiency can be suppressed.
[0709] As described above, by applying method 7, an increase in the load of the decoding process can be suppressed.
[0710] Note that the control parameters of transform skip or adaptive orthogonal transform may be signaled (decoded) for each treeType instead of color component ID (cIdx). That is, [cIdx] of each control parameter may be replaced with [treeType].
[0711] In addition, the above method of assigning a context variable to each bin in the bin sequence of the adaptive orthogonal transform identifier mts_idx can also be applied to other syntax elements related to orthogonal transform, etc. For example, the method can also be applied to the secondary transform identifier st_idx and the transform skip flag ts_flag.
[0712] <8-3. Encoding side>
[0713] <Configuration>
[0714] Next, the encoding side will be described. The configuration of the encoding side in this case is similar to that of the encoding side in the first embodiment. That is, the image encoding device 100 in this case has the same Figure 6In addition, the encoding unit 115 in this case has the same Figure 7 The configuration described is similar to the configuration.
[0715] <Encoding Process Flow>
[0716] In addition, the image encoding device 100 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image encoding processing performed by the image encoding device 100 in this case is performed by comparing the image encoding processing with the reference image encoding processing. Figure 8 The process described in the flowchart is similar to that described in the flowchart.
[0717] Will refer to Fig.45 The flowchart in describes an example of the flow of encoding processing for encoding an adaptive orthogonal transform identifier performed by the encoding unit 115 in this case.
[0718] In this encoding process, Fig. 9 The processing in steps S131 and S132 in 1 is similar to the processing in steps S551 and S552. That is, in this case, the selection unit 132 selects the context setting unit 133 as the supply destination of the bin (that is, the context encoding is selected.). For example, the selection unit 132 selects Fig.42 and Fig.43 Any one of the tables shown (ie, by applying any one of Methods 7-1 to 7-4) selects context coding as the encoding method for the bin.
[0719] In step S553, the context setting unit 133 assigns a context variable ctx (index ctxInc) to a bin according to whether the component is luminance (Y) ((cIdx==0)?).
[0720] For example, when the component is luma (Y) (cIdx==0), the context setting unit 133 assigns the index ctxInc=A0 to the bin. Also, when the component is not luma (Y) (cIdx>0), the context setting unit 133 assigns the index ctxInc=A1 to the bin.
[0721] Then, the context encoding unit 134 performs arithmetic encoding using the context variable. That is, context encoding is performed.
[0722] The processing in steps S554 to S558 is also the same as Fig. 9 The processing in steps S134 to S138 in is similarly performed. In step S558, in the case where it is determined to terminate the encoding, the encoding processing is terminated.
[0723] By performing each process as described above, the encoding unit 115 can encode the adaptive orthogonal transform identifier by applying method 7 (for example, any one of method 7-1 to method 7-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the encoding process can be suppressed.
[0724] <8-4. Decoding side>
[0725] <Configuration>
[0726] Next, the decoding side will be described. The configuration of the decoding side in this case is similar to that of the decoding side in the first embodiment. That is, the image decoding device 200 in this case has the same Fig.10 In addition, the decoding unit 212 in this case has the same Fig.11 The configuration described is similar to the configuration.
[0727] <Decoding Process Flow>
[0728] In addition, the image decoding device 200 in this case performs processing basically similar to that in the case of the first embodiment. That is, the image decoding processing performed by the image decoding device 200 in this case is performed by comparing with the reference Fig.12 The process described in the flowchart is similar to that described in the flowchart.
[0729] Will refer to Fig.46 The flowchart in describes an example of the flow of a decoding process for decoding the encoded data of the adaptive orthogonal transform identifier performed by the decoding unit 212 in this case.
[0730] In this decoding process, Fig.13 The processing of step S571 is performed similarly to the processing of step S231 of . That is, in this case, the selection unit 231 selects the context setting unit 232 as the supply destination of the bin (that is, selects the context decoding.). For example, the selection unit 231 selects Fig.42 and Fig.43 Any one of the tables shown (ie, by applying any one of Methods 7-1 to 7-4) selects context decoding as the decoding method for the bin.
[0731] In step S572, the context setting unit 232 assigns a context variable ctx (index ctxInc) to a bin according to whether the component is luminance (Y) ((cIdx==0)?).
[0732] For example, when the component is luma (Y) (cIdx==0), the context setting unit 232 assigns the index ctxInc=A0 to the bin. Also, when the component is not luma (Y) (cIdx>0), the context setting unit 232 assigns the index ctxInc=A1 to the bin.
[0733] Then, the context decoding unit 233 performs arithmetic decoding using the context variable, that is, performs context decoding.
[0734] The processing in steps S573 to S578 is also the same as Fig.13 The processing in steps S233 to S238 in is similarly performed. In step S578, in the case where it is determined to terminate the decoding, the decoding processing is terminated.
[0735] By performing each process in this manner, the decoding unit 212 can decode the encoded data of the adaptive orthogonal transform identifier by applying method 7 (for example, any one of method 7-1 to method 7-4). Therefore, an increase in memory usage can be suppressed. In addition, an increase in the amount of processing (throughput) can be suppressed. That is, an increase in the load of the decoding process can be suppressed.
[0736] <9. Appendix>
[0737] <Combination>
[0738] The present technology described in each of the above-mentioned embodiments can be applied in combination with the present technology described in any other embodiment as long as there is no contradiction.
[0739] <Computer>
[0740] The above series of processing can be performed by hardware or by software. In the case of performing this series of processing by software, a program configuring the software is installed in the computer. Here, the computer includes a computer incorporated in dedicated hardware, a computer capable of performing various functions by installing various programs, such as a general-purpose personal computer, and the like.
[0741] Fig.47 : is a block diagram showing a configuration example of hardware of a computer that executes the above-described series of processes by a program.
[0742] exist Fig.47 In the computer 800 shown in FIG. 8 , a central processing unit (CPU) 801 , a read only memory (ROM) 802 , and a random access memory (RAM) 803 are connected to each other via a bus 804 .
[0743] An input / output interface 810 is also connected to the bus 804. An input unit 811, an output unit 812, a storage unit 813, a communication unit 814, and a drive 815 are connected to the input / output interface 810.
[0744] The input unit 811 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 812 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 813 includes, for example, a hard disk, a RAM disk, and a nonvolatile memory, etc. The communication unit 814 includes, for example, a network interface. The drive 815 drives a removable medium 821, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0745] In the computer configured as described above, the CPU 801 loads the program stored in the storage unit 813 into the RAM 803 via the input / output interface 810 and the bus 804, for example, and executes the program, so that the above-mentioned series of processing is performed. In addition, the RAM 803 appropriately stores data and the like required for the CPU 801 to perform various types of processing.
[0746] For example, the program to be executed by the computer may be recorded and applied on the removable medium 821 or the like as a package medium and may be provided. In this case, by attaching the removable medium 821 to the drive 815 , the program may be installed to the storage unit 813 via the input / output interface 810 .
[0747] In addition, the program may be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In this case, the program may be received by the communication unit 814 and installed in the storage unit 813.
[0748] In addition to the above-described method, the program may be pre-installed in the ROM 802 or the storage unit 813 .
[0749] <Units of Information and Processing>
[0750] The data units in which the above-mentioned various types of information are set and the data units to be processed by various types of processing are arbitrary and not limited to the above-mentioned examples. For example, this information and processing can be set for each transform unit (TU), transform block (TB), prediction unit (PU), prediction block (PB), coding unit (CU), maximum coding unit (LCU), sub-block, block, tile, slice, picture, sequence or component, or the data in these data units can be used. Of course, the data unit can be set for each information and processing, and there is no need to unify the data units for all information and processing. Note that the storage location of this information is arbitrary and can be stored in the header, parameters, etc. of the above-mentioned data unit. In addition, the information can be stored in multiple locations.
[0751] <Control Information>
[0752] The control information about the present technology described in the above embodiments may be sent from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) for controlling whether to allow (or prohibit) the application of the above-mentioned present technology may be sent. In addition, for example, control information (e.g., present_flag) indicating an object to which the above-mentioned present technology is applied (or an object to which the present technology is not applied) may be sent. For example, control information for specifying a block size (upper limit, lower limit, or both), a frame, a component, a layer, etc. to which the present technology is applied (or to which the application is allowed or prohibited) may be sent.
[0753] <Target audience of this technology>
[0754] The present technology can be applied to any image encoding / decoding method. That is, the specifications of various types of processing such as transformation (inverse transformation), quantization (inverse quantization), encoding (decoding) and prediction regarding image encoding / decoding are arbitrary and not limited to the above examples, as long as they do not conflict with the above-mentioned present technology. In addition, as long as they do not conflict with the above-mentioned present technology, part of the processing can be omitted.
[0755] In addition, the present technology can be applied to a multi-view image encoding / decoding system that performs encoding / decoding of multi-view images including images of multiple viewpoints (views). In this case, the present technology is simply applied to encoding / decoding of each viewpoint (view).
[0756] In addition, the present technology can be applied to a layered image coding (scalable coding) / decoding system that encodes / decodes a layered image of multiple layers (layers) to have a scalability function for a predetermined parameter. In this case, the present technology is simply applied to the coding / decoding of each layer (layer).
[0757] Furthermore, in the above description, the image encoding device 100 and the image decoding device 200 have been described as application examples of the present technology, but the present technology can be applied to an arbitrary configuration.
[0758] The present technology can be applied to, for example, various electronic devices, such as transmitters and receivers in satellite broadcasting (e.g., television receivers and mobile phones), cable broadcasting such as cable television, distribution on the Internet, and distribution to terminals via cellular communications, or devices that record images on media such as optical disks, magnetic disks, and flash memories and reproduce images from these storage media (e.g., hard disk recorders and camera devices).
[0759] In addition, for example, the present technology can be implemented as a configuration of a part of a device, such as a processor as a system large-scale integration (LSI) or the like (e.g., a video processor), a module using multiple processors or the like (e.g., a video module), a unit using multiple modules or the like (e.g., a video unit), or a set in which other functions are added to the unit (e.g., a video set) (i.e., a configuration of a part of a device).
[0760] In addition, for example, the present technology can also be applied to a network system including multiple devices. For example, the present technology can be implemented as cloud computing shared and collaboratively processed by multiple devices via a network. For example, the present technology can be implemented as a cloud service that provides services related to images (moving images) to any terminal such as a computer, an audio-visual (AV) device, a portable information processing terminal, or an Internet of Things (IoT) device.
[0761] Note that in this specification, the term "system" means a collection of multiple configuration elements (devices, modules (components), etc.), and it does not matter whether all the configuration elements are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network and one device housing multiple modules in one housing are both systems.
[0762] <Fields and applications where this technology can be applied>
[0763] Systems, devices, processing units, etc. to which this technology is applied can be used in any field such as transportation, medical care, crime prevention, agriculture, animal husbandry, mining, beauty, factories, home appliances, weather and nature monitoring. In addition, the use in any field is also arbitrary.
[0764] For example, the present technology can be applied to systems and devices provided for providing content for appreciation, etc. In addition, for example, the present technology can also be applied to systems and devices for transportation such as traffic condition monitoring and automatic driving control. In addition, for example, the present technology can also be applied to systems and devices provided for safety. In addition, for example, the present technology can also be applied to systems and devices provided for automatic control of machines, etc. In addition, for example, the present technology can also be applied to systems and devices provided for agriculture or animal husbandry. In addition, the present technology can also be applied to systems and devices for monitoring natural conditions such as volcanoes, forests and oceans, wild animals, etc. In addition, for example, the present technology can also be applied to systems and devices provided for sports.
[0765] <Others>
[0766] Note that the "flag" in this specification is information for identifying multiple states, and includes not only information for identifying the two states of true (1) and false (0), but also information that can identify three or more states. Therefore, the value that the "flag" can take can be, for example, a binary value 1 / 0 or can be a ternary value or more binary values. That is, the number of bits that constitute the "flag" is arbitrary and can be 1 bit or more. In addition, it is assumed that the identification information (including the flag) is not only in the form of including the identification information in the bit stream, but also in the form of including the difference information between the identification information and the specific reference information in the bit stream. Therefore, in this specification, the "flag" and "identification information" include not only the information itself, but also the difference information relative to the reference information.
[0767] In addition, various types of information (metadata, etc.) about the coded data (bitstream) can be sent or recorded in any form, as long as the various types of information are associated with the coded data. Here, the term "association" means that other data can be used (linked) when processing one data, for example. That is, data associated with each other can be collected as one data or can be separate data. For example, information associated with the coded data (image) can be sent on a transmission path different from the transmission path of the coded data (image). In addition, for example, information associated with the coded data (image) can be recorded on a recording medium different from the coded data (image) (or another recording area of the same recording medium). Note that the "association" can be a part of the data, not the entire data. For example, an image and information corresponding to the image can be associated with each other in arbitrary units such as multiple frames, one frame, or a part of a frame.
[0768] Note that in this specification, terms such as "combine", "multiplex", "add", "integrate", "include", "store", and "insert" mean putting multiple things into one thing, such as putting encoded data and metadata into one data, and mean a method of the above-mentioned "association".
[0769] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications may be made without departing from the gist of the present technology.
[0770] For example, a configuration described as one device (or processing unit) may be divided and configured into a plurality of devices (or processing units). Conversely, a configuration described as a plurality of devices (or processing units) may be configured together as one device (or processing unit). Furthermore, a configuration other than the above configuration may be added to the configuration of each device (or each processing unit). Furthermore, a portion of the configuration of a specific device (or processing unit) may be included in the configuration of another device (or another processing unit), as long as the configuration and operation of the system as a whole are substantially the same.
[0771] In addition, for example, the above-mentioned program can be executed by any device. In this case, the device only needs to have necessary functions (functional blocks, etc.) and obtain necessary information.
[0772] In addition, for example, each step of a flowchart may be performed by one device, or may be shared and performed by multiple devices. In addition, in the case where multiple processes are included in one step, the multiple processes may be performed by one device, or may be shared and performed by multiple devices. In other words, the multiple processes included in one step may be performed as processes of multiple steps. Conversely, the processes described as multiple steps may be performed collectively as one step.
[0773] Note that in a program executed by a computer, the processing of the steps describing the program may be performed in chronological order according to the order described in this specification, or may be performed separately in parallel or at a necessary time such as when a call is made. That is, the processing of each step may be performed in an order different from the above order as long as no contradiction occurs. In addition, the processing of the steps describing the program may be performed in parallel with the processing of another program, or may be performed in combination with the processing of another program.
[0774] In addition, for example, as long as there is no contradiction, multiple technologies related to the present technology can be implemented independently as a whole. Of course, any number of the present technologies can be implemented together. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. In addition, part or all of any of the above-mentioned present technologies can be implemented in combination with another technology not described above.
[0775] Note that the present technology may also have the following configurations.
[0776] (1) An image processing device comprising:
[0777] An encoding unit is configured to assign a predetermined context variable to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding, and to perform context encoding on the first bin in the bin sequence.
[0778] (2) The image processing device according to (1), wherein:
[0779] The encoding unit assigns predetermined context variables different from each other to first to fourth bins in the bin sequence, and performs context encoding on the first to fourth bins in the bin sequence.
[0780] (3) The image processing device according to (1), wherein:
[0781] The encoding unit assigns predetermined context variables different from each other to first to third bins in the bin sequence, performs context encoding on the first to third bins in the bin sequence, and performs bypass encoding on a fourth bin in the bin sequence.
[0782] (4) The image processing device according to (1), wherein:
[0783] The encoding unit assigns predetermined context variables different from each other to a first bin and a second bin in the bin sequence, performs context encoding on the first bin and the second bin in the bin sequence, and performs bypass encoding on a third bin and a fourth bin in the bin sequence.
[0784] (5) The image processing device according to (1), wherein:
[0785] The encoding unit assigns a predetermined context variable to a first bin in the bin sequence, performs context encoding on the first bin in the bin sequence, and performs bypass encoding on second to fourth bins in the bin sequence.
[0786] (6) The image processing device according to any one of (1) to (5), wherein:
[0787] The encoding unit binarizes the adaptive orthogonal transform identifier into a bin sequence configured by one bit indicating whether a transform type is different from the transform type DCT2×DCT2 and two bits indicating other transform types, and encodes the bin sequence.
[0788] (7) An image processing method, comprising:
[0789] A predetermined context variable is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0790] (8) An image processing device comprising:
[0791] The encoding unit is configured to assign a context variable based on a parameter regarding a block size to a first bin in a bin sequence, and to perform context encoding on the first bin in the bin sequence, wherein the bin sequence is obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding.
[0792] (9) The image processing device according to (8), wherein:
[0793] The parameter regarding the block size is a difference between a logarithmic value of a long side of a transform block and a logarithmic value of a minimum transform block size to which the adaptive orthogonal transform is applicable.
[0794] (10) The image processing device according to (8), wherein:
[0795] The parameter regarding the block size is a minimum value between a difference between a logarithmic value of a long side of a transform block and a logarithmic value of a minimum transform block size to which adaptive orthogonal transform is applicable and a predetermined threshold.
[0796] (11) The image processing device according to (8), wherein:
[0797] The parameter about the block size is a result of right-shifting the minimum value between the difference between the logarithmic value of the long side of the transform block and the logarithmic value of the minimum transform block size to which the adaptive orthogonal transform is applicable and a predetermined threshold.
[0798] (12) The image processing device according to (8), wherein:
[0799] The encoding unit allocates a context variable according to whether the parameter regarding the block size is equal to or greater than a predetermined threshold and performs context encoding.
[0800] (13) The image processing device according to any one of (8) to (12), wherein:
[0801] The encoding unit binarizes the adaptive orthogonal transform identifier into a bin sequence configured by one bit indicating whether a transform type is different from the transform type DCT2×DCT2 and two bits indicating other transform types, and encodes the bin sequence.
[0802] (14) The image processing device according to (8), wherein:
[0803] The encoding unit binarizes the adaptive orthogonal transform identifier for each component and performs encoding.
[0804] (15) An image processing method, comprising:
[0805] A context variable based on a parameter about a block size is assigned to a first bin in a bin sequence obtained by binarizing an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding, and context encoding is performed on the first bin in the bin sequence.
[0806] Reference numerals list
[0807] 100 Image encoding device
[0808] 115 Coding Unit
[0809] 131 Binarization Unit
[0810] 132 Select Unit
[0811] 133 Context Setting Unit
[0812] 134 Context Coding Unit
[0813] 135 Bypass Encoding Unit
[0814] 200 Image decoding device
[0815] 212 Decoding Unit
[0816] 231 Select Unit
[0817] 232 Context Setting Unit
[0818] 233 Context decoding unit
[0819] 234 Bypass decoding unit
[0820] 235 Debinarization Unit
Claims
1. An image encoding device, comprising: an encoding unit configured to: assign a plurality of predetermined context variables different from each other to each bin in the bin sequence, wherein an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding is encoded using at least four bins, and The encoding unit is configured to select the context variable based on a predefined set of rules.
2. An image encoding method, comprising: assigning a plurality of predetermined context variables different from each other to each bin in the bin sequence, wherein an adaptive orthogonal transform identifier indicating a mode of an adaptive orthogonal transform in image encoding is encoded using at least four bins, and The image encoding method comprises selecting the context variable based on a predefined set of rules.
3. An image decoding device, comprising: a decoding unit configured to perform context decoding on each bin in the bin sequence using a plurality of predetermined context variables assigned to each bin in the bin sequence and different from each other, wherein an adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding is encoded using at least four bins, and Therein, the context variables have been selected based on a predefined set of rules.
4. An image decoding method, comprising: performing context decoding on each bin in the bin sequence using a plurality of predetermined context variables assigned to each bin in the bin sequence that are different from each other, wherein an adaptive orthogonal transform identifier indicating a mode of inverse adaptive orthogonal transform in image decoding is encoded using at least four bins, and Therein, the context variables have been selected based on a predefined set of rules.
Citation Information
Patent Citations
Image processing apparatus and method
US20190104322A1