Image processing apparatus and method
By setting a fixed context variable for the first bit of secondary transform control information, the encoding and decoding processes are optimized, reducing load and memory usage in image processing.
Patent Information
- Application Number
- JP2024074323
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-06
- Filing Date
- 2024-05-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-09-04
AI Technical Summary
The redundant condition 'tu_mts_idx == 0' in the branch of the encoding process for secondary transform in VVC WD6 leads to an increase in encoding and decoding load.
Setting a fixed context variable for the first bit of the secondary transform control information and arithmetically encoding/decoding it, thereby simplifying the derivation of context variables and reducing memory usage.
This approach reduces the load of the encoding and decoding processes while maintaining efficiency by simplifying the derivation of context variables and minimizing memory requirements.
Smart Images

Figure 0007704252000001 
Figure 0007704252000002 
Figure 0007704252000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing apparatus and method, and more particularly, to an image processing apparatus and method capable of suppressing an increase in the load of encoding processing and decoding processing.
Background Art
[0002] Conventionally, an encoding method has been proposed in which a prediction residual of a moving image is derived, coefficient-converted, quantized, and encoded (for example, Non-Patent Document 1). In VVC WD6 described in Non-Patent Document 1, a low-frequency secondary transform (LFST (Low Frequency Secondary Transform)) is performed on the transform coefficients after primary transform, and there is an encoding tool for further improving energy compaction. As mode information regarding this low-frequency secondary transform, there is an ST identifier st_idx (also referred to as lfnst_idx).
[0003] When encoding this ST identifier lfnst_idx, first, this ST identifier lfnst_idx is binarized, and arithmetic coding is performed with reference to a context variable ctx corresponding to each binIdx of the obtained bin sequence (bins). For example, for the bin with binIdx = 0 in this bin sequence, a context variable ctx is assigned according to the tree type treType and the MTS identifier mts_idx.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
[0005] However, since LFSNT is only applied when tu_mts_idx == 0 (DCT2xDCT2), the condition "tu_mts_idx == 0" is always true. Therefore, the branch of "tu_mts_idx==0" is redundant, which may increase the load of the encoding process and the decoding process.
[0006] The present disclosure has been made in view of such a situation, and aims to suppress an increase in the load of the encoding process and the decoding process. Means for Solving the Problems
[0007] An image processing apparatus according to an aspect of the present technology includes an arithmetic coder that refers to a context variable for secondary transform control information, which is control information related to secondary transform, for secondary transform control information in a bin string. corresponding to of the bin string the second bin for preset as one fixed value referring to a context variable for the secondary transform control information the second bin and arithmetically encoding encoding unit the secondary transform control information.
[0008] An image processing method according to an aspect of the present technology refers to a context variable for secondary transform control information, which is control information related to secondary transform, for secondary transform control information in a bin string. corresponding to of the bin string the second bin for preset as one fixed value referring to a context variable for the secondary transform control information the second binIt is an image processing method for arithmetic coding.
[0009] An image processing apparatus according to another aspect of the present technology includes a decoding unit that arithmetic-decodes the secondary conversion control information, which is control information related to secondary conversion corresponding to for the the second bin bin sequence preset in advance as one fixed value with reference to the set context variable. the second bin It is an image processing apparatus including a decoding unit that arithmetic-decodes the secondary conversion control information.
[0010] An image processing method according to another aspect of the present technology is an image processing method that arithmetic-decodes the secondary conversion control information, which is control information related to secondary conversion corresponding to for the the second bin bin sequence preset in advance as one fixed value with reference to the set context variable. the second bin It is an image processing method that arithmetic-decodes the secondary conversion control information.
[0011] In an image processing apparatus and method according to one aspect of the present technology, the secondary conversion control information, which is control information related to secondary conversion corresponding to for the the second bin bin sequence preset in advance as one fixed value is referred to by the set context variable, and the second bin the secondary conversion control information is arithmetic-coded.
[0012] In an image processing apparatus and method according to another aspect of the present technology, the secondary conversion control information, which is control information related to secondary conversion corresponding to for the the second bin bin sequence preset in advance as one fixed value is referred to by the set context variable, and the second bin the secondary conversion control information is arithmetic-decoded.
Brief Description of Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0014] Hereinafter, a mode for carrying out the present disclosure (hereinafter referred to as the embodiment) will be described. The description will be made in the following order. 1. Encoding of ST Identifier 2. First Embodiment (Symbolization Device) 3. Second Embodiment (Decoding Device) 4. Third Embodiment (Image Decoding Device) 5. Fourth Embodiment (Image Symbolization Device) 6. Supplementary Note
[0015] <1. Encoding of ST Identifier> <Literature etc. Supporting Technical Content and Technical Terms> The scope disclosed in this technology includes not only the content described in the embodiment, but also the content described in the following non-patent documents etc. that were publicly known at the time of filing, and the content of other documents referred to in the following non-patent documents.
[0016] Non-Patent Document 1: (above-mentioned) Non-Patent Document 2: Jianle Chen, Yan Ye, Seung Hwan Kim, "Algorithm description for Versatile Video Coding and Test Model 6 (VTM 6)", JVET-O2002-v2, m49914, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 15th Meeting: Gothenburg, SE, 3-12 July 2019 Non-Patent Document 3: Recommendation ITU-T H.264 (04 / 2017) "Advanced video coding for generic audiovisual services", April 2017 Non-Patent Document 4: Recommendation ITU-T H.265 (02 / 18) "High efficiency video coding", february 2018
[0017] That is, the content described in the above non-patent documents also serves as a basis for determining the support requirements. For example, even if there is no direct description in the embodiments of the Quad-Tree Block Structure and QTBT (Quad Tree Plus Binary Tree) Block Structure described in the above non-patent documents, they are within the disclosure scope of the present technology and are considered to meet the support requirements of the claims. Also, for example, with regard to technical terms such as Parsing, Syntax, and Semantics, even if there is no direct description in the embodiments, they are within the disclosure scope of the present technology and are considered to meet the support requirements of the claims.
[0018] In addition, in this specification, the "block" (not the block indicating a processing unit) used in the description as a partial region or a processing unit of an image (picture) indicates an arbitrary partial region within the picture, and its size, shape, characteristics, etc. are not limited. For example, the "block" includes any partial region (processing unit) such as TB (Transform Block), TU (Transform Unit), PB (Prediction Block), PU (Prediction Unit), SCU (Smallest Coding Unit), CU (Coding Unit), LCU (Largest Coding Unit), CTB (Coding Tree Block), CTU (Coding Tree Unit), sub-block, macro-block, tile, or slice described in the above-mentioned non-patent documents.
[0019] In addition, when specifying the size of such a block, not only directly specify the block size, but also indirectly specify the block size. For example, the block size may be specified using identification information for identifying the size. Also, for example, the block size may be specified by the ratio or difference from the size of a reference block (such as LCU or SCU). For example, when transmitting information specifying the block size as a syntax element, such indirectly specified size information as described above may be used as the information. By doing so, the amount of information of the information can be reduced, and the coding efficiency may be improved in some cases. In addition, the specification of the block size includes the specification of the range of the block size (for example, the specification of the allowable range of the block size, etc.).
[0020] Also, in this specification, encoding includes not only the entire process of converting an image into a bitstream but also some partial processes. For example, it includes processes such as prediction processing, orthogonal transformation, quantization, arithmetic coding, etc., and not only includes processes that collectively refer to quantization and arithmetic coding, processes that include prediction processing, quantization, and arithmetic coding, etc. Similarly, decoding includes not only the entire process of converting a bitstream into an image but also some partial processes. For example, it includes processes such as inverse arithmetic decoding, inverse quantization, inverse orthogonal transformation, prediction processing, etc., and not only includes processes that include inverse arithmetic decoding and inverse quantization, processes that include inverse arithmetic decoding, inverse quantization, and prediction processing, etc.
[0021] <ST identifier> In the VVC WD6 described in Non-Patent Document 1, for the transform coefficients after the primary transformation, there is an encoding tool that performs low-frequency secondary transform (LFST (Low Frequency Secondary Transform)) to further improve energy compaction. As mode information (control information related to secondary transform) regarding this low-frequency secondary transform, there is an ST identifier st_idx (also referred to as lfnst_idx) that indicates the type of secondary transform. The encoding of this ST identifier is performed according to the following procedure.
[0022] First, the ST identifier lfnst_idx is binarized to obtain a bit sequence bins. The ST identifier lfnst_idx takes values as shown in the table of A in FIG. 1, and each value indicates a transform type as shown in the same table. Also, by such binarization of the ST identifier lfnst_idx, a bit sequence (binarization) as shown in the same table is obtained.
[0023] Note that the binarization of this ST identifier lfnst_idx is binarized by truncated unary code (TU (Truncated Unary)) to obtain a bit sequence bins. Note that this TU code is equivalent to a truncated Rice code (TR (Truncated Rice)) with a Rice parameter cRiceParam = 0 as shown in the table in B of FIG. 1.
[0024] Next, arithmetic coding is performed by referring to the context variable ctx corresponding to each binIdx of the obtained bin sequence bins. The index for identifying the context variable ctx is also referred to as ctxInc (or ctxIdx).
[0025] Specifically, for the bin with binIdx = 0 in the bin sequence, as shown in the table indicated by C in FIG. 1, a context variable ctx is assigned according to the tree type treType and the MTS identifier mts_idx. In the example of this table, ctxInc corresponding to binIdx = 0 is derived by the value of ctxInc = (mts_idx == 0 && treeType != SINGLE_TREE). Also, for binIdx = 1 in the bin sequence bins, bypass coding is applied.
[0026] Note that each bin of the bin sequence bins of the ST identifier lfnst_idx can also be interpreted as a flag corresponding to the conversion type. In that case, the value of the bin with binIdx = 0 corresponds to a flag indicating the availability of secondary conversion (0: Yes, 1: No), and the value of the bin with binIdx = 1 corresponds to a flag indicating whether it is the first secondary conversion (0: Yes, 1: No).
[0027] During decoding, each bin is arithmetically decoded by referring to the context variable ctx set as described above. Then, the bin sequence is made multi-valued, and the ST identifier lfnst_idx is obtained.
[0028] However, since LFSNT is only applied when tu_mts_idx == 0 (DCT2xDCT2), in the table shown by C in FIG. 1, the condition "tu_mts_idx == 0" is always true. Therefore, the branch of "tu_mts_idx == 0" is redundant, which may increase the load of the encoding process and the decoding process. In order to suppress the implementation cost, it is required to reduce the context variables and simplify the derivation of the context variables while maintaining the encoding efficiency.
[0029] <Assignment of Fixed Contexts> Therefore, a fixed context variable is set for the first bit of the bit string of the secondary conversion control information, which is control information related to the secondary conversion.
[0030] For example, in image processing (particularly, image encoding), a fixed context variable is set for the first bit of the bit string of the secondary conversion control information, which is control information related to the secondary conversion, and each bit of the secondary conversion control information is arithmetically encoded with reference to the set fixed context variable.
[0031] For example, in an image processing apparatus (particularly, an image encoding apparatus that encodes an image), a fixed context variable is set for the first bit of the bit string of the secondary conversion control information, which is control information related to the secondary conversion, and an encoding unit that arithmetically encodes each bit of the secondary conversion control information with reference to the set fixed context variable is provided.
[0032] By doing so, it is possible to simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, it is possible to reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, an increase in the load of the encoding process can be suppressed.
[0033] Also, for example, in image processing (decoding of encoded image data), arithmetic decoding is performed with reference to the fixed context variable set for the first bit of the bit string of the secondary conversion control information, which is control information related to the secondary conversion.
[0034] For example, in an image processing apparatus (particularly, an image decoding apparatus that decodes encoded image data), a decoding unit that performs arithmetic decoding with reference to the fixed context variable set for the first bit of the bit string of the secondary conversion control information, which is control information related to the secondary conversion, is provided.
[0035] By doing so, it is possible to simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, while maintaining the encoding efficiency, it is possible to reduce the memory size for holding the context variable of the ST identifier. That is, an increase in the load of the decoding process can be suppressed.
[0036] For example, as in the second row (Method 1) from the top of the table shown in A of FIG. 2, for the bin with binIdx = 0 (the first bin) in the bin sequence, “0” may be assigned as a fixed context variable ctx. That is, “0” may be set as the context variable ctx for the first bin of the bin sequence of the ST identifier, and arithmetic coding may be performed with reference to that value. In other words, arithmetic decoding may be performed with reference to “0” set as the context variable ctx for the first bin of the bin sequence of the ST identifier. By doing so, in the derivation of the context variable in the first bin of the bin sequence of the ST identifier, the branching process related to the selection of the context variable by treeType (and tu_mts_idx == 0) can be omitted. Thereby, an increase in the load of the encoding process and the decoding process can be suppressed. Also, as shown in the table in B of FIG. 2, compared with the case of VTM, Method 1 can reduce the number of contexts. Therefore, the memory size for holding the context variable of the ST identifier can be reduced.
[0037] Also, as in the table (Method 1) shown in A of FIG. 2, a bypass flag may be assigned to the bin with binIdx = 1 (the second bin) in the bin sequence. That is, a bypass flag (bypass) may be set for the first bin of the bin sequence of the ST identifier, and arithmetic coding may be performed based on that bypass flag. In other words, arithmetic decoding may be performed based on the bypass flag set for the first bin of the bin sequence of the ST identifier.
[0038] Also, for example, as in the third row from the top in the table shown at A in FIG. 2 (Method 2), a common context variable ctx may be assigned to the bin with binIdx = 0 (the first bin) and the bin with binIdx = 1 (the second bin) in the bin column. In the case of the table shown at A in FIG. 2, “0” is assigned to both the bin with binIdx = 0 (the first bin) and the bin with binIdx = 1 (the second bin) in the bin column. That is, a common context variable ctx (for example, “0”) may be set for the first and second bins in the bin column of the ST identifier, and arithmetic coding of each bin may be performed with reference to its value. In other words, arithmetic decoding of each bin may be performed with reference to the common context variable ctx (for example, “0”) set for the first and second bins in the bin column of the ST identifier. By doing so, as in the table shown at B in FIG. 2, in Method 2, the number of bypass-coded bins can be reduced while maintaining the number of contexts compared to the case of Method 1. Therefore, the amount of code for the ST identifier can be reduced.
[0039] <2. First Embodiment> <Encoding Device> The technology described in <1. Encoding of ST Identifier> can be applied to any device. Application examples of this technology will be described below. FIG. 3 is a block diagram showing an example of the configuration of an encoding device which is an aspect of an image processing device to which this technology is applied. The encoding device 100 shown in FIG. 3 is a device that encodes the ST identifier lfnst_idx. The encoding device 100 performs encoding by applying, for example, CABAC (Context-based Adaptive Binary Arithmetic Code).
[0040] Note that in FIG. 3, main components such as processing units and data flows are shown, and not all of the components shown in FIG. 3 are necessarily included. That is, in the encoding device 100, there may be a processing unit not shown as a block in FIG. 3, or there may be a process or data flow not shown as an arrow or the like in FIG. 3.
[0041] As shown in FIG. 3, the encoding device 100 includes a binarization unit 121, a selection unit 122, a context model 123, an arithmetic encoding unit 124, an arithmetic encoding unit 125, and a selection unit 126.
[0042] The binarization unit 121 acquires the syntax element value supplied to the encoding device 100, binarizes it in a method defined for each syntax element, and generates a binarized bit sequence. The binarization unit 121 supplies the binarized bit sequence to the selection unit 122.
[0043] The selection unit 122 acquires the binarized bit sequence supplied from the binarization unit 121 and the flag information isBypass. The selection unit 122 selects the supply destination of the binarized bit sequence based on the value of isBypass. For example, when isBypass = 0, the selection unit 122 determines that it is in the normal mode and supplies the binarized bit sequence to the context model 123. Also, when isBypass = 1, the selection unit 122 determines that it is in the bypass mode and supplies the binarized bit sequence to the arithmetic encoding unit 125.
[0044] The context model 123 dynamically switches the context model to be applied according to the encoding target and the surrounding situation. For example, the context model 123 holds the context variable ctx, and when it acquires the binarized bit sequence from the selection unit 122, it reads out the context variable ctx corresponding to each bin position (binIdx) of the bin sequence defined for each syntax element. The context model 123 supplies the binarized bit sequence and the read context variable ctx to the arithmetic encoding unit 124.
[0045] When the arithmetic coding unit 124 acquires the binarized bit sequence and the context variable ctx supplied from the context model 123, it refers to the probability state of the context variable ctx and arithmetically codes (performs context coding) the value of the bin at binIdx of the binarized bit sequence in the normal mode of CABAC. The arithmetic coding unit 124 supplies the coded data generated by the context coding to the selection unit 126. Further, the arithmetic coding unit 124 supplies the context variable ctx after the context coding process to the context model 123 and causes it to be held.
[0046] The arithmetic coding unit 125 arithmetically codes (performs bypass coding) the binarized bit sequence supplied from the selection unit 122 in the bypass mode of CABAC. The arithmetic coding unit 125 supplies the coded data generated by the bypass coding to the selection unit 126.
[0047] The selection unit 126 acquires the flag information isBypass and selects the coded data to be output based on the value of the isBypass. For example, when isBypass = 0, the selection unit 126 determines that it is the normal mode, acquires the coded data supplied from the arithmetic coding unit 124, and outputs it to the outside of the CABAC 102 (coding device 100). Also, when isBypass = 1, the selection unit 126 determines that it is the bypass mode, acquires the coded data supplied from the arithmetic coding unit 125, and outputs it to the outside of the coding device 100.
[0048] Note that these processing units (the binarization unit 121 to the selection unit 126) may have any configuration. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Further, each processing unit may have, for example, a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc., and realize the above-described processing by executing a program using them. Of course, each processing unit may have both of these configurations, and realize a part of the above-described processing by a logic circuit and the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, some of the processing units may realize a part of the above-described processing by a logic circuit, some other processing units may realize the above-described processing by executing a program, and still other processing units may realize the above-described processing by both a logic circuit and program execution.
[0049] In such an encoding device 100, the present technology described in <1. Encoding of ST Identifier> is applied. That is, the binarization unit 121 binarizes the secondary conversion control information (ST identifier lfnst_idx) to generate a bit string. The selection unit 122 sets a fixed context variable ctx (for example, "0") for the first bit of the bit string. The arithmetic encoding unit 124 refers to the fixed context variable ctx (for example, "0") and performs arithmetic encoding in the normal mode of CABAC.
[0050] By doing so, the encoding device 100 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Further, the encoding device 100 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the encoding device 100 can suppress an increase in the load of the encoding process.
[0051] At that time, the selection unit 122 may set a bypass flag for the second bin of the bin sequence (Method 1). In that case, the arithmetic coding unit 125 arithmetically codes the second bin in the bypass mode of CABAC.
[0052] Alternatively, the selection unit 122 may set a common context variable ctx (e.g., "0") for the second bin of the bin sequence with the first bin (Method 2). In that case, the arithmetic coding unit 124 refers to the fixed context variable ctx (e.g., "0") and arithmetically codes in the normal mode of CABAC.
[0053] <Flow of Encoding Process> Next, an example of the flow of the encoding process executed by this encoding device 100 will be described with reference to the flowchart of FIG. 4.
[0054] When the encoding process is started, the binarization unit 121 of the encoding device 100 inputs a syntax element value (syncVal) to be processed in step S101.
[0055] In step S102, the binarization unit 121 performs binarization processing defined for each syntax element and derives a bin sequence (synBins) of the syntax element value (syncVal).
[0056] In step S103, the selection unit 122 reads out a context variable ctx corresponding to each bin position (binIdx) of the bin sequence defined for each syntax element and a flag isBypass indicating whether it is in the bypass mode.
[0057] In step S104, the selection unit 122 determines whether it is in the bypass mode. If isBypass = 0 and it is determined that it is in the normal mode, the process proceeds to step S105.
[0058] In step S105, the arithmetic coding unit 124 performs context coding. That is, the arithmetic coding unit 124 refers to the probability state of the context variable ctx and codes the value of the bin at the bin position (binIdx) of the bin sequence (synBins) in the normal mode of CABAC. When the process of step S105 ends, the process proceeds to step S107.
[0059] Also, in step S104, if isBypass = 1 and it is determined that the bypass mode is enabled, the process proceeds to step S106.
[0060] In step S106, the arithmetic coding unit 125 performs bypass coding. That is, the arithmetic coding unit 125 codes the value of the bin at the bin position (binIdx) of the bin sequence (synBins) in the bypass mode of CABAC. When the process of step S106 ends, the process proceeds to step S107.
[0061] In step S107, the selection unit 126 determines whether a predetermined break condition A is satisfied. The break condition A is defined based on the values of the bin sequence from binIdx = 0 to binIdx = k (the position k of the current binIdx) and the binarization method for each syntax element.
[0062] If it is determined that the break condition A is not satisfied, the process returns to step S103, and the subsequent process is executed for the next bin position (binIdx). That is, the processes of steps S103 to S107 are executed for each bin position (binIdx).
[0063] Then, in step S107, if it is determined that the break condition A is satisfied, the CABAC process ends.
[0064] In such an encoding process, the present technology described in <1. Encoding of ST Identifier> is applied. That is, in step S102, the binarization unit 121 binarizes the secondary conversion control information (ST identifier lfnst_idx) to generate a bit string. In step S103, the selection unit 122 sets a fixed context variable ctx (for example, "0") for the first bit of the bit string. In step S105, the arithmetic coding unit 124 refers to the fixed context variable ctx (for example, "0") and performs arithmetic coding in the normal mode of CABAC.
[0065] By doing so, the encoding device 100 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the encoding device 100 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the encoding device 100 can suppress an increase in the load of the encoding process.
[0066] Also, in step S103, the selection unit 122 may set a bypass flag for the second bit of the bit string (Method 1). In that case, in step S106, the arithmetic coding unit 125 performs arithmetic coding on the second bit in the bypass mode of CABAC.
[0067] Also, in step S103, the selection unit 122 may set a context variable ctx (for example, "0") common to the first bit for the second bit of the bit string (Method 2). In that case, in step S105, the arithmetic coding unit 124 refers to the fixed context variable ctx (for example, "0") and performs arithmetic coding in the normal mode of CABAC.
[0068] <3. Second Embodiment> <Decoder> FIG. 5 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an image processing apparatus to which the present technology is applied. The decoding device 200 shown in FIG. 5 is a device that decodes the encoded data of the ST identifier lfnst_idx. The decoding device 200 performs decoding by applying a decoding method (e.g., CABAC) corresponding to the encoding method of the encoding device 100. For example, the decoding device 200 decodes the encoded data of the ST identifier lfnst_idx generated by the encoding device 100.
[0069] Note that in FIG. 5, main components such as the processing unit and the data flow are shown, and what is shown in FIG. 5 is not necessarily all. That is, in the decoding device 200, there may be a processing unit that is not shown as a block in FIG. 5, or there may be a processing or data flow that is not shown as an arrow or the like in FIG. 5.
[0070] As shown in FIG. 5, the decoding device 200 includes a selection unit 221, a context model 222, an arithmetic decoding unit 223, an arithmetic decoding unit 224, a selection unit 225, and a quantization unit 226.
[0071] The selection unit 221 acquires the encoded data input to the decoding device 200 and the flag information isBypass. The selection unit 221 selects the supply destination of the encoded data based on the value of the isBypass. For example, when isBypass = 0, the selection unit 221 determines that it is in the normal mode and supplies the encoded data to the context model 222. When isBypass = 1, the selection unit 221 determines that it is in the bypass mode and supplies the binary bit sequence to the arithmetic decoding unit 224.
[0072] The context model 222 dynamically switches the context model to be applied according to the decoding target and the surrounding situation. For example, when the context model 222 holds the context variable ctx and acquires the encoded data from the selection unit 221, it reads out the context variable ctx corresponding to each bin position (binIdx) of the bin sequence defined for each syntax element. The context model 222 supplies the encoded data and the read context variable ctx to the arithmetic decoding unit 223.
[0073] When the arithmetic decoding unit 223 acquires the encoded data and the context variable ctx supplied from the context model 222, it refers to the probability state of the context variable ctx and arithmetically decodes the value of the bin at binIdx in the binarized bit sequence in the normal mode of CABAC (performs context decoding). The arithmetic decoding unit 223 supplies the binarized bit sequence generated by the context decoding to the selection unit 225. Also, the arithmetic decoding unit 223 supplies the context variable ctx after the context decoding process to the context model 222 for retention.
[0074] The arithmetic decoding unit 224 arithmetically decodes the encoded data supplied from the selection unit 221 in the bypass mode of CABAC (performs bypass decoding). The arithmetic decoding unit 224 supplies the binarized bit sequence generated by the bypass decoding to the selection unit 225.
[0075] The selection unit 225 acquires the flag information isBypass and selects the binarized bit sequence to be supplied to the multivalue conversion unit 226 based on the value of isBypass. For example, when isBypass = 0, the selection unit 225 determines that it is the normal mode, acquires the binarized bit sequence supplied from the arithmetic decoding unit 223, and supplies it to the multivalue conversion unit 226. Also, when isBypass = 1, the selection unit 225 determines that it is the bypass mode, acquires the binarized bit sequence supplied from the arithmetic decoding unit 224, and supplies it to the multivalue conversion unit 226.
[0076] The multi-valued conversion unit 226 acquires the binary bit string supplied from the selection unit 225, multi-valued converts it by a method defined for each syntax element, and generates a syntax element value. The multi-valued conversion unit 226 outputs the syntax element value to the outside of the decoding device 200.
[0077] Note that these processing units (selection unit 221 to multi-valued conversion unit 226) have an arbitrary configuration. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Also, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and realize the above-described processing by executing a program using them. Of course, each processing unit may have both configurations, realize a part of the above-described processing by a logic circuit, and realize the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, a part of the processing units may realize a part of the above-described processing by a logic circuit, another part of the processing units may realize the above-described processing by executing a program, and still another processing unit may realize the above-described processing by both a logic circuit and program execution.
[0078] In such a decoding device 200, the present technology described in <1. Encoding of ST Identifier> is applied. That is, the arithmetic decoding unit 223 refers to a fixed context variable ctx (for example, "0") set for the first bit of the bin string of the secondary conversion control information (ST identifier lfnst_idx), and performs arithmetic decoding in the normal mode of CABAC.
[0079] By doing so, the decoding device 200 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the decoding device 200 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the decoding device 200 can suppress an increase in the load of the decoding process.
[0080] At that time, the arithmetic decoding unit 224 may arithmetically decode the second bit of the bit string in the bypass mode of CABAC based on the bypass flag set for the second bit of the bit string.
[0081] Further, the arithmetic decoding unit 223 may refer to the context variable ctx (for example, "0") common to the first bit set for the second bit of the bit string and arithmetically decode it in the normal mode of CABAC.
[0082] <Flow of Decoding Process> Next, an example of the flow of the decoding process executed by this decoding apparatus 200 will be described with reference to the flowchart of FIG. 6.
[0083] When the decoding process is started, the selection unit 221 of the decoding apparatus 200 reads out, in step S201, the context variable ctx corresponding to each bit position (binIdx) of the bin string defined for each syntax element, and the flag isBypass indicating whether it is in the bypass mode.
[0084] In step S202, the selection unit 221 determines whether it is in the bypass mode. If isBypass = 0 and it is determined that it is in the normal mode, the process proceeds to step S203.
[0085] In step S203, the arithmetic decoding unit 223 performs context decoding. That is, the arithmetic decoding unit 223 refers to the probability state of the context variable ctx and decodes the encoded data in the normal mode of CABAC to generate the value of the bit at the bit position (binIdx) of the bin string (synBins). When the process of step S203 ends, the process proceeds to step S205.
[0086] Also, if isBypass = 1 in step S202 and it is determined that it is in the bypass mode, the process proceeds to step S204.
[0087] In step S204, the arithmetic decoder 224 performs bypass decoding. That is, the arithmetic decoder 224 decodes the encoded data in the bypass mode of CABAC and generates the value of the bin at the bin position (binIdx) of the bin sequence (synBins). When the process of step S204 ends, the process proceeds to step S205.
[0088] In step S205, the selection unit 225 determines whether a predetermined break condition A is satisfied. The break condition A is defined based on the values of the bin sequence from binIdx = 0 to binIdx = k (the position k of the current binIdx) and the binarization method for each syntax element.
[0089] If it is determined that the break condition A is not satisfied, the process returns to step S201, and the subsequent processes for generating the value of the next bin position (binIdx) are executed. That is, the processes of steps S201 to S205 are executed for each bin position (binIdx).
[0090] And in step S205, if it is determined that the break condition A is satisfied, the process proceeds to step S206.
[0091] In step S206, the multi - value conversion unit 226 derives the syntax element value (syncVal) from the bin sequence (synBins) by the multi - value conversion process defined for each syntax element.
[0092] In step S207, the multi - value conversion unit 226 outputs the derived syntax element value (syncVal) to the outside of the decoder 200.
[0093] In such decoding processing, the present technology described in <1. Encoding of ST identifier> is applied. That is, in step S203, the arithmetic decoder 223 refers to the fixed context variable ctx (for example, "0") set for the first bit of the bin sequence of the secondary conversion control information (ST identifier lfnst_idx), and performs arithmetic decoding in the normal mode of CABAC.
[0094] By doing so, the decoding device 200 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the decoding device 200 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the decoding device 200 can suppress an increase in the load of the decoding process.
[0095] Also, in step S204, the arithmetic decoding unit 224 may arithmetically decode the second bin of the bin sequence in the bypass mode of CABAC based on the bypass flag set for the second bin of the bin sequence (Method 1).
[0096] Also, in step S203, the arithmetic decoding unit 223 may refer to the context variable ctx (for example, "0") common to the first bin set for the second bin of the bin sequence and arithmetically decode it in the normal mode of CABAC (Method 2).
[0097] <4. Third Embodiment> <Image Encoding Device> FIG. 7 is a block diagram showing an example of the configuration of an image encoding device which is an aspect of an image processing device to which the present technology is applied. The image encoding device 300 shown in FIG. 7 is a device that encodes moving image data. For example, the image encoding device 300 can encode moving image data by an encoding method described in any of the above-mentioned non-patent documents.
[0098] Note that in FIG. 7, main components such as processing units (blocks) and data flows are shown, and not all of the components shown in FIG. 7 are necessarily included. That is, in the image encoding device 300, there may be a processing unit not shown as a block in FIG. 7, or a process or data flow not shown as an arrow or the like in FIG. 7.
[0099] As shown in FIG. 7, the image encoding apparatus 300 includes a control unit 301, a rearrangement buffer 311, an arithmetic unit 312, an orthogonal transformation unit 313, a quantization unit 314, an encoding unit 315, an accumulation buffer 316, an inverse quantization unit 317, an inverse orthogonal transformation unit 318, an arithmetic unit 319, an in-loop filter unit 320, a frame memory 321, a prediction unit 322, and a rate control unit 323.
[0100] <Control Unit> Based on the external or the block size of a preset processing unit, the control unit 301 divides the moving image data held by the rearrangement buffer 311 into blocks (such as CU, PU, transformation blocks, etc.) of the processing unit. Further, the control unit 301 determines encoding parameters (header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, filter information Finfo, etc.) to be supplied to each block based on, for example, RDO (Rate-Distortion Optimization).
[0101] Details of these encoding parameters will be described later. When the control unit 301 determines the encoding parameters as described above, it supplies them to each block. Specifically, it is as follows.
[0102] The header information Hinfo is supplied to each block. The prediction mode information Pinfo is supplied to the encoding unit 315 and the prediction unit 322. The transformation information Tinfo is supplied to the encoding unit 315, the orthogonal transformation unit 313, the quantization unit 314, the inverse quantization unit 317, and the inverse orthogonal transformation unit 318. The filter information Finfo is supplied to the in-loop filter unit 320.
[0103] <Rearrangement Buffer> Each field (input image) of the moving image data is input to the image encoding device 300 in its playback order (display order). The rearrangement buffer 311 acquires and holds (stores) each input image in its playback order (display order). Based on the control of the control unit 301, the rearrangement buffer 311 rearranges the input images in the encoding order (decoding order) or divides them into blocks of the processing unit. The rearrangement buffer 311 supplies each processed input image to the arithmetic unit 312. Also, the rearrangement buffer 311 supplies each input image (original image) to the prediction unit 322 and the in-loop filter unit 320 as well.
[0104] <Arithmetic unit> The arithmetic unit 312 takes as input the image I corresponding to the block of the processing unit and the prediction image P supplied from the prediction unit 322, subtracts the prediction image P from the image I as shown in the following equation to derive the prediction residual D, and supplies it to the orthogonal transformation unit 313.
[0105] D = I - P
[0106] <Orthogonal transformation unit> The orthogonal transformation unit 313 takes as input the prediction residual D supplied from the arithmetic unit 312 and the transformation information Tinfo supplied from the control unit 301, and based on the transformation information Tinfo, performs an orthogonal transformation on the prediction residual D to derive the transformation coefficient Coeff. For example, the orthogonal transformation unit 313 performs a primary transformation on the prediction residual D to generate a primary transformation coefficient, and based on the ST identifier, performs a secondary transformation on the primary transformation coefficient to generate a secondary transformation coefficient. The orthogonal transformation unit 313 supplies the obtained secondary transformation coefficient to the quantization unit 314 as the transformation coefficient Coeff.
[0107] <Quantization unit> The quantization unit 314 takes as inputs the conversion coefficient Coeff supplied from the orthogonal conversion unit 313 and the conversion information Tinfo supplied from the control unit 301, and scales (quantizes) the conversion coefficient Coeff based on the conversion information Tinfo. Note that the rate of this quantization is controlled by the rate control unit 323. The quantization unit 314 supplies the quantized conversion coefficient obtained by such quantization, that is, the quantization conversion coefficient level level, to the encoding unit 315 and the inverse quantization unit 317.
[0108] <Encoding Unit> The encoding unit 315 takes as inputs the quantization conversion coefficient level level supplied from the quantization unit 314, various encoding parameters (header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, filter information Finfo, etc.) supplied from the control unit 301, information related to the filter such as the filter coefficient supplied from the in-loop filter unit 320, and information related to the optimal prediction mode supplied from the prediction unit 322. The encoding unit 315 performs variable-length encoding (for example, arithmetic encoding) on the quantization conversion coefficient level level to generate a bit sequence (encoded data).
[0109] Also, the encoding unit 315 derives residual information Rinfo from the quantization conversion coefficient level level, encodes the residual information Rinfo, and generates a bit sequence.
[0110] Furthermore, the encoding unit 315 includes the information related to the filter supplied from the in-loop filter unit 320 in the filter information Finfo, and includes the information related to the optimal prediction mode supplied from the prediction unit 322 in the prediction mode information Pinfo. Then, the encoding unit 315 encodes the above-mentioned various encoding parameters (header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, filter information Finfo, etc.) and generates a bit sequence.
[0111] Also, the encoding unit 315 multiplexes the bit sequences of the various types of information generated as described above to generate encoded data. The encoding unit 315 supplies the encoded data to the storage buffer 316.
[0112] <Accumulation buffer> The accumulation buffer 316 temporarily holds the encoded data obtained in the encoding unit 315. The accumulation buffer 316 outputs the held encoded data to the outside of the image encoding apparatus 300 as, for example, a bit stream or the like at a predetermined timing. For example, this encoded data is transmitted to the decoding side via an arbitrary recording medium, an arbitrary transmission medium, an arbitrary information processing apparatus, or the like. That is, the accumulation buffer 316 is also a transmission unit that transmits the encoded data (bit stream).
[0113] <Inverse quantization unit> The inverse quantization unit 317 performs processing related to inverse quantization. For example, the inverse quantization unit 317 takes as inputs the quantized transform coefficient level level supplied from the quantization unit 314 and the transform information Tinfo supplied from the control unit 301, and scales (inverse quantizes) the value of the quantized transform coefficient level level based on the transform information Tinfo. Note that this inverse quantization is the inverse process of the quantization performed in the quantization unit 314. The inverse quantization unit 317 supplies the transform coefficient Coeff_IQ obtained by such inverse quantization to the inverse orthogonal transform unit 318.
[0114] <Inverse orthogonal transform unit> The inverse orthogonal transform unit 318 performs processing related to inverse orthogonal transform. For example, the inverse orthogonal transform unit 318 takes as inputs the transform coefficient Coeff_IQ supplied from the inverse quantization unit 317 and the transform information Tinfo supplied from the control unit 301, and performs an inverse orthogonal transform on the transform coefficient Coeff_IQ based on the transform information Tinfo to derive the prediction residual D'. Note that this inverse orthogonal transform is the inverse process of the orthogonal transform performed in the orthogonal transform unit 313. That is, the inverse orthogonal transform unit 318 can perform an adaptive inverse orthogonal transform (AMT) that adaptively selects the type (transform coefficient) of the inverse orthogonal transform.
[0115] The inverse orthogonal transformation unit 318 supplies the prediction residual D' obtained by such inverse orthogonal transformation to the arithmetic unit 319. Since the inverse orthogonal transformation unit 318 is the same as the inverse orthogonal transformation unit on the decoding side (described later), the description of the inverse orthogonal transformation unit 318 on the decoding side (described later) can be applied.
[0116] <Arithmetic unit> The arithmetic unit 319 takes as inputs the prediction residual D' supplied from the inverse orthogonal transformation unit 318 and the predicted image P supplied from the prediction unit 322. The arithmetic unit 319 adds the prediction residual D' and the predicted image P corresponding to the prediction residual D' to derive a local decoded image Rlocal. The arithmetic unit 319 supplies the derived local decoded image Rlocal to the in-loop filter unit 320 and the frame memory 321.
[0117] <In-loop filter unit> The in-loop filter unit 320 performs processing related to in-loop filter processing. For example, the in-loop filter unit 320 takes as inputs the local decoded image Rlocal supplied from the arithmetic unit 319, the filter information Finfo supplied from the control unit 301, and the input image (original image) supplied from the rearrangement buffer 311. Note that the information input to the in-loop filter unit 320 is arbitrary, and information other than these may be input. For example, if necessary, information such as the prediction mode, motion information, coding amount target value, quantization parameter QP, picture type, and information on blocks (CU, CTU, etc.) may be input to the in-loop filter unit 320.
[0118] The in-loop filter unit 320 appropriately performs filter processing on the local decoded image Rlocal based on the filter information Finfo. The in-loop filter unit 320 uses the input image (original image) and other input information in the filter processing as necessary.
[0119] For example, as described in Non-Patent Document 11, the in-loop filter unit 320 applies four in-loop filters, namely, a bilateral filter, a deblocking filter (DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF), in this order. Note that which filter to apply and in what order are arbitrary and can be selected as appropriate.
[0120] Of course, the filtering process performed by the in-loop filter unit 320 is arbitrary and is not limited to the above example. For example, the in-loop filter unit 320 may apply a Wiener filter or the like.
[0121] The in-loop filter unit 320 supplies the filtered local decoded image Rlocal to the frame memory 321. When transmitting information related to the filter, such as filter coefficients, to the decoding side, for example, the in-loop filter unit 320 supplies the information related to the filter to the encoding unit 315.
[0122] <Frame Memory> The frame memory 321 performs processing related to storing data related to the image. For example, the frame memory 321 takes as input the local decoded image Rlocal supplied from the arithmetic unit 319 and the filtered local decoded image Rlocal supplied from the in-loop filter unit 320, and holds (stores) them. Further, the frame memory 321 reconstructs and holds the decoded image R for each picture unit using the local decoded image Rlocal (stores it in the buffer in the frame memory 321). The frame memory 321 supplies the decoded image R (or a part thereof) to the prediction unit 322 in response to a request from the prediction unit 322.
[0123] <Prediction Unit> The prediction unit 322 performs processing related to the generation of a predicted image. For example, the prediction unit 322 takes as inputs the prediction mode information Pinfo supplied from the control unit 301, the input image (original image) supplied from the rearrangement buffer 311, and the decoded image R (or a part thereof) read from the frame memory 321. The prediction unit 322 performs prediction processing such as inter prediction and intra prediction using the prediction mode information Pinfo and the input image (original image), makes a prediction with reference to the decoded image R as a reference image, performs motion compensation processing based on the prediction result, and generates a predicted image P. The prediction unit 322 supplies the generated predicted image P to the arithmetic unit 312 and the arithmetic unit 319. Also, the prediction unit 322 supplies information regarding the prediction mode selected by the above processing, that is, the optimal prediction mode, to the encoding unit 315 as necessary.
[0124] <Rate control unit> The rate control unit 323 performs processing related to rate control. For example, the rate control unit 323 controls the rate of the quantization operation of the quantization unit 314 so that overflow or underflow does not occur based on the amount of code of the encoded data stored in the accumulation buffer 316.
[0125] Note that these processing units (control unit 301, rearrangement buffer 311 to rate control unit 323) have arbitrary configurations. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Also, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and realize the above-described processing by executing a program using them. Of course, each processing unit may have both configurations, realize a part of the above-described processing by a logic circuit, and realize the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, a part of the processing units may realize a part of the above-described processing by a logic circuit, another part of the processing units may realize the above-described processing by executing a program, and still another processing unit may realize the above-described processing by both a logic circuit and program execution.
[0126] In the image encoding apparatus 300 configured as described above, the present technology described in <1. Encoding of ST Identifier> can be applied to the encoding unit 315. That is, the encoding unit 315 has the same configuration as the encoding apparatus 100 shown in FIG. 3 and performs the same processing.
[0127] By doing so, the image encoding apparatus 300 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the image encoding apparatus 300 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the image encoding apparatus 300 can suppress an increase in the load of the encoding process.
[0128] <Flow of Image Encoding Process> Next, an example of the flow of the image encoding process executed by the image encoding apparatus 300 configured as described above will be described with reference to the flowchart of FIG. 8.
[0129] When the image encoding process is started, in step S301, the rearrangement buffer 311 rearranges the order of the frames of the input moving image data from the display order to the encoding order under the control of the control unit 301.
[0130] In step S302, the control unit 301 sets a processing unit (performs block division) for the input image held by the rearrangement buffer 311.
[0131] In step S303, the control unit 301 determines (sets) the encoding parameters for the input image held by the rearrangement buffer 311.
[0132] In step S304, the prediction unit 322 performs prediction processing to generate a prediction image or the like in an optimal prediction mode. For example, in this prediction processing, the prediction unit 322 performs intra prediction to generate a prediction image or the like in an optimal intra prediction mode, performs inter prediction to generate a prediction image or the like in an optimal inter prediction mode, and selects an optimal prediction mode from among them based on a cost function value or the like.
[0133] In step S305, the arithmetic unit 312 calculates the difference between the input image and the prediction image in the optimal mode selected by the prediction processing in step S304. That is, the arithmetic unit 312 generates a prediction residual D between the input image and the prediction image. The prediction residual D obtained in this way has a reduced data amount compared to the original image data. Therefore, the data amount can be compressed compared to the case where the image is encoded as it is.
[0134] In step S306, the orthogonal transformation unit 313 performs orthogonal transformation processing on the prediction residual D generated by the processing in step S305 to derive transformation coefficients Coeff.
[0135] In step S307, the quantization unit 314 quantizes the transformation coefficients Coeff obtained by the processing in step S306 by using the quantization parameter calculated by the control unit 301 or the like, and derives a quantized transformation coefficient level level.
[0136] In step S308, the inverse quantization unit 317 inversely quantizes the quantized transformation coefficient level level generated by the processing in step S307 with a characteristic corresponding to the quantization characteristic in step S307 to derive transformation coefficients Coeff_IQ.
[0137] In step S309, the inverse orthogonal transform unit 318 performs an inverse orthogonal transform on the transform coefficient Coeff_IQ obtained by the process of step S308 in a manner corresponding to the orthogonal transform process of step S306 to derive a prediction residual D'. Note that since this inverse orthogonal transform process is the same as the inverse orthogonal transform process (described later) performed on the decoding side, the description of the inverse orthogonal transform process of this step S309 can apply the description for the decoding side (described later).
[0138] In step S310, the arithmetic unit 319 generates a locally decoded decoded image by adding the predicted image obtained by the prediction process of step S304 to the prediction residual D' derived by the process of step S309.
[0139] In step S311, the in-loop filter unit 320 performs an in-loop filter process on the locally decoded decoded image derived by the process of step S310.
[0140] In step S312, the frame memory 321 stores the locally decoded decoded image derived by the process of step S310 and the locally decoded decoded image filtered in step S311.
[0141] In step S313, the encoding unit 315 encodes the quantization transform coefficient level level obtained by the process of step S307. For example, the encoding unit 315 encodes the quantization transform coefficient level level, which is information about the image, by arithmetic coding or the like to generate encoded data. At this time, the encoding unit 315 also encodes various encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo). Furthermore, the encoding unit 315 derives residual information RInfo from the quantization transform coefficient level level and encodes the residual information RInfo.
[0142] In step S314, the accumulation buffer 316 accumulates the thus obtained encoded data and outputs it to the outside of the image encoding apparatus 300 as, for example, a bit stream. This bit stream is transmitted to the decoding side via, for example, a transmission path or a recording medium. Also, the rate control unit 323 performs rate control as necessary.
[0143] When the process of step S314 ends, the image encoding process ends.
[0144] In the image encoding process with the above flow, the present technology is applied to the encoding process of step S313. That is, in this step S313, an encoding process with the same flow as in FIG. 4 is performed. That is, the encoding unit 315 performs an encoding process to which the present technology described in <1. Encoding of ST identifier> is applied.
[0145] By doing so, the image encoding apparatus 300 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the image encoding apparatus 300 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the image encoding apparatus 300 can suppress an increase in the load of the encoding process.
[0146] <5. Fourth Embodiment> <Image Decoding Apparatus> FIG. 9 is a block diagram showing an example of the configuration of an image decoding apparatus which is an aspect of an image processing apparatus to which the present technology is applied. The image decoding apparatus 400 shown in FIG. 9 is an apparatus for decoding encoded data of a moving image. For example, the image decoding apparatus 400 can decode the encoded data by a decoding method described in any of the above non-patent documents. For example, the image decoding apparatus 400 decodes the encoded data (bit stream) generated by the above-described image encoding apparatus 300.
[0147] Note that in FIG. 9, the main components such as the processing unit (block) and the data flow are shown, but not all of the components shown in FIG. 9 are necessarily included. That is, in the image decoding apparatus 400, there may be a processing unit that is not shown as a block in FIG. 9, or there may be a process or data flow that is not shown as an arrow or the like in FIG. 9.
[0148] In FIG. 9, the image decoding apparatus 400 includes an accumulation buffer 411, a decoding unit 412, an inverse quantization unit 413, an inverse orthogonal transformation unit 414, an arithmetic unit 415, an in-loop filter unit 416, a rearrangement buffer 417, a frame memory 418, and a prediction unit 419. Note that the prediction unit 419 includes an intra prediction unit and an inter prediction unit (not shown). The image decoding apparatus 400 is an apparatus for generating moving image data by decoding encoded data (bit stream).
[0149] <Accumulation Buffer> The accumulation buffer 411 acquires and holds (stores) the bit stream input to the image decoding apparatus 400. The accumulation buffer 411 supplies the stored bit stream to the decoding unit 412 at a predetermined timing or when a predetermined condition is satisfied.
[0150] <Decoding Unit> The decoding unit 412 performs processing related to image decoding. For example, the decoding unit 412 takes the bit stream supplied from the accumulation buffer 411 as an input, and performs variable length decoding of the syntax values of the respective syntax elements from the bit string in accordance with the definition of the syntax table, and derives parameters.
[0151] The parameters derived from the syntax elements and the syntax values of the syntax elements include, for example, information such as header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, residual information Rinfo, and filter information Finfo. That is, the decoding unit 412 parses (analyzes and acquires) this information from the bit stream. These pieces of information will be described below.
[0152] <Header information Hinfo> The header information Hinfo includes header information such as VPS (Video Parameter Set) / SPS (Sequence Parameter Set) / PPS (Picture Parameter Set) / SH (slice header). The header information Hinfo includes information that defines, for example, the image size (horizontal width PicWidth, vertical height PicHeight), bit depth (luminance bitDepthY, chrominance bitDepthC), chroma array type ChromaArrayType, maximum value MaxCUSize / minimum value MinCUSize of the CU size, maximum depth MaxQTDepth / minimum depth MinQTDepth of the quadtree split (also called quad-tree split), maximum depth MaxBTDepth / minimum depth MinBTDepth of the binary tree split (binary-tree split), maximum value MaxTSSize of the transform skip block (also called the maximum transform skip block size), on / off flags (also called enable flags) of each encoding tool, etc.
[0153] For example, as the on / off flags of the encoding tools included in the header information Hinfo, there are on / off flags related to the following conversion and quantization processes. Note that the on / off flag of the encoding tool can also be interpreted as a flag indicating whether the syntax related to the encoding tool exists in the encoded data. Also, when the value of the on / off flag is 1 (true), it indicates that the encoding tool is available, and when the value of the on / off flag is 0 (false), it indicates that the encoding tool is not available. Note that the interpretation of the flag value may be reversed.
[0154] Cross-component prediction enable flag (ccp_enabled_flag): It is flag information indicating whether cross-component prediction (CCP (Cross-Component Prediction), also called CC prediction) is available. For example, when this flag information is "1" (true), it indicates that it is available, and when it is "0" (false), it indicates that it is not available.
[0155] Note that this CCP is also referred to as component - to - component linear prediction (CCLM or CCLMP).
[0156] <Prediction mode information Pinfo> The prediction mode information Pinfo includes, for example, information such as the size information PBSize (prediction block size) of the PB (prediction block) to be processed, intra - prediction mode information IPinfo, motion prediction information MVinfo, etc.
[0157] The intra - prediction mode information IPinfo includes, for example, prev_intra_luma_pred_flag, mpm_idx, rem_intra_pred_mode in JCTVC - W1005, 7.3.8.5 Coding Unit syntax, and the luminance intra - prediction mode IntraPredModeY derived from its syntax, etc.
[0158] Also, the intra - prediction mode information IPinfo includes, for example, the component - to - component prediction flag (ccp_flag (cclmp_flag)), multi - class linear prediction mode flag (mclm_flag), chroma sample position type identifier (chroma_sample_loc_type_idx), chroma MPM identifier (chroma_mpm_idx), and the luminance intra - prediction mode (IntraPredModeC) derived from these syntaxes, etc.
[0159] The component - to - component prediction flag (ccp_flag (cclmp_flag)) is flag information indicating whether to apply component - to - component linear prediction. For example, when ccp_flag == 1, it indicates that component - to - component prediction is applied, and when ccp_flag == 0, it indicates that component - to - component prediction is not applied.
[0160] The multi-class linear prediction mode flag (mclm_flag) is information regarding the mode of linear prediction (linear prediction mode information). More specifically, the multi-class linear prediction mode flag (mclm_flag) is flag information indicating whether to use the multi-class linear prediction mode or not. For example, when it is "0", it indicates that it is in the 1-class mode (single-class mode) (e.g., CCLMP), and when it is "1", it indicates that it is in the 2-class mode (multi-class mode) (e.g., MCLMP).
[0161] The chroma sample position type identifier (chroma_sample_loc_type_idx) is an identifier that identifies the type of the pixel position of the chroma component (also referred to as the chroma sample position type). For example, when the chroma array type (ChromaArrayType), which is information regarding the color format, indicates the 420 format, the chroma sample position type identifier is assigned as in the following formula.
[0162] chroma_sample_loc_type_idx == 0:Type2 chroma_sample_loc_type_idx == 1:Type3 chroma_sample_loc_type_idx == 2:Type0 chroma_sample_loc_type_idx == 3:Type1
[0163] Note that this chroma sample position type identifier (chroma_sample_loc_type_idx) is transmitted as (stored as) information (chroma_sample_loc_info()) regarding the pixel position of the chroma component.
[0164] The chroma MPM identifier (chroma_mpm_idx) is an identifier that indicates which prediction mode candidate in the chroma intra prediction mode candidate list (intraPredModeCandListC) is specified as the chroma intra prediction mode.
[0165] The motion prediction information MVinfo includes, for example, information such as merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X = {0,1}, mvd, etc. (see, for example, JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax).
[0166] Of course, the information included in the prediction mode information Pinfo is arbitrary, and information other than these may be included.
[0167] <Transformation information Tinfo> The transformation information Tinfo includes, for example, the following information. Of course, the information included in the transformation information Tinfo is arbitrary, and information other than these may be included.
[0168] The horizontal size TBWSize and vertical size TBHSize of the transformation block to be processed (or the logarithmic values log2TBWSize and log2TBHSize of each TBWSize and TBHSize with base 2 may also be used). The transformation skip flag (ts_flag): A flag indicating whether to skip the (inverse) primary transformation and the (inverse) secondary transformation. The scan identifier (scanIdx) The quantization parameter (qp) The quantization matrix (scaling_matrix (see, for example, JCTVC-W1005, 7.3.4 Scaling list data syntax))
[0169] <Residual information Rinfo> The residual information Rinfo (see, for example, 7.3.8.11 Residual Coding syntax of JCTVC-W1005) includes, for example, the following syntax.
[0170] cbf (coded_block_flag): A flag indicating the presence or absence of residual data last_sig_coeff_x_pos: Last non-zero coefficient X coordinate last_sig_coeff_y_pos: Last non-zero coefficient Y coordinate coded_sub_block_flag: Sub-block non-zero coefficient presence flag sig_coeff_flag: Non-zero coefficient presence flag gr1_flag: Flag indicating whether the level of the non-zero coefficient is greater than 1 (also called GR1 flag) gr2_flag: Flag indicating whether the level of the non-zero coefficient is greater than 2 (also called GR2 flag) sign_flag: Sign indicating the positive or negative of the non-zero coefficient (also called sign bit) coeff_abs_level_remaining: Remaining level of the non-zero coefficient (also called non-zero coefficient remaining level) etc.
[0171] Of course, the information included in the residual information Rinfo is arbitrary, and information other than these may also be included.
[0172] <Filter information Finfo> The filter information Finfo includes, for example, control information regarding each of the following filter processes.
[0173] Control information regarding the deblocking filter (DBF) Control information regarding the sample adaptive offset (SAO) Control information regarding the adaptive loop filter (ALF) Control information regarding other linear and non-linear filters
[0174] More specifically, for example, it includes information specifying the picture to which each filter is applied, the area within the picture, CU unit filter On / Off control information, filter On / Off control information regarding the boundaries of slices and tiles, etc. Of course, the information included in the filter information Finfo is arbitrary, and information other than these may also be included.
[0175] Return to the description of the decoding unit 412. The decoding unit 412 refers to the residual information Rinfo and derives the quantized transform coefficient level level at each coefficient position within each transform block. The decoding unit 412 supplies the quantized transform coefficient level level to the inverse quantization unit 413.
[0176] Also, the decoding unit 412 supplies the parsed header information Hinfo, prediction mode information Pinfo, quantized transform coefficient level level, transform information Tinfo, and filter information Finfo to each block. Specifically, it is as follows.
[0177] The header information Hinfo is supplied to the inverse quantization unit 413, inverse orthogonal transform unit 414, prediction unit 419, and in-loop filter unit 416. The prediction mode information Pinfo is supplied to the inverse quantization unit 413 and the prediction unit 419. The transform information Tinfo is supplied to the inverse quantization unit 413 and the inverse orthogonal transform unit 414. The filter information Finfo is supplied to the in-loop filter unit 416.
[0178] Of course, the above example is just an example and is not limited to this example. For example, each encoding parameter may be supplied to any processing unit. Also, other information may be supplied to any processing unit.
[0179] <Inverse Quantization Unit> The inverse quantization unit 413 has at least a configuration necessary for performing processing related to inverse quantization. For example, the inverse quantization unit 413 takes as inputs the transform information Tinfo and the quantized transform coefficient level level supplied from the decoding unit 412, scales (inverse quantizes) the value of the quantized transform coefficient level level based on the transform information Tinfo, and derives the inverse quantized transform coefficient Coeff_IQ.
[0180] Note that this inverse quantization is performed as an inverse process of the quantization by the quantization unit 314. Also, this inverse quantization is the same process as the inverse quantization by the inverse quantization unit 317. That is, the inverse quantization unit 317 performs the same process (inverse quantization) as the inverse quantization unit 413.
[0181] The inverse quantization unit 413 supplies the derived conversion coefficient Coeff_IQ to the inverse orthogonal transform unit 414.
[0182] <Inverse orthogonal transform unit> The inverse orthogonal transform unit 414 performs processing related to inverse orthogonal transform. For example, the inverse orthogonal transform unit 414 takes as inputs the conversion coefficient Coeff_IQ supplied from the inverse quantization unit 413 and the conversion information Tinfo supplied from the decoding unit 412, and based on the conversion information Tinfo, performs an inverse orthogonal transform process on the conversion coefficient Coeff_IQ to derive the prediction residual D'. For example, the inverse orthogonal transform unit 414 performs an inverse secondary transform on the conversion coefficient Coeff_IQ based on the ST identifier to generate a primary conversion coefficient, performs a primary transform on the primary conversion coefficient, and generates the prediction residual D'.
[0183] Note that this inverse orthogonal transform is performed as an inverse process of the orthogonal transform by the orthogonal transform unit 313. Also, this inverse orthogonal transform is the same process as the inverse orthogonal transform by the inverse orthogonal transform unit 318. That is, the inverse orthogonal transform unit 318 performs the same process (inverse orthogonal transform) as the inverse orthogonal transform unit 414.
[0184] The inverse orthogonal transform unit 414 supplies the derived prediction residual D' to the arithmetic unit 415.
[0185] <Arithmetic unit> The arithmetic unit 415 performs processing related to addition of information regarding the image. For example, the arithmetic unit 415 takes as inputs the prediction residual D' supplied from the inverse orthogonal transform unit 414 and the predicted image P supplied from the prediction unit 419. The arithmetic unit 415 adds the prediction residual D' and the predicted image P (prediction signal) corresponding to the prediction residual D' as shown in the following formula to derive the local decoded image Rlocal.
[0186] Rlocal = D' + P
[0187] The arithmetic unit 415 supplies the derived local decoded image Rlocal to the in-loop filter unit 416 and the frame memory 418.
[0188] <In-loop filter unit> The in-loop filter unit 416 performs processing related to in-loop filter processing. For example, the in-loop filter unit 416 takes as inputs the local decoded image Rlocal supplied from the arithmetic unit 415 and the filter information Finfo supplied from the decoding unit 412. Note that the information input to the in-loop filter unit 416 is arbitrary, and information other than these may be input.
[0189] The in-loop filter unit 416 appropriately performs filter processing on the local decoded image Rlocal based on the filter information Finfo.
[0190] For example, as described in Non-Patent Document 11, the in-loop filter unit 416 applies four in-loop filters, namely, a bilateral filter, a deblocking filter (DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF), in this order. Note that which filter to apply and in what order are arbitrary and can be selected as appropriate.
[0191] The in-loop filter unit 416 performs filter processing corresponding to the filter processing performed by the encoding side (for example, the in-loop filter unit 320 of the image encoding device 300). Of course, the filter processing performed by the in-loop filter unit 416 is arbitrary and is not limited to the above example. For example, the in-loop filter unit 416 may apply a Wiener filter or the like.
[0192] The in-loop filter unit 416 supplies the filtered local decoded image Rlocal to the rearrangement buffer 417 and the frame memory 418.
[0193] <Rearrangement buffer> The rearrangement buffer 417 takes the local decoded image Rlocal supplied from the in-loop filter unit 416 as input and holds (stores) it. The rearrangement buffer 417 reconstructs the decoded image R for each picture unit using the local decoded image Rlocal and holds (stores it in the buffer). The rearrangement buffer 417 rearranges the obtained decoded image R from the decoding order to the playback order. The rearrangement buffer 417 outputs the rearranged group of decoded images R as moving image data to the outside of the image decoding apparatus 400.
[0194] <Frame memory> The frame memory 418 performs processing related to the storage of data related to the image. For example, the frame memory 418 takes the local decoded image Rlocal supplied from the arithmetic unit 415 as input, reconstructs the decoded image R for each picture unit, and stores it in the buffer in the frame memory 418.
[0195] Also, the frame memory 418 takes the in-loop filtered local decoded image Rlocal supplied from the in-loop filter unit 416 as input, reconstructs the decoded image R for each picture unit, and stores it in the buffer in the frame memory 418. The frame memory 418 appropriately supplies the stored decoded image R (or a part thereof) to the prediction unit 419 as a reference image.
[0196] Note that the frame memory 418 may store header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, filter information Finfo, etc. related to the generation of the decoded image.
[0197] <Prediction unit> The prediction unit 419 performs processing related to the generation of a predicted image. For example, the prediction unit 419 takes the prediction mode information Pinfo supplied from the decoding unit 412 as an input, performs prediction by the prediction method specified by the prediction mode information Pinfo, and derives a predicted image P. At the time of the derivation, the prediction unit 419 uses the decoded image R (or a part thereof) before or after the filter stored in the frame memory 418, which is specified by the prediction mode information Pinfo, as a reference image. The prediction unit 419 supplies the derived predicted image P to the arithmetic unit 415.
[0198] Note that these processing units (accumulation buffer 411 to prediction unit 419) have an arbitrary configuration. For example, each processing unit may be configured by a logic circuit that realizes the above-described processing. Also, each processing unit may have, for example, a CPU, a ROM, a RAM, etc., and realize the above-described processing by executing a program using them. Of course, each processing unit may have both configurations, realize a part of the above-described processing by a logic circuit, and realize the other part by executing a program. The configurations of the respective processing units may be independent of each other. For example, a part of the processing units may realize a part of the above-described processing by a logic circuit, another part of the processing units may realize the above-described processing by executing a program, and still another processing unit may realize the above-described processing by both a logic circuit and program execution.
[0199] In the image decoding apparatus 400 having the above configuration, the present technology described in <1. Encoding of ST identifier> can be applied to the decoding unit 412. That is, the decoding unit 412 has the same configuration as the decoding apparatus 200 shown in FIG. 5 and performs the same processing.
[0200] By doing so, the image decoding apparatus 400 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the image decoding apparatus 400 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the image decoding apparatus 400 can suppress an increase in the load of the decoding process.
[0201] <Flow of Image Encoding Process> Next, an example of the flow of the image decoding process executed by the image decoding apparatus 400 having the above configuration will be described with reference to the flowchart of FIG. 10.
[0202] When the image decoding process is started, the accumulation buffer 411 acquires and holds (accumulates) the encoded data (bit stream) supplied from the outside of the image decoding apparatus 400 in step S401.
[0203] In step S402, the decoding unit 412 decodes the encoded data (bit stream) to obtain the quantized transform coefficient level level. Further, the decoding unit 412 parses (analyzes and acquires) various encoding parameters from the encoded data (bit stream) by this decoding.
[0204] In step S403, the inverse quantization unit 413 performs inverse quantization, which is the inverse process of quantization performed on the encoding side, on the quantized transform coefficient level level obtained by the process of step S402, to obtain the transform coefficient Coeff_IQ.
[0205] In step S404, the inverse orthogonal transformation unit 414 performs an inverse orthogonal transformation process, which is the inverse process of the orthogonal transformation process performed on the encoding side, on the transform coefficient Coeff_IQ obtained in step S403, to obtain the prediction residual D'.
[0206] In step S405, the prediction unit 419 executes a prediction process by the prediction method specified from the encoding side based on the information parsed in step S402, and generates a prediction image P by referring to the reference image stored in the frame memory 418 or the like.
[0207] In step S406, the arithmetic unit 415 adds the prediction residual D' obtained in step S404 and the prediction image P obtained in step S405 to derive a local decoded image Rlocal.
[0208] In step S407, the in-loop filter unit 416 performs an in-loop filter process on the local decoded image Rlocal obtained by the process of step S406.
[0209] In step S408, the rearrangement buffer 417 derives the decoded image R using the filtered local decoded image Rlocal obtained by the process of step S407, and rearranges the order of the decoded image R group from the decoding order to the playback order. The group of decoded images R rearranged in the playback order is output to the outside of the image decoder 400 as a moving image.
[0210] Also, in step S409, the frame memory 418 stores at least one of the local decoded image Rlocal obtained by the process of step S406 and the local decoded image Rlocal after the filter process obtained by the process of step S407.
[0211] When the process of step S409 ends, the image decoding process ends.
[0212] In the image decoding process with the above flow, this technology is applied to the decoding process of step S402. That is, in this step S402, a decoding process with the same flow as in FIG. 6 is performed. That is, the decoding unit 412 performs a decoding process applying this technology described in <1. Encoding of ST identifier>.
[0213] By doing so, the image decoder 400 can simplify the derivation of the context variable of the ST identifier while maintaining the encoding efficiency. Also, the image decoder 400 can reduce the memory size for holding the context variable of the ST identifier while maintaining the encoding efficiency. That is, the image decoder 400 can suppress an increase in the load of the decoding process.
[0214] <6. Supplementary Note> <Computer> The above-described series of processes can be executed either by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer. Here, the computer includes a computer incorporated in dedicated hardware, or a general-purpose personal computer or the like that can execute various functions by installing various programs.
[0215] FIG. 11 is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.
[0216] In the computer 800 shown in FIG. 11, a CPU (Central Processing Unit) 801, a ROM (Read Only Memory) 802, and a RAM (Random Access Memory) 803 are interconnected via a bus 804.
[0217] An input / output interface 810 is also connected to the bus 804. Connected to the input / output interface 810 are an input unit 811, an output unit 812, a storage unit 813, a communication unit 814, and a drive 815.
[0218] The input unit 811 includes, for example, a keyboard, a mouse, a microphone, a touch panel, input terminals, and the like. The output unit 812 includes, for example, a display, a speaker, output terminals, and the like. The storage unit 813 includes, for example, a hard disk, a RAM disk, a non-volatile memory, and the like. The communication unit 814 includes, for example, a network interface. The drive 815 drives a removable medium 821 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0219] In the computer configured as described above, the CPU 801 loads and executes, for example, a program stored in the storage unit 813 via the input / output interface 810 and the bus 804 into the RAM 803, thereby performing the series of processes described above. The RAM 803 also appropriately stores data and the like necessary for the CPU 801 to execute various processes.
[0220] The program executed by the computer can be recorded and applied, for example, on a removable medium 821 such as a package medium. In that case, the program can be installed in the storage unit 813 via the input / output interface 810 by mounting the removable medium 821 on the drive 815.
[0221] Also, this program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 814 and installed in the storage unit 813.
[0222] In addition, this program can be installed in advance in the ROM 802 or the storage unit 813.
[0223] <Applicable Object of the Present Technology> The present technology can be applied to any image encoding / decoding method. That is, as long as it does not conflict with the present technology described above, the specifications of various processes related to image encoding / decoding such as conversion (inverse conversion), quantization (inverse quantization), encoding (decoding), prediction, etc. are arbitrary and are not limited to the examples described above. Also, as long as it does not conflict with the present technology described above, some of these processes may be omitted.
[0224] Also, the present technology can be applied to a multi-viewpoint image encoding / decoding system that performs encoding / decoding of images of a plurality of viewpoints (views). In that case, the present technology may be applied to the encoding / decoding of each viewpoint (view).
[0225] Furthermore, the present technology can be applied to a hierarchical image encoding (scalable encoding) / decoding system that encodes and decodes hierarchical images that are layered (hierarchical) to have a scalability function for predetermined parameters. In that case, the present technology may be applied to the encoding and decoding of each layer.
[0226] Also, in the above, as an application example of the present technology, the encoding device 100, the decoding device 200, the image encoding device 300, and the image decoding device 400 have been described, but the present technology can be applied to any configuration.
[0227] For example, the present technology can be applied to various electronic devices such as transmitters and receivers (e.g., television receivers and mobile phones) in satellite broadcasting, cable broadcasting such as cable TV, distribution on the Internet, and distribution to terminals by cellular communication, or devices that record images on media such as optical disks, magnetic disks, and flash memories, or reproduce images from these storage media (e.g., hard disk recorders and cameras).
[0228] Also, for example, the present technology can be implemented as a part of a device such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using a plurality of processors, etc., a unit (e.g., a video unit) using a plurality of modules, etc., or a set (e.g., a video set) with other functions added to the unit.
[0229] Furthermore, for example, the present technology can also be applied to a network system composed of multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and jointly process via a network. For example, the present technology may be implemented in a cloud service that provides services related to images (moving images) to any terminal such as a computer, an AV (Audio Visual) device, a portable information processing terminal, an IoT (Internet of Things) device, etc.
[0230] In this specification, a system means a collection of multiple components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a single device in which a plurality of modules are housed in one housing are both systems.
[0231] <Fields of Application and Uses Applicable to the Present Technology> The system, device, processing unit, etc. to which the present technology is applied can be used in any field such as, for example, transportation, medical care, crime prevention, agriculture, livestock industry, mining, beauty, factories, home appliances, meteorology, natural monitoring, etc. Also, its use is arbitrary.
[0232] For example, the present technology can be applied to a system or device used for providing ornamental content or the like. Also, for example, the present technology can also be applied to a system or device used for transportation such as monitoring of traffic conditions and automatic driving control. Furthermore, for example, the present technology can also be applied to a system or device used for security. Also, for example, the present technology can be applied to a system or device used for automatic control of machines and the like. Furthermore, for example, the present technology can also be applied to a system or device used for agriculture and the livestock industry. Also, the present technology can be applied to a system or device for monitoring the state of nature such as volcanoes, forests, oceans, etc. and wild animals. Furthermore, for example, the present technology can also be applied to a system or device used for sports.
[0233] <Others> In addition, in this specification, a "flag" is information for identifying a plurality of states, and includes not only information used for identifying two states of true (1) or false (0), but also information capable of identifying three or more states. Therefore, the values that this "flag" can take may be, for example, two values of 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary and may be 1 bit or a plurality of bits. Also, identification information (including flags) is assumed to be included in the bit stream not only in the form of including the identification information itself, but also in the form of including the difference information of the identification information with respect to a certain reference information. Therefore, in this specification, "flags" and "identification information" include not only the information itself, but also the difference information with respect to the reference information.
[0234] Also, various information (such as metadata) related to the encoded data (bit stream) may be transmitted or recorded in any form as long as it is associated with the encoded data. Here, the term "associate" means, for example, making it possible to use (link) the other data when processing one data. That is, the data associated with each other may be grouped as one data or may be individual data. For example, the information associated with the encoded data (image) may be transmitted on a transmission path different from that of the encoded data (image). Also, for example, the information associated with the encoded data (image) may be recorded on a recording medium different from that of the encoded data (image) (or a different recording area of the same recording medium). Note that this "association" may be a part of the data, not the entire data. For example, an image and the information corresponding to the image may be associated with each other in any unit such as a plurality of frames, one frame, or a part within a frame.
[0235] In addition, in this specification, terms such as "synthesize", "multiplex", "add", "integrate", "include", "store", "inject", "insert", "plug in", etc. mean to combine multiple things into one, for example, combining encoded data and metadata into one piece of data, and it means one method of the above-mentioned "associate".
[0236] Also, the embodiments of the present technology are not limited to the above-described embodiments, and various changes can be made without departing from the gist of the present technology.
[0237] For example, the configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). Conversely, the configurations described as a plurality of devices (or processing units) above may be combined and configured as one device (or processing unit). Also, of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, if the overall configuration and operation of the system are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or another processing unit).
[0238] Also, for example, the above-described program may be executed on any device. In that case, it suffices if the device has necessary functions (such as function blocks) and can obtain necessary information.
[0239] Also, for example, each step of one flowchart may be executed by one device, or may be executed by a plurality of devices in cooperation. Furthermore, when a plurality of processes are included in one step, the plurality of processes may be executed by one device, or may be executed by a plurality of devices in cooperation. In other words, the plurality of processes included in one step can also be executed as processes of a plurality of steps. Conversely, the processes described as a plurality of steps can also be executed as one step.
[0240] In addition, for example, the program executed by a computer may be such that the processing of the steps of describing the program is executed in time series along the order described in this specification, or may be executed in parallel, or may be executed individually at a necessary timing such as when a call is made. That is, as long as there is no contradiction, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of describing this program may be executed in parallel with the processing of other programs, or may be executed in combination with the processing of other programs.
[0241] In addition, for example, as long as there is no contradiction, a plurality of technologies related to this technology can be implemented independently and individually. Of course, any plurality of this technology can also be implemented in combination. For example, a part or all of the technology described in any one of the embodiments can also be implemented in combination with a part or all of the technology described in other embodiments. In addition, a part or all of any of the above-described technologies can also be implemented in combination with other technologies not described above.
[0242] Note that this technology can also have the following configuration. (1) An encoding unit that sets a fixed context variable for the first bit of the bit string of secondary conversion control information, which is control information related to secondary conversion, and arithmetic-encodes each bit of the secondary conversion control information with reference to the set fixed context variable An image processing apparatus comprising the same. (2) The fixed context variable is "0" The image processing apparatus according to (1). (3) The encoding unit further sets a bypass flag for the second bit of the bit string The image processing apparatus according to (1). (4) The encoding unit further sets a context variable common to the first bit for the second bit of the bit string The image processing apparatus according to (1). (5) The common context is "0" The image processing apparatus according to (4). (6) The secondary conversion control information includes a secondary conversion identifier indicating the type of the secondary conversion. (1) The image processing apparatus according to (1). (7) The encoding unit binarizes the secondary conversion control information to generate the bit string, and sets the fixed context variable for the first bit of the generated bit string. (1) The image processing apparatus according to (1). (8) The apparatus further includes a coefficient conversion unit that performs secondary conversion on the primary conversion coefficients of the conversion block corresponding to the secondary conversion control information to generate secondary conversion coefficients. (1) The image processing apparatus according to (1). (9) The apparatus further includes a quantization unit that quantizes the secondary conversion coefficients generated by the coefficient conversion unit to generate quantization coefficients. (8) The image processing apparatus according to (8). (10) A fixed context variable is set for the first bit of the bit string of the secondary conversion control information, which is control information regarding secondary conversion, and each bit of the secondary conversion control information is arithmetic-coded with reference to the set fixed context variable. An image processing method.
[0243] (11) A decoding unit that performs arithmetic decoding with reference to the fixed context variable set for the first bit of the bit string of the secondary conversion control information, which is control information regarding secondary conversion. An image processing apparatus comprising the same. (12) The fixed context variable is "0". (11) The image processing apparatus according to (11). (13) The decoding unit further performs arithmetic decoding based on a bypass flag set for the second bit of the bit string. (11) The image processing apparatus according to (11). (14) The decoding unit further performs arithmetic decoding with reference to a context variable common to the first bit set for the second bit of the bit string. (11) The image processing apparatus according to (11). (15) The common context is "0". The image processing apparatus according to (14). (16) The secondary conversion control information includes a secondary conversion identifier indicating the type of the secondary conversion. The image processing apparatus according to (11). (17) The decoding unit generates the secondary conversion control information by multi-valuing the bit string obtained by arithmetic decoding. The image processing apparatus according to (11). (18) The image processing apparatus further includes an inverse coefficient conversion unit that performs an inverse secondary conversion on the secondary conversion coefficient of the conversion block corresponding to the secondary conversion control information to generate a primary conversion coefficient. The image processing apparatus according to (11). (19) The image processing apparatus further includes an inverse quantization unit that inverse quantizes quantization coefficients to generate the secondary conversion coefficients, and the inverse coefficient conversion unit performs the inverse secondary conversion on the secondary conversion coefficients generated by the inverse quantization unit to generate the primary conversion coefficients. The image processing apparatus according to (18). (20) Perform arithmetic decoding with reference to a fixed context variable set for the first bit of the bit string of the secondary conversion control information, which is control information regarding secondary conversion. Image processing method.
Description of Signs
[0244] 100 Encoding device, 121 Binarization unit, 122 Selection unit, 123 Context model, 124 Arithmetic encoding unit, 125 Arithmetic encoding unit, 126 Selection unit, 200 Decoding device, 221 Selection unit, 222 Context model, 223 Arithmetic decoding unit, 224 Arithmetic decoding unit, 225 Selection unit, 226 Multi-valuing unit, 300 Image decoding device, 315 Encoding unit, 400 Image encoding device, 412 Decoding unit
Claims
1. An encoding unit that arithmetic-encodes the second bit of the binary sequence corresponding to the secondary conversion control information, which is control information regarding secondary conversion, by referring to a context variable preset as one fixed value for the second bit of the binary sequence An image processing apparatus comprising the same.
2. The encoding unit arithmetic-encodes the first bit of the secondary conversion control information by referring to a context variable preset as one fixed value for the first bit of the binary sequence of the secondary conversion control information The image processing apparatus according to claim 1.
3. The context variable preset as one fixed value for the first bit is 0 The image processing apparatus according to claim 2.
4. The secondary conversion control information includes a secondary conversion identifier indicating the type of the secondary conversion The image processing apparatus according to claim 1.
5. The apparatus further comprises a coefficient conversion unit that performs the secondary conversion on the primary conversion coefficients of the block corresponding to the secondary conversion control information to generate secondary conversion coefficients The image processing apparatus according to claim 1.
6. An image processing method of arithmetic-encoding the second bit of the secondary conversion control information, which is control information regarding secondary conversion, by referring to a context variable preset as one fixed value for the second bit of the binary sequence corresponding to the secondary conversion control information
7. A decoding unit that arithmetic-decodes the second bit of the secondary conversion control information, which is control information regarding secondary conversion, by referring to a context variable preset as one fixed value for the second bit of the binary sequence corresponding to the secondary conversion control information An image processing apparatus comprising the same.
8. The decoding unit arithmetic-decodes the first bit of the secondary conversion control information by referring to a context variable preset as one fixed value for the first bit of the binary sequence of the secondary conversion control information The image processing apparatus according to claim 7.
9. The context variable preset as one fixed value for the first bit is 0 The image processing apparatus according to claim 8.
10. The secondary conversion control information includes a secondary conversion identifier indicating the type of the secondary conversion The image processing apparatus according to claim 7.
11. Further comprising an inverse coefficient conversion unit that performs an inverse secondary conversion on the secondary conversion coefficients of the block corresponding to the secondary conversion control information to generate primary conversion coefficients The image processing apparatus according to claim 7
12. Refer to a context variable preset as one fixed value for the second bin of the bin string corresponding to the secondary conversion control information, which is control information regarding secondary conversion, and arithmetically decode the second bin of the secondary conversion control information Image processing method