Symbolizing device, decoding device, and bit stream generating device
The encoding device enhances encoding efficiency by applying different quantizations to primary and secondary coefficients based on secondary conversion, addressing the limitations of existing technologies in maintaining subjective image quality.
Patent Information
- Application Number
- JP2024126095
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-08-31
- Filing Date
- 2024-08-01
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2038-07-25
AI Technical Summary
Existing encoding and decoding technologies, such as those in the HEVC standard, require further improvements to enhance encoding efficiency while maintaining subjective image quality.
An encoding device that performs primary conversion of image block residuals into primary coefficients, determines whether to apply secondary conversion, and if so, performs secondary conversion to obtain secondary coefficients. Different quantizations are applied based on the application of secondary conversion, using distinct quantization matrices for primary and secondary coefficients.
This approach improves encoding efficiency by allowing different quantization strategies for primary and secondary coefficients, thereby reducing loss in low-frequency bands and increasing loss in high-frequency bands, which helps in maintaining subjective image quality.
Smart Images

Figure 0007690097000003 
Figure 0007690097000004 
Figure 0007690097000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to an encoding device that encodes image blocks and the like.
Background Art
[0002] A video coding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding).
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In such encoding and decoding technologies, further improvements are required.
[0005] Therefore, an object of the present disclosure is to provide an encoding device and the like that can achieve further improvements.
Means for Solving the Problems
[0006] An encoding device according to one aspect of the present disclosure is an encoding device that encodes an image block, and includes a circuit and a memory. The circuit performs a primary conversion from a residual of the image block to primary coefficients in a primary space using the memory, determines whether to apply a secondary conversion to the image block, (i) when the secondary conversion is not applied, calculates quantized primary coefficients by performing a first quantization on the primary coefficients, and (ii) when the secondary conversion is applied, performs a secondary conversion from the primary coefficients to secondary coefficients in a secondary space different from the primary space, and calculates quantized secondary coefficients by performing a second quantization different from the first quantization on the secondary coefficients.
[0007] Note that these general or specific aspects may be implemented in a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, method, integrated circuit, computer program, and recording medium.
Advantages of the Invention
[0008] The present disclosure can provide an encoding device and the like that can achieve further improvement.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
DETAILED DESCRIPTION OF THE INVENTION
[0010] (Knowledge underlying the present disclosure) In the next-generation moving image compression standard, in order to further remove spatial redundancy, secondary conversion for coefficients obtained by primary conversion of residuals is being studied. Even when such secondary conversion is performed, it is expected to improve the encoding efficiency while suppressing a decrease in subjective image quality.
[0011] Therefore, an encoding apparatus according to an aspect of the present disclosure is an encoding apparatus that encodes an encoding target block of an image, and includes a circuit and a memory. The circuit performs primary conversion from the residual of the encoding target block to primary coefficients using the memory, determines whether to apply secondary conversion to the encoding target block, (i) when the secondary conversion is not applied, calculates quantized primary coefficients by performing first quantization on the primary coefficients, (ii) when the secondary conversion is applied, performs secondary conversion from the primary coefficients to secondary coefficients, calculates quantized secondary coefficients by performing second quantization different from the first quantization on the secondary coefficients, and generates an encoded bit stream by encoding the quantized primary coefficients or the quantized secondary coefficients.
[0012] According to this, different quantization can be performed according to the application / non-application of the secondary transformation to the block to be encoded. The secondary coefficients obtained by the secondary transformation from the primary coefficients represented in the first space will be represented in a secondary space that is not the primary space. Therefore, even if the quantization for the primary coefficients is applied to the secondary coefficients, it is difficult to improve the coding efficiency while suppressing the degradation of the subjective image quality. For example, quantization for reducing the loss of components in the low-frequency band to suppress the degradation of the subjective image quality and increasing the loss of components in the high-frequency band to improve the coding efficiency is different between the primary space and the secondary space. Therefore, by performing different quantization according to the application / non-application of the secondary transformation to the block to be encoded, it is possible to improve the coding efficiency while suppressing the degradation of the subjective image quality as compared with the case where common quantization is performed.
[0013] Also, in the encoding device according to one aspect of the present disclosure, for example, the first quantization may be weighted quantization using a first quantization matrix, and the second quantization may be weighted quantization using a second quantization matrix different from the first quantization matrix.
[0014] According to this, as the first quantization, weighted quantization using a first quantization matrix can be performed. Further, as the second quantization, weighted quantization using a second quantization matrix different from the first quantization matrix can be performed. Therefore, the first quantization matrix corresponding to the primary space can be used for the quantization of the primary coefficients, and the second quantization matrix corresponding to the secondary space can be used for the quantization of the secondary coefficients. Therefore, in both the application and non-application of the secondary transformation, it is possible to improve the coding efficiency while suppressing the degradation of the subjective image quality.
[0015] Also, in the encoding device according to one aspect of the present disclosure, for example, the circuit may further write the first quantization matrix and the second quantization matrix into the encoded bit stream.
[0016] According to this, the first quantization matrix and the second quantization matrix can be included in the encoded bit stream. Therefore, the first quantization matrix and the second quantization matrix can be adaptively determined according to the original image, and it is possible to improve the encoding efficiency while suppressing a decrease in subjective image quality.
[0017] Also, in the encoding apparatus according to one aspect of the present disclosure, for example, the primary coefficients include one or more first primary coefficients and one or more second primary coefficients, the secondary conversion is applied to the one or more first primary coefficients and not applied to the one or more second primary coefficients, the second quantization matrix includes one or more first component values corresponding to the one or more first primary coefficients and one or more second component values corresponding to the one or more second primary coefficients, each of the one or more second component values of the second quantization matrix coincides with the corresponding component value of the first quantization matrix, and in writing the second quantization matrix, only the one or more first component values among the one or more first component values and the one or more second component values may be written into the encoded bit stream.
[0018] According to this, each of the one or more second component values of the second quantization matrix can be made to coincide with the corresponding component value of the first quantization matrix. Therefore, it is not necessary to write the one or more second component values of the second quantization matrix into the encoded bit stream, and the encoding efficiency can be improved.
[0019] Also, in the encoding apparatus according to one aspect of the present disclosure, for example, in the secondary conversion, a plurality of predetermined bases are selectively used, the encoded bit stream includes a plurality of second quantization matrices corresponding to the plurality of bases, and in the second quantization, a second quantization matrix corresponding to the base used in the secondary conversion may be selected from among the plurality of second quantization matrices.
[0020] According to this, second quantization can be performed using a second quantization matrix corresponding to the basis used in the second transformation. The characteristics of the quadratic space representing the quadratic coefficients differ depending on the basis used in the second transformation. Therefore, by performing second quantization using the second quantization matrix corresponding to the basis used in the second transformation, second quantization can be performed using a quantization matrix corresponding more to the quadratic space, and the coding efficiency can be improved while suppressing a decrease in subjective image quality.
[0021] Also, in an encoding device according to an aspect of the present disclosure, for example, the first quantization matrix and the second quantization matrix may be predefined in a standardization standard.
[0022] According to this, the first quantization matrix and the second quantization matrix are predefined in a standardization standard. Therefore, the first quantization matrix and the second quantization matrix may not be included in the encoded bit stream, and the amount of code for the first quantization matrix and the second quantization matrix can be reduced.
[0023] Also, in an encoding device according to an aspect of the present disclosure, for example, the circuit may further derive the second quantization matrix from the first quantization matrix.
[0024] According to this, the second quantization matrix can be derived from the first quantization matrix. Therefore, since it is not necessary to transmit the second quantization matrix to the decoding device, the coding efficiency can be improved.
[0025] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the primary coefficient includes one or more first primary coefficients and one or more second primary coefficients, the secondary transformation is applied to the one or more first primary coefficients and not applied to the one or more second primary coefficients, the second quantization matrix includes one or more first component values corresponding to the one or more first primary coefficients and one or more second component values corresponding to the one or more second primary coefficients, each of the one or more second component values of the second quantization matrix coincides with the corresponding component value of the first quantization matrix, and in the derivation of the second quantization matrix, the one or more first component values of the second quantization matrix may be derived from the first quantization matrix.
[0026] According to this, each of the one or more second component values of the second quantization matrix can be made to coincide with the corresponding component value of the first quantization matrix. Therefore, it is not necessary to derive the one or more second component values of the second quantization matrix from the first quantization matrix, and the processing load can be reduced.
[0027] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the second quantization matrix may be derived by applying the secondary transformation to the first quantization matrix.
[0028] According to this, the second quantization matrix can be derived by applying the secondary transformation to the first quantization matrix. Therefore, the first quantization matrix corresponding to the primary space can be converted into the second quantization matrix corresponding to the secondary space, and the encoding efficiency can be improved while suppressing a decrease in subjective image quality.
[0029] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the circuit further derives a third quantization matrix from the first quantization matrix, and each component value of the third quantization matrix is larger as the corresponding component value of the first quantization matrix is smaller. By applying the secondary transformation to the third quantization matrix, a fourth quantization matrix is derived, and a fifth quantization matrix is derived from the fourth quantization matrix as the second quantization matrix. Each component value of the fifth quantization matrix may be larger as the corresponding component value of the fourth quantization matrix is smaller.
[0030] According to this, it is possible to reduce the influence of the rounding error during the secondary transformation on the components having relatively small values included in the first quantization matrix. That is, it is possible to reduce the influence of the rounding error on the values of the components applied to the coefficients for which it is desired to reduce the loss in order to suppress the degradation of the subjective image quality. Therefore, it is possible to further suppress the degradation of the subjective image quality.
[0031] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, each component value of the third quantization matrix may be the reciprocal of the corresponding component value of the first quantization matrix, and each component value of the fifth quantization matrix may be the reciprocal of the corresponding component value of the fourth quantization matrix.
[0032] According to this, the reciprocal of the corresponding component value of the first quantization matrix / fourth quantization matrix can be used as each component value of the third quantization matrix / fifth quantization matrix. Therefore, the component values can be derived by simple calculation, and the processing load or processing time for deriving the second quantization matrix can be reduced.
[0033] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the first quantization may be weighted quantization using a quantization matrix, and the second quantization may be unweighted quantization without using a quantization matrix.
[0034] According to this, as second quantization, unweighted quantization without using a quantization matrix can be used. Therefore, while preventing a deterioration in subjective image quality due to using the first quantization matrix for first quantization in second quantization, the amount of code or the derivation process of the quantization matrix for second quantization can be omitted.
[0035] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the first quantization is weighted quantization using a first quantization matrix, and in the secondary conversion, (i) by multiplying each of the primary coefficients by the corresponding component value of a weight matrix, a weighted primary coefficient is calculated, and (ii) the weighted primary coefficient is converted into a secondary coefficient. In the second quantization, each of the secondary coefficients may be divided by a quantization step common to the secondary coefficients.
[0036] According to this, a weighted primary coefficient can be calculated by multiplying each of the primary coefficients by the corresponding component value of the weight matrix. Weighting related to quantization can be performed on the primary coefficients before the secondary conversion. Therefore, when the secondary conversion is applied, quantization equivalent to weighted quantization can be performed without newly preparing a quantization matrix corresponding to the secondary space. As a result, it is possible to improve the encoding efficiency while suppressing a deterioration in subjective image quality.
[0037] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the circuit may further derive the weight matrix from the first quantization matrix.
[0038] According to this, the weight matrix can be derived from the first quantization matrix. Therefore, the amount of code for the weight matrix can be reduced, and it is possible to improve the encoding efficiency while suppressing a deterioration in subjective image quality.
[0039] Also, in an encoding apparatus according to an aspect of the present disclosure, for example, the circuit may further derive the common quantization step from the quantization parameter for the encoding target block.
[0040] According to this, a quantization step common to the secondary coefficients of the block to be encoded can be derived from the quantization parameter. Therefore, it is not necessary to include new information in the encoded stream for the common quantization step, and the amount of code for the common quantization step can be reduced.
[0041] An encoding method according to an aspect of the present disclosure is an encoding method for encoding a block to be encoded in an image, the method including performing a primary conversion from a residual of the block to be encoded to a primary coefficient, determining whether to apply a secondary conversion to the block to be encoded, (i) when the secondary conversion is not applied, calculating a quantized primary coefficient by performing a first quantization on the primary coefficient, (ii) when the secondary conversion is applied, performing a secondary conversion from the primary coefficient to a secondary coefficient, calculating a quantized secondary coefficient by performing a second quantization different from the first quantization on the secondary coefficient, and generating an encoded bit stream by encoding the quantized primary coefficient or the quantized secondary coefficient.
[0042] According to this, the same effect as the above encoding apparatus can be realized.
[0043] An encoding apparatus according to an aspect of the present disclosure is an encoding apparatus for encoding a block to be encoded in an image, the apparatus including a circuit and a memory, the circuit using the memory to perform a primary conversion of a residual of the block to be encoded to a primary coefficient, determining whether to apply a secondary conversion to the block to be encoded, (i) when the secondary conversion is not applied, calculating a first quantized primary coefficient by performing a first quantization on the primary coefficient, (ii) when the secondary conversion is applied, calculating a second quantized primary coefficient by performing a second quantization on the primary coefficient, performing a secondary conversion from the second quantized primary coefficient to a quantized secondary coefficient, and generating an encoded bit stream by encoding the first quantized primary coefficient or the quantized secondary coefficient.
[0044] According to this, quantization can be performed before the second conversion. Therefore, when the second conversion process is lossless, the second conversion can be removed from the loop of the prediction process. Thus, the load on the processing pipeline can be reduced. Further, by performing quantization before the second conversion, it is not necessary to separate the first quantization matrix and the second quantization matrix, so that the process can be simplified.
[0045] An encoding method according to an aspect of the present disclosure is an encoding method for encoding an encoding target block of an image, comprising: performing a first conversion on a residual of the encoding target block into first-order coefficients; determining whether to apply a second conversion to the encoding target block; (i) when the second conversion is not applied, calculating first-quantized first-order coefficients by performing first quantization on the first-order coefficients; (ii) when the second conversion is applied, calculating second-quantized first-order coefficients by performing second quantization on the first-order coefficients; performing a second conversion from the second-quantized first-order coefficients to quantized second-order coefficients; and generating an encoded bitstream by encoding the first-quantized first-order coefficients or the quantized second-order coefficients.
[0046] According to this, the same effect as the above encoding device can be realized.
[0047] A decoding device according to an aspect of the present disclosure is a decoding device for decoding a decoding target block of an image, comprising a circuit and a memory, wherein the circuit uses the memory to decode the quantized coefficients of the decoding target block from an encoded bitstream, determines whether to apply an inverse second conversion to the decoding target block, calculates first-order coefficients by performing first inverse quantization on the quantized coefficients when the inverse second conversion is not applied, performs an inverse first conversion from the first-order coefficients to the residual of the decoding target block, calculates second-order coefficients by performing a second inverse quantization different from the first inverse quantization on the quantized coefficients when the inverse second conversion is applied, performs an inverse second conversion from the second-order coefficients to first-order coefficients, and performs an inverse first conversion from the first-order coefficients to the residual of the decoding target block.
[0048] According to this, different inverse quantization can be performed according to whether inverse second-order transformation is applied or not to the block to be decoded. The second-order coefficients obtained by second-order transformation from the first-order coefficients represented in the first space will be represented in the second space which is not the first space. Therefore, even if the inverse quantization for the first-order coefficients is applied to the second-order coefficients, it is difficult to improve the coding efficiency while suppressing the degradation of subjective image quality. For example, quantization for reducing the loss of components in the low-frequency band to suppress the degradation of subjective image quality and increasing the loss of components in the high-frequency band to improve the coding efficiency is different between the first space and the second space. Thus, by performing different inverse quantization according to whether inverse second-order transformation is applied or not to the block to be decoded, it is possible to improve the coding efficiency while suppressing the degradation of subjective image quality as compared with the case where common inverse quantization is performed.
[0049] Also, in the decoding apparatus according to one aspect of the present disclosure, for example, the first inverse quantization may be weighted inverse quantization using a first quantization matrix, and the second inverse quantization may be weighted inverse quantization using a second quantization matrix different from the first quantization matrix.
[0050] According to this, as the first inverse quantization, weighted inverse quantization using a first quantization matrix can be performed. Further, as the second inverse quantization, weighted inverse quantization using a second quantization matrix different from the first quantization matrix can be performed. Therefore, the first quantization matrix corresponding to the first space can be used for the inverse quantization for the first-order coefficients, and the second quantization matrix corresponding to the second space can be used for the inverse quantization for the second-order coefficients. Therefore, in both the application and non-application of the inverse second-order transformation, it is possible to improve the coding efficiency while suppressing the degradation of subjective image quality.
[0051] Also, in the decoding apparatus according to one aspect of the present disclosure, for example, the circuit may further decode the first quantization matrix and the second quantization matrix from the encoded bit stream.
[0052] According to this, the first quantization matrix and the second quantization matrix can be included in the encoded bitstream. Therefore, the first quantization matrix and the second quantization matrix can be adaptively determined according to the original image, and it is possible to improve the encoding efficiency while suppressing the degradation of the subjective image quality.
[0053] Also, in the decoding device according to one aspect of the present disclosure, for example, the quadratic coefficient includes one or more first quadratic coefficients and one or more second quadratic coefficients, and the inverse quadratic transform is applied to the one or more first quadratic coefficients and not applied to the one or more second quadratic coefficients. The second quantization matrix includes one or more first component values corresponding to the one or more first quadratic coefficients and one or more second component values corresponding to the one or more second quadratic coefficients. Each of the one or more second component values of the second quantization matrix matches the corresponding component value of the first quantization matrix. In the decoding of the second quantization matrix, only the one or more first component positions among the one or more first component values and the one or more second component values may be decoded from the encoded bitstream.
[0054] According to this, each of the one or more second component values of the second quantization matrix can be made to match the corresponding component value of the first quantization matrix. Therefore, it is not necessary to decode the one or more second component values of the second quantization matrix from the encoded bitstream, and the encoding efficiency can be improved.
[0055] Also, in the decoding device according to one aspect of the present disclosure, for example, in the inverse quadratic transform, a plurality of predetermined bases are selectively used, the encoded bitstream includes a plurality of second quantization matrices corresponding to the plurality of bases, and in the second inverse quantization, a second quantization matrix corresponding to the base used in the inverse quadratic transform may be selected from among the plurality of second quantization matrices.
[0056] According to this, depending on the basis used for the inverse quadratic transformation, the characteristics of the quadratic space representing the quadratic coefficients are different. Therefore, by performing the second inverse quantization using the second quantization matrix corresponding to the basis used for the inverse quadratic transformation, it is possible to perform the second inverse quantization using a quantization matrix corresponding to the quadratic space, and it is possible to improve the coding efficiency while suppressing the degradation of the subjective image quality.
[0057] Also, in the decoding device according to one aspect of the present disclosure, for example, the first quantization matrix and the second quantization matrix may be predefined in a standardization standard.
[0058] According to this, the first quantization matrix and the second quantization matrix are predefined in the standardization standard. Therefore, the first quantization matrix and the second quantization matrix may not be included in the coded bit stream, and the amount of bits for the first quantization matrix and the second quantization matrix can be reduced.
[0059] Also, in the decoding device according to one aspect of the present disclosure, for example, the circuit may further derive the second quantization matrix from the first quantization matrix.
[0060] According to this, the second quantization matrix can be derived from the first quantization matrix. Therefore, since it is not necessary to receive the second quantization matrix from the coding device, the coding efficiency can be improved.
[0061] Also, in the decoding device according to one aspect of the present disclosure, for example, the secondary coefficient includes one or more first secondary coefficients and one or more second secondary coefficients, the inverse secondary transformation is applied to the one or more first secondary coefficients and not applied to the one or more second secondary coefficients, the second quantization matrix includes one or more first component values corresponding to the one or more first secondary coefficients and one or more second component values corresponding to the one or more second secondary coefficients, each of the one or more second component values of the second quantization matrix coincides with the corresponding component value of the first quantization matrix, and in the derivation of the second quantization matrix, the one or more first component positions may be derived from the first quantization matrix.
[0062] According to this, each of the one or more second component values of the second quantization matrix can be made to coincide with the corresponding component value of the first quantization matrix. Therefore, it is not necessary to derive the one or more second component values of the second quantization matrix from the first quantization matrix, and the processing load can be reduced.
[0063] Also, in the decoding device according to one aspect of the present disclosure, for example, the second quantization matrix may be derived by applying a secondary transformation to the first quantization matrix.
[0064] According to this, the second quantization matrix can be derived by applying a secondary transformation to the first quantization matrix. Therefore, the first quantization matrix corresponding to the primary space can be converted into the second quantization matrix corresponding to the secondary space, and the coding efficiency can be improved while suppressing a decrease in subjective image quality.
[0065] Also, in a decoding apparatus according to an aspect of the present disclosure, for example, the circuit further derives a third quantization matrix from the first quantization matrix, and each component value of the third quantization matrix is larger as the corresponding component value of the first quantization matrix is smaller. By applying a secondary transformation to the third quantization matrix, a fourth quantization matrix is derived, and a fifth quantization matrix is derived from the fourth quantization matrix as the second quantization matrix. Each component value of the fifth quantization matrix may be larger as the corresponding component value of the fourth quantization matrix is smaller.
[0066] According to this, it is possible to reduce the influence of rounding errors during the secondary transformation on components having relatively small values included in the first quantization matrix. That is, it is possible to reduce the influence of rounding errors on the values of the components applied to the coefficients for which it is desired to reduce the loss in order to suppress the degradation of the subjective image quality. Therefore, it is possible to further suppress the degradation of the subjective image quality.
[0067] Also, in a decoding apparatus according to an aspect of the present disclosure, for example, each component value of the third quantization matrix may be the reciprocal of the corresponding component value of the first quantization matrix, and each component value of the fifth quantization matrix may be the reciprocal of the corresponding component value of the fourth quantization matrix.
[0068] According to this, the reciprocal of the corresponding component value of the first quantization matrix / fourth quantization matrix can be used as each component value of the third quantization matrix / fifth quantization matrix. Therefore, the component values can be derived by simple calculation, and the processing load or processing time for deriving the second quantization matrix can be reduced.
[0069] Also, in a decoding apparatus according to an aspect of the present disclosure, for example, the first inverse quantization may be weighted inverse quantization using a quantization matrix, and the second inverse quantization may be unweighted inverse quantization without using a quantization matrix.
[0070] According to this, as the second inverse quantization, unweighted inverse quantization without using a quantization matrix can be used. Therefore, while preventing a decrease in subjective image quality caused by using the first quantization matrix for the first inverse quantization in the second inverse quantization, the encoding or derivation process of the quantization matrix for the second quantization can be omitted.
[0071] Also, in a decoding apparatus according to an aspect of the present disclosure, for example, the first inverse quantization is weighted inverse quantization using a first quantization matrix, and in the second inverse quantization, a common quantization step for the quantization coefficients is multiplied by each of the quantization coefficients to calculate the secondary coefficients. In the inverse secondary transformation, (i) the secondary coefficients may be inversely transformed into weighted primary coefficients, and (ii) the primary coefficients may be calculated by dividing each of the weighted primary coefficients by the corresponding component value of the weight matrix.
[0072] According to this, the primary coefficients can be calculated by dividing each of the weighted primary coefficients by the corresponding component value of the weight matrix. That is, weighting related to quantization can be performed on the primary coefficients before the secondary transformation. Therefore, when the inverse secondary transformation is applied, inverse quantization equivalent to weighted inverse quantization can be performed without newly preparing a quantization matrix corresponding to the secondary space. As a result, it is possible to improve the coding efficiency while suppressing a decrease in subjective image quality.
[0073] Also, in a decoding apparatus according to an aspect of the present disclosure, for example, the circuit may further derive the weight matrix from the first quantization matrix.
[0074] According to this, the weight matrix can be derived from the first quantization matrix. Therefore, the amount of code for the weight matrix can be reduced, and it is possible to improve the coding efficiency while suppressing a decrease in subjective image quality.
[0075] Also, in the decoding device according to one aspect of the present disclosure, for example, the circuit may further derive the common quantization step from the quantization parameter for the block to be decoded.
[0076] According to this, the common quantization step for the secondary coefficients of the block to be decoded can be derived from the quantization parameter. Therefore, it is not necessary to include new information for the common quantization step in the coded stream, and the amount of code for the common quantization step can be reduced.
[0077] A decoding method according to one aspect of the present disclosure is a decoding method for decoding a block to be decoded in an image, which decodes the quantization coefficients of the block to be decoded from a coded bit stream, determines whether to apply an inverse secondary transform to the block to be decoded, and when the inverse secondary transform is not applied, calculates primary coefficients by performing a first inverse quantization on the quantization coefficients, performs an inverse primary transform from the primary coefficients to the residual of the block to be decoded, and when the inverse secondary transform is applied, calculates secondary coefficients by performing a second inverse quantization different from the first inverse quantization on the quantization coefficients, performs an inverse secondary transform from the secondary coefficients to primary coefficients, and performs an inverse primary transform from the primary coefficients to the residual of the block to be decoded.
[0078] According to this, the same effect as the above decoding device can be achieved.
[0079] A decoding apparatus according to one aspect of the present disclosure is a decoding apparatus that decodes a decoding target block of an image, and includes a circuit and a memory. The circuit uses the memory to decode quantization coefficients of the decoding target block from a coded bit stream, determines whether to apply an inverse secondary transformation to the decoding target block, and when the inverse secondary transformation is not applied, calculates primary coefficients by performing a first inverse quantization on the quantization coefficients, performs an inverse primary transformation from the primary coefficients to a residual of the decoding target block, and when the inverse secondary transformation is applied, performs an inverse secondary transformation from the quantization coefficients to quantization primary coefficients, calculates primary coefficients by performing a second inverse quantization on the quantization primary coefficients, and performs an inverse primary transformation from the primary coefficients to a residual of the decoding target block.
[0080] According to this, in an encoding apparatus, quantization can be performed before secondary transformation. Therefore, when the secondary transformation process is lossless, the secondary transformation can be removed from the loop of the prediction process. Thus, the load on the processing pipeline can be reduced. Further, by performing quantization before secondary transformation, it is not necessary to divide the first quantization matrix and the second quantization matrix, so that the process can be simplified.
[0081] A decoding method according to one aspect of the present disclosure is a decoding method that decodes a decoding target block of an image, decodes quantization coefficients of the decoding target block from a coded bit stream, determines whether to apply an inverse secondary transformation to the decoding target block, and when the inverse secondary transformation is not applied, calculates primary coefficients by performing a first inverse quantization on the quantization coefficients, performs an inverse primary transformation from the primary coefficients to a residual of the decoding target block, and when the inverse secondary transformation is applied, performs an inverse secondary transformation from the quantization coefficients to quantization primary coefficients, calculates primary coefficients by performing a second inverse quantization on the quantization primary coefficients, and performs an inverse primary transformation from the primary coefficients to a residual of the decoding target block.
[0082] According to this, the same effects as those of the above decoding apparatus can be achieved.
[0083] Note that these general or specific aspects may be implemented in a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0084] Hereinafter, embodiments will be specifically described with reference to the drawings.
[0085] Note that all of the embodiments described below show general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Also, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.
[0086] (Embodiment 1) First, as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure to be described later are applicable, an overview of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.
[0087] When applying the processes and / or configurations described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.
[0088] (1) Replacing, in the encoding device or the decoding device of Embodiment 1, among the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure with the component described in each aspect of the present disclosure (2) With respect to the encoding device or decoding device of Embodiment 1, after making any changes such as addition, replacement, or deletion of functions or processes performed by some of the plurality of components constituting the encoding device or decoding device, replace the components corresponding to the components described in each aspect of the present disclosure with the components described in each aspect of the present disclosure. (3) With respect to the method performed by the encoding device or decoding device of Embodiment 1, after making any changes such as addition of processes and / or replacement or deletion of some of the plurality of processes included in the method, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure. (4) Implement some of the plurality of components constituting the encoding device or decoding device of Embodiment 1 in combination with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components performing a part of the processes performed by the components described in each aspect of the present disclosure. (5) Implement components having a part of the functions provided by some of the plurality of components constituting the encoding device or decoding device of Embodiment 1, or components performing a part of the processes performed by some of the plurality of components constituting the encoding device or decoding device of Embodiment 1 in combination with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components performing a part of the processes performed by the components described in each aspect of the present disclosure. (6) With respect to the method performed by the encoding device or decoding device of Embodiment 1, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure among the plurality of processes included in the method. (7) Implement some of the plurality of processes included in the method performed by the encoding device or decoding device of Embodiment 1 in combination with the processes described in each aspect of the present disclosure.
[0089] Note that the implementation methods of the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or the moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Further, the processes and / or configurations described in different aspects may be implemented in combination.
[0090] [Outline of the Encoding Device] First, the outline of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.
[0091] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0092] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0093] The components included in the encoding device 100 will be described below.
[0094] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These blocks of fixed size are sometimes called Coding Tree Units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These blocks of variable size are sometimes called Coding Units (CUs), Prediction Units (PUs), or Transformation Units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks within the picture may be the processing units of CUs, PUs, and TUs.
[0095] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, the solid lines represent block boundaries by quadtree block splitting, and the dashed lines represent block boundaries by binary tree block splitting.
[0096] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).
[0097] The upper-left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper-left 64x64 block is split into two 16x64 blocks 11, 12 and a 32x64 block 13.
[0098] The upper right 64x64 block is horizontally divided into two rectangular 64x32 blocks 14 and 15 (binary tree block division).
[0099] The lower left 64x64 block is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary tree block division). As a result, the lower left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.
[0100] The lower right 64x64 block 23 is not divided.
[0101] As described above, in FIG. 2, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.
[0102] Note that in FIG. 2, one block was divided into four or two blocks (quad-tree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). A division including such a ternary tree block division is sometimes called MBT (multi type tree) division.
[0103] [Subtraction unit] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in block units divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also referred to as the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.
[0104] The original signal is the input signal of the encoding device 100 and is a signal representing the image of each picture constituting the moving image (for example, a luminance signal and two chrominance signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0105] [Conversion Unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0106] Note that the conversion unit 106 may adaptively select a conversion type from among a plurality of conversion types and convert the prediction error into conversion coefficients using a conversion basis function corresponding to the selected conversion type. Such a conversion is sometimes referred to as an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).
[0107] The plurality of conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the conversion basis functions corresponding to each conversion type. In FIG. 3, N indicates the number of input pixels. The selection of the conversion type from among these plurality of conversion types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.
[0108] Information indicating whether to apply such EMT or AMT (for example, called an AMT flag) and information indicating the selected conversion type are signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0109] Also, the conversion unit 106 may re-convert the conversion coefficients (conversion results). Such re-conversion may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (for example, 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0110] Here, separable conversion is a method of performing multiple conversions by separating in each direction by the number of dimensions of the input, and non-separable conversion is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing conversion collectively.
[0111] For example, as an example of non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.
[0112] Also, after similarly regarding a 4×4 input block as an array having 16 elements, a method of performing a plurality of Givens rotations on the array (Hypercube Givens Transform) is also an example of a non-separable transform.
[0113] [Quantization unit] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the scanned transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0114] The predetermined order is an order for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).
[0115] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. That is, as the value of the quantization parameter increases, the quantization error increases.
[0116] [Entropy encoding unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable-length encoding on the quantization coefficients that are input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantization coefficients and performs arithmetic encoding on the binary signal, for example.
[0117] [Inverse quantization unit] The inverse quantization unit 112 inverse quantizes the quantization coefficients which are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0118] [Inverse transform unit] The inverse transform unit 114 restores the prediction error by inverse-transforming the transform coefficients which are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0119] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes quantization error.
[0120] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error which is the input from the inverse transform unit 114 and the prediction sample which is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.
[0121] [Block memory] The block memory 118 is a storage unit for storing blocks within the picture to be coded (hereinafter referred to as the current picture) which are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.
[0122] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0123] In the ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 sub-block within the current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.
[0124] Specifically, first, sub-blocks (for example, 2x2 sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (for example, 15 or 25 classes).
[0125] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantizing the addition result.
[0126] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.
[0127] For the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. The information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0128] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance difference, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or the CU level. Note that the signaling of the information indicating the on / off of ALF is not necessarily limited to the picture level or the CU level and may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0129] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets is not necessarily limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0130] [Frame Memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction and may also be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0131] [Intra Prediction Unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also referred to as in-picture prediction) of the current block with reference to the blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0132] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0133] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0134] The plurality of directional prediction modes includes, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.
[0135] In addition, in the intra prediction of the color difference block, a luminance block may be referred to. That is, based on the luminance component of the current block, the color difference component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra prediction mode of the color difference block that refers to a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the color difference block.
[0136] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. Such intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called the PDPC flag) is signaled at, for example, the CU level. Note that the signaling of this information is not necessarily limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0137] [Inter Prediction Unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0138] The motion information used for motion compensation is signaled. For the signaling of motion vectors, a motion vector predictor may be used. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0139] Note that, not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks may be used to generate an inter-prediction signal. Specifically, an inter-prediction signal may be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter-prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0140] In such an OBMC mode, information indicating the size of sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be at other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).
[0141] The OBMC mode will be described in more detail. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process by OBMC processing.
[0142] First, a predicted image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.
[0143] Next, apply the motion vector (MV_L) of the encoded left adjacent block to the block to be encoded to obtain a predicted image (Pred_L), and perform the first correction of the predicted image by weighted addition of the predicted image and Pred_L.
[0144] Similarly, apply the motion vector (MV_U) of the encoded upper adjacent block to the block to be encoded to obtain a predicted image (Pred_U), and perform the second correction of the predicted image by weighted addition of the predicted image after the first correction and Pred_U, and use it as the final predicted image.
[0145] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described, but it is also possible to adopt a configuration in which corrections are performed more times than two stages using the right adjacent block or the lower adjacent block.
[0146] Note that the region for addition may be only a partial region near the block boundary, rather than the pixel region of the entire block.
[0147] Here, the predicted image correction process from a single reference picture has been described, but the same applies to the case of correcting the predicted image from multiple reference pictures. After obtaining the predicted images corrected from each reference picture, the final predicted image is obtained by further adding the obtained predicted images together.
[0148] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.
[0149] As a method for determining whether to apply OBMC processing, for example, there is a method of using an obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, a value 1 is set as the obmc_flag, and OBMC processing is applied for encoding. If it does not belong to a region with complex motion, a value 0 is set as the obmc_flag, and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding the obmc_flag described in the stream, decoding is performed while switching whether to apply OBMC processing according to the value.
[0150] Note that the motion information may be derived on the decoding device side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also, for example, the motion information may be derived by performing motion search on the decoding device side. In this case, the motion search is performed without using the pixel values of the current block.
[0151] Here, the mode of performing motion search on the decoding device side will be described. This mode of performing motion search on the decoding device side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.
[0152] An example of FRUC processing is shown in FIG. 5D. First, by referring to the motion vectors of encoded blocks spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0153] Then, based on the motion vectors of the selected candidates, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, a motion vector for the current block may be derived. That is, search is performed in the same way for the region around the best candidate MV, and if there is an MV with a better evaluation value, the best candidate MV may be updated to the said MV and used as the final MV of the current block. Note that it is also possible to adopt a configuration in which such processing is not performed.
[0154] When processing is performed in units of sub-blocks, it may be the same processing.
[0155] Note that the evaluation value is calculated by obtaining the difference value of the reconstructed image by pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that in addition to the difference value, other information may be used to calculate the evaluation value.
[0156] As the pattern matching, first pattern matching or second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.
[0157] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the above-mentioned candidate, the region in another reference picture along the motion trajectory of the current block is used.
[0158] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is preferable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.
[0159] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0160] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidate.
[0161] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that most closely matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or either of the left and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.
[0162] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the pattern matching method (the first pattern matching or the second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0163] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode may be called the BIO (bi-directional optical flow) mode.
[0164] FIG. 8 is a diagram for explaining a model assuming a uniform linear motion. In FIG. 8, (vx, vy) indicates a velocity vector, and τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.
[0165] At this time, under the assumption of the uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (1) holds.
[0166]
Equation
[0167] Here, I(k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the block-unit motion vectors obtained from the merge list, etc. are corrected in pixel units.
[0168] Note that the motion vectors may be derived on the decoder side by a method different from the derivation of the motion vectors based on the model assuming uniform linear motion. For example, the motion vectors may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0169] Here, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.
[0170] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, a current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of the adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (vx, vy) of each sub-block within the current block is derived by the following equation (2).
[0171]
Equation
[0172] Here, x and y indicate the horizontal position and the vertical position of the sub-block, respectively, and w indicates a predetermined weight coefficient.
[0173] Such an affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right control points are different. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode is not necessarily limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0174] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.
[0175] Here, an example of deriving the motion vector of the picture to be coded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.
[0176] First, a prediction MV list in which candidates for the prediction MV are registered is generated. Examples of candidates for the prediction MV include a spatial adjacent prediction MV which is an MV of a plurality of coded blocks located spatially adjacent to the block to be coded, a temporal adjacent prediction MV which is an MV of a nearby block obtained by projecting the position of the block to be coded in the coded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV which is an MV with a value of zero.
[0177] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be coded.
[0178] Furthermore, in the variable length coding unit, a merge_idx which is a signal indicating which prediction MV is selected is described in the stream and coded.
[0179] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are just an example, and the number may be different from that in the figure, or the configuration may not include some of the types of prediction MVs in the figure, or the configuration may include prediction MVs other than the types of prediction MVs in the figure.
[0180] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be coded derived in the merge mode.
[0181] Here, an example of determining the MV using the DMVR process will be described.
[0182] FIG. 9C is a conceptual diagram for explaining the outline of the DMVR process.
[0183] First, using the optimal MVP set for the processing target block as a candidate MV, according to the candidate MV, reference pixels are respectively obtained from the first reference picture which is the processed picture in the L0 direction and the second reference picture which is the processed picture in the L1 direction, and a template is generated by taking the average of each reference pixel.
[0184] Next, using the template, the peripheral areas of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, and the MV value, etc.
[0185] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.
[0186] Note that even if it is not the processing itself described here, other processing may be used as long as it is a process capable of searching the periphery of the candidate MV to derive the final MV.
[0187] Here, the mode of generating a predicted image using the LIC process will be described.
[0188] FIG. 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction process by the LIC process.
[0189] First, an MV for obtaining a reference image corresponding to the block to be encoded is derived from the reference picture which is the encoded picture.
[0190] Next, for the block to be encoded, information indicating how the luminance values change between the reference picture and the picture to be encoded is extracted using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, and the luminance correction parameter is calculated.
[0191] By performing luminance correction processing on the reference image in the reference picture specified by the MV using the luminance correction parameter, a predicted image for the block to be encoded is generated.
[0192] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.
[0193] Also, although the process of generating a predicted image from a single reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures. Luminance correction processing is performed on the reference images obtained from each reference picture in the same manner, and then the predicted image is generated.
[0194] As a method for determining whether to apply the LIC process, for example, there is a method using a lic_flag, which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding apparatus, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding apparatus, the lic_flag described in the stream is decoded, and decoding is performed by switching whether to apply the LIC process according to the value.
[0195] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the peripheral encoded blocks selected during the derivation of the MV in the merge mode process are encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. In the case of this example, the processing in decoding is also exactly the same.
[0196] [Outline of Decoder] Next, an outline of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing a functional configuration of a decoder 200 according to Embodiment 1. The decoder 200 is a moving image / image decoder that decodes moving images / images in units of blocks.
[0197] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0198] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0199] Each component included in the decoder 200 will be described below.
[0200] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes, for example, the encoded bit stream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in block units.
[0201] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0202] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients, which is the input from the inverse quantization unit 204.
[0203] For example, when the information decoded from the encoded bit stream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.
[0204] Also, for example, when the information decoded from the encoded bit stream indicates that NSST is to be applied, the inverse transform unit 206 applies an inverse reverse transform to the transform coefficients.
[0205] [Addition Unit] The adder 208 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 206, and the prediction sample, which is the input from the predictive control unit 220. Then, the adder 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0206] [Block Memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture), which are blocks referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the adder 208.
[0207] [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device, etc.
[0208] When the information indicating the on / off of the ALF read from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.
[0209] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.
[0210] [Intra Prediction Unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0211] Note that when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0212] Also, when the information decoded from the encoded bit stream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0213] [Inter Prediction Unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vectors) decoded from the encoded bit stream, and outputs the inter prediction signal to the prediction control unit 220.
[0214] Note that when the information decoded from the encoded bit stream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.
[0215] Also, when the information decoded from the encoded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.
[0216] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives motion vectors in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0217] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the addition unit 208 as the prediction signal.
[0218] [Transformation processing, quantization processing, and encoding processing in the encoding device] Next, the conversion processing, quantization processing, and encoding processing performed by the conversion unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding device 100 configured as described above will be specifically described with reference to the drawings.
[0219] FIG. 11 is a flowchart showing an example of the conversion processing, quantization processing, and encoding processing in Embodiment 1. Each step shown in FIG. 11 is performed by the conversion unit 106, quantization unit 108, or entropy encoding unit 110 of the encoding device 100 according to Embodiment 1.
[0220] First, the conversion unit 106 performs a primary conversion from the residual of the block to be encoded to the primary coefficients (S101). The primary conversion is, for example, a separable conversion. Specifically, the primary conversion is, for example, DCT or DST.
[0221] Next, the conversion unit 106 determines whether to perform a second conversion on the primary coefficients (S102). That is, the conversion unit 106 determines whether to apply a second conversion to the block to be encoded. For example, the conversion unit 106 determines whether to perform a second conversion based on the cost based on the difference between the original image and the reconstructed image and / or the amount of coding. Note that the determination of whether to perform a second conversion is not limited to the determination based on such a cost. For example, whether to perform a second conversion may be determined based on the prediction mode, block size, picture type, or any combination thereof.
[0222] Here, when it is determined not to perform a second conversion (No in S102), the quantization unit 108 calculates quantized primary coefficients by performing a first quantization on the primary coefficients (S103). The first quantization is weighted quantization using a first quantization matrix. In the first quantization, quantization steps weighted by the first quantization matrix are used for each coefficient. The first quantization matrix is a weight matrix for adjusting the magnitude of the quantization step for each coefficient. The size of the first quantization matrix matches the size of the block to be encoded. That is, the number of components of the first quantization matrix matches the number of coefficients of the block to be encoded.
[0223] On the other hand, when it is determined to perform a second conversion (Yes in S102), the conversion unit 106 performs a second conversion from the primary coefficients to the secondary coefficients (S104). In the second conversion, a basis different from the basis of the first conversion is used. For example, the second conversion is a non-separable conversion. The basis used in the second conversion is, for example, predefined by a standard specification.
[0224] Thereafter, the quantization unit 108 calculates quantized secondary coefficients by performing a second quantization different from the first quantization on the secondary coefficients (S105). The second quantization different from the first quantization means that the parameters used for quantization or the quantization method are different between the first quantization and the second quantization. The parameters used for quantization are, for example, a quantization matrix or quantization parameters.
[0225] In this embodiment, the second quantization is different from the first quantization in terms of the quantization parameter used for quantization. Specifically, the second quantization is a weighted quantization that uses a second quantization matrix different from the first quantization matrix. In the second quantization, quantization steps weighted for each coefficient by the second quantization matrix are used.
[0226] The second quantization matrix is a weight matrix for adjusting the magnitude of the quantization step for each coefficient. The second quantization matrix has component values different from those of the first quantization matrix. The size of the second quantization matrix matches the size of the block to be encoded. That is, the number of components of the second quantization matrix matches the number of coefficients of the block to be encoded.
[0227] The entropy encoding unit 110 generates an encoded bit stream by entropy encoding the quantized primary coefficients or the quantized secondary coefficients (S106). At this time, the entropy encoding unit 110 writes the first quantization matrix and the second quantization matrix into the encoded bit stream.
[0228] When the secondary conversion is performed only on one or more first primary coefficients included in the primary coefficients within the block to be encoded, only one or more first component values of the second quantization matrix corresponding to the one or more first primary coefficients on which the secondary conversion is performed are written into the encoded bit stream, and for each of one or more second component values of the second quantization matrix corresponding to one or more second primary coefficients on which the secondary conversion is not performed, they are not written into the encoded bit stream and may be common with the corresponding component values of the first quantization matrix. That is, each of one or more second component values of the second quantization matrix may match the corresponding component values of the first quantization matrix. Note that the one or more first primary coefficients are, for example, coefficients in the low-frequency region, and the one or more second primary coefficients are, for example, coefficients in the high-frequency region.
[0229] Further, the entropy encoding unit 110 may write information indicating whether to apply the secondary conversion to the block to be encoded into the encoded bit stream.
[0230] Note that the positions of the first quantization matrix and the second quantization matrix in the encoded bit stream are not particularly limited. For example, as shown in FIG. 12, the first quantization matrix and the second quantization matrix may be written into (i) a video parameter set (VPS), (ii) a sequence parameter set (SPS), (iii) a picture parameter set (PPS), (iv) a slice header, or (v) video system configuration parameters.
[0231] As described above, in this embodiment, different quantization is performed depending on whether the secondary conversion is performed or not. That is, the quantization unit 108 switches between the first quantization and the second quantization based on the application / non-application of the secondary conversion to the block to be encoded. In particular, in this embodiment, different quantization matrices are used depending on whether the secondary conversion is performed or not. That is, in this embodiment, the quantization unit 108 switches between the first quantization matrix and the second quantization matrix based on the application / non-application of the secondary conversion to the block to be encoded and performs quantization.
[0232] [Decoding process, inverse quantization process, and inverse conversion process in the decoding apparatus] Next, the decoding process, inverse quantization process, and inverse conversion process performed by the entropy decoding unit 202, inverse quantization unit 204, and inverse conversion unit 206 of the decoding apparatus 200 according to this embodiment will be specifically described with reference to the drawings.
[0233] FIG. 13 is a flowchart showing an example of the decoding process, inverse quantization process, and inverse conversion process in Embodiment 1. Each step shown in FIG. 13 is performed by the entropy decoding unit 202, inverse quantization unit 204, or inverse conversion unit 206.
[0234] First, the entropy decoding unit 202 entropy-decodes the encoded quantization coefficients of the block to be decoded from the encoded bit stream (S201). The quantization coefficients decoded here are quantization primary coefficients or quantization secondary coefficients. Here, the entropy decoding unit 202 further decodes the first quantization matrix and the second quantization matrix from the encoded bit stream. When the inverse secondary transformation is performed only on one or more first secondary coefficients included in the secondary coefficients in the block to be decoded, only one or more first component values of the second quantization matrix corresponding to the one or more first secondary coefficients on which the inverse secondary transformation is performed are decoded from the encoded bit stream, and for each of one or more second component values of the second quantization matrix corresponding to one or more second secondary coefficients on which the inverse secondary transformation is not performed, it may be common with the corresponding component value of the first quantization matrix. That is, each of one or more second component values of the second quantization matrix may match the corresponding component value of the first quantization matrix. Further, the entropy decoding unit 202 may decode information indicating whether to apply the inverse secondary transformation to the block to be decoded from the encoded bit stream.
[0235] The inverse transformation unit 206 determines whether to perform the inverse secondary transformation based on the encoded bit stream (S202). That is, the inverse transformation unit 206 determines whether to apply the inverse secondary transformation to the block to be decoded. For example, the inverse transformation unit 206 determines whether to perform the inverse secondary transformation based on the information indicating whether to apply the inverse secondary transformation decoded from the encoded bit stream.
[0236] Here, when it is determined not to perform the inverse secondary transformation (No in S202), the inverse quantization unit 204 calculates the primary coefficients by performing the first inverse quantization on the decoded quantization coefficients (S203). The first inverse quantization is the inverse quantization of the first quantization in the encoding apparatus 100. In the present embodiment, the first inverse quantization is a weighted inverse quantization using the first quantization matrix.
[0237] On the other hand, when it is determined that inverse quadratic transformation is to be performed (Yes in S202), the inverse quantization unit 204 calculates the quadratic coefficients by performing a second inverse quantization different from the first inverse quantization on the decoded quantization coefficients (S204). The second inverse quantization is the inverse quantization of the second quantization in the encoding device 100. In the present embodiment, the second inverse quantization is weighted inverse quantization using a second quantization matrix. Then, the inverse transformation unit 206 performs an inverse quadratic transformation from the quadratic coefficients calculated by the second inverse quantization to the primary coefficients (S205). The inverse quadratic transformation is the inverse transformation of the quadratic transformation in the encoding device 100.
[0238] The inverse transformation unit 206 performs an inverse primary transformation from the primary coefficients obtained by the inverse quadratic transformation or the first inverse quantization to the residual of the block to be decoded (S206). The inverse primary transformation is the inverse transformation of the primary transformation in the encoding device 100.
[0239] Thus, in the present embodiment, different inverse quantizations are performed depending on whether inverse quadratic transformation is performed or not. That is, the inverse quantization unit 204 switches between the first inverse quantization and the second inverse quantization based on the application / non-application of the inverse quadratic transformation to the block to be decoded. In particular, in the present embodiment, different quantization matrices are used depending on whether inverse quadratic transformation is performed or not. That is, in the present embodiment, the inverse quantization unit 204 switches between the first quantization matrix and the second quantization matrix based on the application / non-application of the inverse quadratic transformation and performs inverse quantization.
[0240] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to the present embodiment, different quantization / inverse quantization can be performed according to the application / non-application of the secondary conversion / inverse secondary conversion to the current block. The secondary coefficients obtained by secondary conversion from the primary coefficients represented in the first space are represented in a secondary space that is not the primary space. Therefore, even if the quantization / inverse quantization for the primary coefficients is applied to the secondary coefficients, it is difficult to improve the encoding efficiency while suppressing the degradation of the subjective image quality. For example, quantization for reducing the loss of components in the low-frequency region to suppress the degradation of the subjective image quality and increasing the loss of components in the high-frequency region to improve the encoding efficiency is different between the primary space and the secondary space. Therefore, by performing different quantization / inverse quantization according to the application / non-application of the secondary conversion / inverse secondary conversion to the current block, it is possible to improve the encoding efficiency while suppressing the degradation of the subjective image quality as compared with the case where common quantization / inverse quantization is performed.
[0241] Further, according to the encoding device 100 and the decoding device 200 according to the present embodiment, as the first quantization / first inverse quantization, weighted quantization / inverse quantization using a first quantization matrix can be performed. Furthermore, as the second quantization / second inverse quantization, weighted quantization / inverse quantization using a second quantization matrix different from the first quantization matrix can be performed. Therefore, the first quantization matrix corresponding to the primary space can be used for the quantization / inverse quantization for the primary coefficients, and the second quantization matrix corresponding to the secondary space can be used for the quantization / inverse quantization for the secondary coefficients. Therefore, in both the application and non-application of the secondary conversion / inverse secondary conversion, it is possible to improve the encoding efficiency while suppressing the degradation of the subjective image quality.
[0242] Further, according to the encoding device 100 and the decoding device 200 according to the present embodiment, the first quantization matrix and the second quantization matrix can be included in the encoded bit stream. Therefore, the first quantization matrix and the second quantization matrix can be adaptively determined according to the original image, and furthermore, it is possible to improve the encoding efficiency while suppressing the degradation of the subjective image quality.
[0243] In addition, in this embodiment, the first quantization matrix and the second quantization matrix were included in the encoded bit stream, but it is not limited to this. For example, the first quantization matrix and the second quantization matrix may be transmitted from the encoding device to the decoding device separately from the encoded bit stream. Also, for example, the first quantization matrix and the second quantization matrix may be predefined in a standardization specification. At this time, the first quantization matrix and the second quantization matrix may also be called default matrices. Also, for example, the first quantization matrix and the second quantization matrix may be selected from a plurality of default matrices based on a given profile or level, etc.
[0244] As a result, the first quantization matrix and the second quantization matrix do not have to be included in the encoded bit stream, and the amount of code for the first quantization matrix and the second quantization matrix can be reduced.
[0245] In addition, in this embodiment, the case where one basis is fixedly used in the second conversion / inverse second conversion has been described, but it is not limited to this. For example, in the second conversion / inverse second conversion, a plurality of predefined bases may be selectively used. In this case, for example, the encoded bit stream may include a plurality of second quantization matrices corresponding to the plurality of bases. Then, in the second quantization / second inverse quantization, the second quantization matrix corresponding to the basis used in the second conversion / inverse second conversion may be selected from among the plurality of second quantization matrices.
[0246] Thus, second quantization / second inverse quantization can be performed using a second quantization matrix corresponding to the basis used in the second conversion / inverse second conversion. The characteristics of the second-order space representing the second-order coefficients differ depending on the basis used in the second conversion / inverse second conversion. Therefore, by performing second quantization / second inverse quantization using the second quantization matrix corresponding to the basis used in the second conversion / inverse second conversion, second quantization / second inverse quantization can be performed using a quantization matrix corresponding more closely to the second-order space, and the coding efficiency can be improved while suppressing a degradation in subjective image quality.
[0247] (Embodiment 2) Next, Embodiment 2 will be described. In this embodiment, it is different from Embodiment 1 above in that the second quantization matrix used in the second quantization is derived from the first quantization matrix used in the first quantization. Hereinafter, this embodiment will be specifically described with reference to the drawings, centering on the differences from Embodiment 1 above.
[0248] [Conversion Processing, Quantization Processing, and Encoding Processing in the Encoding Apparatus] The conversion processing, quantization processing, and encoding processing performed by the conversion unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding apparatus 100 according to Embodiment 2 will be specifically described with reference to the drawings.
[0249] FIG. 14 is a flowchart showing an example of the conversion processing, quantization processing, and encoding processing in Embodiment 2. In FIG. 14, for processes substantially the same as those in FIG. 11, the same reference numerals are given, and the description will be omitted as appropriate.
[0250] In this embodiment, after the second conversion is performed (S104), the quantization unit 108 derives the second quantization matrix from the first quantization matrix (S111). For example, the quantization unit 108 derives the second quantization matrix by applying the second conversion to the first quantization matrix. That is, the quantization unit 108 converts the first quantization matrix using the basis used for the second conversion of the block to be encoded.
[0251] In addition, when the second conversion is performed only on one or more first primary coefficients included in the primary coefficients within the block to be coded, one or more first component values of the second quantization matrix corresponding to the one or more first primary coefficients on which the second conversion is performed are derived from the first quantization matrix, and for each of one or more second component values of the second quantization matrix corresponding to one or more second primary coefficients on which the second conversion is not performed, they may be common to the corresponding component values of the first quantization matrix. That is, each of the one or more second component values of the second quantization matrix may coincide with the corresponding component value of the first quantization matrix.
[0252] The quantization unit 108 calculates quantized secondary coefficients by performing second quantization on the secondary coefficients (S105). In this second quantization, the second quantization matrix derived in step S111 is used.
[0253] The entropy coding unit 110 generates a coded bit stream by performing entropy coding on the quantized primary coefficients or the quantized secondary coefficients (S112). Further, in the present embodiment, the entropy coding unit 110 writes the first quantization matrix into the coded bit stream. Conversely, the entropy coding unit 110 does not write the second quantization matrix into the coded bit stream.
[0254] In this way, the quantization unit 108 switches between the first quantization matrix and the second quantization matrix based on the application / non-application of the second conversion to the block to be coded, and performs quantization. At this time, the quantization unit 108 derives the second quantization matrix from the first quantization matrix.
[0255] [Decoding process, inverse quantization process, and inverse conversion process in the decoding device] Next, the decoding process, inverse quantization process, and inverse conversion process performed by the entropy decoding unit 202, inverse quantization unit 204, and inverse conversion unit 206 of the decoding device 200 according to the present embodiment will be specifically described with reference to the drawings.
[0256] FIG. 15 is a flowchart showing an example of decoding processing, inverse quantization processing, and inverse transformation processing in Embodiment 2. In FIG. 15, for processes substantially the same as those in FIG. 13, the same reference numerals are given and the description thereof is omitted as appropriate.
[0257] First, the entropy decoding unit 202 decodes the encoded quantization coefficients included in the encoded bit stream (S211). At this time, the entropy decoding unit 202 deciphers the first quantization matrix from the encoded bit stream.
[0258] The inverse transformation unit 206 determines whether to perform an inverse secondary transformation based on the encoded bit stream in the same manner as in Embodiment 1 (S202). Here, when it is determined not to perform the inverse secondary transformation (No in S202), the inverse quantization unit 204 performs the first inverse quantization on the decoded quantization coefficients in the same manner as in Embodiment 1 (S203).
[0259] On the other hand, when it is determined to perform the inverse secondary transformation (Yes in S202), the inverse quantization unit 204 derives a second quantization matrix from the first quantization matrix (S212). Specifically, the inverse quantization unit 204 derives the second quantization matrix in the same manner as the encoding apparatus 100. For example, the inverse quantization unit 204 derives the second quantization matrix by applying a secondary transformation to the first quantization matrix. That is, the inverse quantization unit 204 transforms the first quantization matrix using the basis used for the secondary transformation of the block to be decoded. When the inverse secondary transformation is performed only on one or more first secondary coefficients included in the secondary coefficients within the block to be decoded, one or more first component values of the second quantization matrix corresponding to the one or more first secondary coefficients on which the inverse secondary transformation is performed are derived from the first quantization matrix, and for each of one or more second component values of the second quantization matrix corresponding to one or more second secondary coefficients on which the inverse secondary transformation is not performed, it may be common with the corresponding component value of the first quantization matrix. That is, each of one or more second component values of the second quantization matrix may coincide with the corresponding component value of the first quantization matrix.
[0260] Thereafter, the processes after step S204 are performed.
[0261] As described above, the inverse quantization unit 204 switches between the first quantization matrix and the second quantization matrix based on the application / non-application of the inverse secondary transformation to the block to be decoded, and performs inverse quantization. At this time, the inverse quantization unit 204 derives the second quantization matrix from the first quantization matrix. Therefore, even if the second quantization matrix is not included in the encoded bit stream, the decoding device 200 can perform the second inverse quantization.
[0262] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to the present embodiment, the second quantization matrix can be derived from the first quantization matrix. Therefore, since it is not necessary to transmit the second quantization matrix to the decoding device, the encoding efficiency can be improved.
[0263] Also, according to the encoding device 100 and the decoding device 200 according to the present embodiment, the second quantization matrix can be derived by applying the secondary transformation to the first quantization matrix. Therefore, the first quantization matrix corresponding to the primary space can be converted into the second quantization matrix corresponding to the secondary space, and the encoding efficiency can be improved while suppressing a decrease in subjective image quality.
[0264] (Modification of Embodiment 2) In the present embodiment, as an example of a method for deriving the second quantization matrix, an example in which the secondary transformation is directly applied to the first quantization matrix has been described, but the present invention is not limited to this. Other examples of the method for deriving the second quantization matrix will be described below with reference to FIG. 16.
[0265] FIG. 16 is a flowchart showing an example of the process for deriving the second quantization matrix in the modification of Embodiment 2. This flowchart shows an example of the processes of step S111 in FIG. 14 and step S212 in FIG. 15.
[0266] In FIG. 16, the quantization unit 108 or the inverse quantization unit 204 derives a third quantization matrix from the first quantization matrix (S301). At this time, for each component value of the third quantization matrix, the smaller the corresponding component value of the first quantization matrix, the larger it is. That is, as the component value of the first quantization matrix increases, the corresponding component value of the third quantization matrix decreases. In other words, the component value of the first quantization matrix and the component value of the third quantization matrix have a monotonically decreasing relationship. For example, each component value of the third quantization matrix is the reciprocal of the corresponding component value of the first quantization matrix.
[0267] Next, the quantization unit 108 or the inverse quantization unit 204 derives a fourth quantization matrix by applying a secondary transformation to the third quantization matrix (S302). That is, the quantization unit 108 or the inverse quantization unit 204 transforms the third quantization matrix using the basis used for the secondary transformation of the current block.
[0268] Finally, the quantization unit 108 or the inverse quantization unit 204 derives a fifth quantization matrix from the fourth quantization matrix as the second quantization matrix (S303). At this time, for each component value of the fifth quantization matrix, the smaller the corresponding component value of the fourth quantization matrix, the larger it is. That is, as the component value of the fourth quantization matrix increases, the corresponding component value of the fifth quantization matrix decreases. In other words, the component value of the fourth quantization matrix and the component value of the fifth quantization matrix have a monotonically decreasing relationship. For example, each component value of the fifth quantization matrix is the reciprocal of the corresponding component value of the fourth quantization matrix.
[0269] By deriving the second quantization matrix as described above, it is possible to reduce the influence of rounding errors during secondary transformation on components having relatively small values included in the first quantization matrix. That is, it is possible to reduce the influence of rounding errors on the values of the components applied to the coefficients for which it is desired to reduce the loss in order to suppress the degradation of subjective image quality. Therefore, it is possible to further suppress the degradation of subjective image quality.
[0270] Also, as each component value of the third quantization matrix / fifth quantization matrix, the reciprocal of the corresponding component value of the first quantization matrix / fourth quantization matrix can be used. Therefore, component values can be derived by simple calculation, and the processing load or processing time for deriving the second quantization matrix can be reduced.
[0271] (Embodiment 3) Next, Embodiment 3 will be described. In this embodiment, when secondary conversion is performed, the difference from Embodiment 1 above is that the secondary coefficients are quantized without using a quantization matrix. Hereinafter, this embodiment will be specifically described with reference to the drawings, centering on the differences from Embodiment 1 above.
[0272] [Conversion Processing, Quantization Processing, and Encoding Processing in the Encoding Device] The conversion processing, quantization processing, and encoding processing performed by the conversion unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding device 100 according to Embodiment 3 will be specifically described with reference to the drawings.
[0273] FIG. 17 is a flowchart showing an example of the conversion processing, quantization processing, and encoding processing in Embodiment 3. In FIG. 17, for processes substantially the same as those in FIG. 11, the same reference numerals are given, and the description will be omitted as appropriate.
[0274] After the secondary conversion from the primary coefficients to the secondary coefficients is performed (S104), the quantization unit 108 calculates quantized secondary coefficients by performing a second quantization different from the first quantization on the secondary coefficients (S121). In this embodiment, the second quantization is non-weighted quantization without using a quantization matrix. That is, in the second quantization, each secondary coefficient is divided by a quantization step common to the secondary coefficients of the block to be encoded. The common quantization step is derived from the quantization parameter for the block to be encoded. Specifically, the common quantization step is one constant fixed for all secondary coefficients of the block to be encoded. That is, the common quantization step is a constant independent of the position or order of the secondary coefficients.
[0275] As described above, in this embodiment, the quantization unit 108 switches between weighted quantization and unweighted quantization based on the application / non-application of the secondary conversion to the block to be encoded.
[0276] [Decoding process, inverse quantization process, and inverse conversion process in the decoding device] Next, the decoding process, inverse quantization process, and inverse conversion process performed by the entropy decoding unit 202, inverse quantization unit 204, and inverse conversion unit 206 of the decoding device 200 according to Embodiment 3 will be specifically described with reference to the drawings.
[0277] FIG. 18 is a flowchart showing an example of the decoding process, inverse quantization process, and inverse conversion process in Embodiment 3. In FIG. 18, for processes substantially the same as those in FIG. 13, the same reference numerals are given, and the description thereof will be omitted as appropriate.
[0278] When it is determined that inverse secondary conversion is to be performed (Yes in S202), the inverse quantization unit 204 calculates secondary coefficients by performing a second inverse quantization different from the first inverse quantization on the decoded quantization coefficients (S221). The second inverse quantization is the inverse quantization of the second quantization in the encoding device 100. In this embodiment, the second inverse quantization is unweighted inverse quantization without using a quantization matrix.
[0279] As described above, in this embodiment, the inverse quantization unit 204 switches between weighted inverse quantization and unweighted inverse quantization based on the application / non-application of inverse secondary conversion to the block to be decoded.
[0280] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to this embodiment, unweighted quantization / inverse quantization can be used as the second quantization / second inverse quantization. Therefore, it is possible to prevent a decrease in subjective image quality due to using the first quantization matrix for the first quantization / first inverse quantization in the second quantization / second inverse quantization, and to omit the encoding or derivation process of the quantization matrix for the second quantization.
[0281] (Embodiment 4) Next, Embodiment 4 will be described. In this embodiment, it is different from Embodiment 3 above in that each primary coefficient obtained by primary conversion is multiplied by the corresponding component value of the weight matrix and then the conversion is performed. Hereinafter, this embodiment will be specifically described with reference to the drawings centering on the differences from Embodiment 3 above.
[0282] [Conversion Processing, Quantization Processing, and Encoding Processing in the Encoding Apparatus] The conversion processing, quantization processing, and encoding processing performed by the conversion unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding apparatus 100 according to Embodiment 4 will be specifically described with reference to the drawings.
[0283] FIG. 19 is a flowchart showing an example of the conversion processing, quantization processing, and encoding processing in Embodiment 4. FIG. 20 is a diagram for explaining an example of the secondary conversion in Embodiment 4. In FIG. 19, for the processing substantially the same as that in FIG. 17, the same reference numerals are given and the description will be omitted as appropriate.
[0284] When it is determined that the secondary conversion is to be performed (Yes in S102), the conversion unit 106 performs the secondary conversion from the primary coefficients to the secondary coefficients (S130). Specifically, as shown in FIG. 20, in the secondary conversion, the conversion unit 106 calculates the weighted primary coefficients by multiplying each of the primary coefficients by the corresponding component value of the weight matrix (S131). The weight matrix is a matrix for weighting the primary coefficients. Then, the conversion unit 106 converts the weighted primary coefficients into secondary coefficients (S132). This conversion is substantially the same as the secondary conversion in each of the above embodiments, for example, except that the weighted primary coefficients are converted instead of the primary coefficients.
[0285] The weight matrix may be derived from the first quantization matrix. In this case, for example, the component values of the weight matrix and the component values of the first quantization matrix may have a monotonically decreasing relationship. Specifically, for example, each component value of the weight matrix may be the reciprocal of the corresponding component value of the first quantization matrix. Thereby, the first quantization matrix for weighting quantization steps can be converted into a weight matrix for weighting coefficients.
[0286] Also, the weight matrix may be included in the encoded bit stream or may be predefined in a standard. Also, in the standard, a plurality of weight matrices may be defined. In this case, the weight matrix may be selected from among the plurality of predefined weight matrices based on a given profile or level or the like.
[0287] The quantization unit 108 calculates quantized secondary coefficients by performing a second quantization different from the first quantization on the secondary coefficients (S121). In the present embodiment, the second quantization is non-weighted quantization that does not use a quantization matrix. That is, in the second quantization, each secondary coefficient is divided by a common quantization step for the secondary coefficients of the block to be encoded. At this time, the common quantization step may be derived, for example, from quantization parameters for the block to be encoded by the quantization unit 108.
[0288] [Decoding process, inverse quantization process, and inverse conversion process in the decoding device] Next, the decoding process, inverse quantization process, and inverse conversion process performed by the entropy decoding unit 202, inverse quantization unit 204, and inverse conversion unit 206 of the decoding device 200 according to Embodiment 4 will be specifically described with reference to the drawings.
[0289] FIG. 21 is a flowchart showing an example of the decoding process, inverse quantization process, and inverse conversion process in Embodiment 4. FIG. 22 is a diagram for explaining an example of inverse secondary conversion in Embodiment 4. In FIG. 21, for processes substantially the same as those in FIG. 18, the same reference numerals are given and the description will be omitted as appropriate.
[0290] When it is determined that inverse quadratic transformation is to be performed (Yes in S202), the inverse quantization unit 204 calculates secondary coefficients by performing a second inverse quantization different from the first inverse quantization on the decoded quantization coefficients (S221). The second inverse quantization is the inverse quantization of the second quantization in the encoding device 100. In the present embodiment, the second inverse quantization is a non-weighted inverse quantization that does not use a quantization matrix. That is, as shown in FIG. 22, the inverse quantization unit 204 calculates secondary coefficients by multiplying each of the quantization coefficients of the block to be decoded by a common quantization step. At this time, the common quantization step may be derived from, for example, the quantization parameter for the block to be decoded.
[0291] Next, the inverse transformation unit 206 performs an inverse quadratic transformation from the secondary coefficients to the primary coefficients (S230). Specifically, as shown in FIG. 22, the inverse transformation unit 206 performs an inverse transformation from the secondary coefficients to the weighted primary coefficients (S231). This inverse transformation is the inverse transformation of the transformation (S132) from the weighted primary coefficients to the secondary coefficients in the encoding device 100.
[0292] Furthermore, the inverse transformation unit 206 calculates the primary coefficients by dividing each of the weighted primary coefficients by the corresponding component value of the weight matrix (S232). This weight matrix matches the weight matrix used in the encoding device 100.
[0293] The weight matrix may be derived from the first quantization matrix by the inverse quantization unit 204. In this case, for example, the component values of the weight matrix and the component values of the first quantization matrix may have a monotonically decreasing relationship. For example, each component value of the weight matrix may be the reciprocal of the corresponding component value of the first quantization matrix.
[0294] Also, the weight matrix may be included in the encoded bit stream or may be predefined in a standard. Also, a plurality of weight matrices may be predefined in the standard. In this case, the weight matrix may be selected from among a plurality of predefined weight matrices based on a given profile or level or the like.
[0295] [Effects, etc.] As described above, according to the encoding apparatus 100 according to the present embodiment, weighted primary coefficients can be calculated by multiplying each of the primary coefficients by the corresponding component value of the weight matrix. Also, according to the decoding apparatus 200 according to the present embodiment, the primary coefficients can be calculated by dividing each of the weighted primary coefficients by the corresponding component value of the weight matrix. That is, according to the encoding apparatus 100 and the decoding apparatus 200 according to the present embodiment, weighting related to quantization can be performed on the primary coefficients before the secondary conversion. Therefore, when the secondary conversion / inverse secondary conversion is applied, quantization / inverse quantization equivalent to weighted quantization / inverse quantization can be performed without newly preparing a quantization matrix corresponding to the secondary space. As a result, it is possible to improve the encoding efficiency while suppressing a decrease in subjective image quality.
[0296] Also, according to the encoding apparatus 100 and the decoding apparatus 200 according to the present embodiment, the weight matrix can be derived from the first quantization matrix. Therefore, the amount of code for the weight matrix can be reduced, and it is possible to improve the encoding efficiency while suppressing a decrease in subjective image quality.
[0297] Also, according to the encoding apparatus 100 and the decoding apparatus 200 according to the present embodiment, a quantization step common to the secondary coefficients of the current block can be derived from the quantization parameter. Therefore, new information does not have to be included in the encoded stream for the common quantization step, and the amount of code for the common quantization step can be reduced.
[0298] (Embodiment 5) Next, Embodiment 5 will be described. In this embodiment, when secondary conversion is applied, quantization is performed on the primary coefficients before the secondary conversion, which is different from the above-described embodiments. Hereinafter, this embodiment will be specifically described with reference to the drawings, centering on the differences from the above-described embodiments.
[0299] [Conversion Processing, Quantization Processing, and Encoding Processing in the Encoding Device] The conversion processing, quantization processing, and encoding processing performed by the conversion unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding device 100 according to Embodiment 5 will be specifically described with reference to the drawings.
[0300] FIG. 23 is a flowchart showing an example of the conversion processing, quantization processing, and encoding processing in Embodiment 5. In FIG. 23, for the processing that is substantially the same as that in FIG. 11, the same reference numerals are given, and the description will be omitted as appropriate.
[0301] When it is determined that the secondary conversion is not performed (No in S102), the quantization unit 108 calculates the first quantized primary coefficients by performing the first quantization on the primary coefficients (S141).
[0302] On the other hand, when it is determined that the secondary conversion is to be performed (Yes in S102), the quantization unit 108 calculates the second quantized primary coefficients by performing the second quantization on the primary coefficients (S142). In this embodiment, the first quantization and the second quantization may be different from each other or the same as in the above-described Embodiments 1 to 4. That is, in this embodiment, in the first quantization and the second quantization, the same quantization matrix may be used to perform the same processing. Subsequently, the conversion unit 106 performs the secondary conversion from the second quantized primary coefficients to the quantized secondary coefficients (S143).
[0303] The entropy encoding unit 110 generates an encoded bit stream by encoding the first quantized primary coefficients or the quantized secondary coefficients (S106).
[0304] As described above, in this embodiment, when the second transformation is applied to the block to be encoded, the second quantization is performed before the second transformation. That is, the second quantization is performed on the primary coefficients.
[0305] [Decoding Process, Inverse Quantization Process, and Inverse Transformation Process in the Decoder] Next, the decoding process, inverse quantization process, and inverse transformation process performed by the entropy decoding unit 202, inverse quantization unit 204, and inverse transformation unit 206 of the decoder 200 according to Embodiment 5 will be specifically described with reference to the drawings.
[0306] FIG. 24 is a flowchart showing an example of the decoding process, inverse quantization process, and inverse transformation process in Embodiment 5. In FIG. 24, for processes substantially the same as those in FIG. 13, the same reference numerals are assigned and the description is omitted as appropriate.
[0307] When it is determined that the inverse second transformation is not to be performed (No in S202), the inverse quantization unit 204 calculates the primary coefficients by performing the first inverse quantization on the decoded quantization coefficients (S241). The first inverse quantization is the inverse quantization of the first quantization in the encoder 100.
[0308] On the other hand, when it is determined that the inverse second transformation is to be performed (Yes in S202), the inverse transformation unit 206 performs the inverse second transformation from the decoded quantization coefficients to the quantized primary coefficients (S242). The inverse second transformation is the inverse transformation of the second transformation in the encoder 100. Subsequently, the inverse quantization unit 204 calculates the primary coefficients by performing the second inverse quantization on the quantized primary coefficients (S243). The second inverse quantization is the inverse quantization of the second quantization in the encoder 100. Therefore, when the second quantization is the same as the first quantization, the second inverse quantization is the same as the first inverse quantization.
[0309] As described above, in this embodiment, when the inverse second transformation is applied to the block to be decoded, the second inverse quantization is performed after the inverse second transformation. That is, the second inverse quantization is performed on the quantized primary coefficients.
[0310] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to the present embodiment, quantization can be performed before the second conversion. Therefore, when the second conversion process is lossless, the second conversion can be removed from the loop of the prediction process. Thus, the load on the processing pipeline can be reduced. Further, by performing quantization before the second conversion, it is not necessary to separate the first quantization matrix and the second quantization matrix, so that the process can be simplified.
[0311] (Modification example) As described above, the encoding device and the decoding device according to one or more aspects of the present disclosure have been described based on the embodiments. However, the present disclosure is not limited to these embodiments. Without departing from the spirit of the present disclosure, various modifications conceived by those skilled in the art applied to the present embodiment or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects of the present disclosure.
[0312] For example, in the above-described Embodiment 2, the derivation of the second quantization matrix is performed after the second conversion or after the determination of the inverse second conversion, but it is not limited thereto. The derivation of the second quantization matrix may be performed at any time after the first quantization matrix is obtained and before the second quantization / inverse quantization. For example, the second quantization matrix may be derived at the start of encoding or decoding of the current picture including the current block. In this case, the second quantization matrix may not be derived for each block.
[0313] In each of the above embodiments, the encoding / decoding of one encoding / decoding target block has been mainly described. However, the above-described conversion process, quantization process, and encoding process, or decoding process, inverse quantization process, and inverse conversion process can be applied to a plurality of blocks included in the encoding / decoding target picture. In this case, the first quantization matrix and the second quantization matrix corresponding to the prediction mode (e.g., intra prediction or inter prediction), the type of pixel value (e.g., luminance or color difference), the size of the block, or any combination thereof may be used.
[0314] Note that the quantization switching process based on the application / non-application of the secondary conversion in each of the above embodiments may be turned on / off at the slice level, tile level, CTU level, or CU level. Also, the on / off may be determined according to the frame type (I-frame, P-frame, B-frame) and / or prediction mode.
[0315] Also, the quantization switching process based on the application / non-application of the secondary conversion in each of the above embodiments may be performed on one or both of the luminance block and the chrominance block.
[0316] Note that in each of the above embodiments, the determination of whether to perform the secondary conversion is performed after the primary conversion, but it is not limited to this. The determination of whether to perform the secondary conversion may be performed in advance before the processing of the coding target block.
[0317] Also, in Embodiment 1 above, the first quantization matrix and the second quantization matrix do not always have to be different.
[0318] (Embodiment 6) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, etc. Also, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (dedicated circuit).
[0319] Also, the processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing or distributed processing may be performed.
[0320] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible, and these are also included within the scope of the aspects of the present disclosure.
[0321] Furthermore, here, an application example of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, and an image encoding / decoding device having both. Other configurations in the system can be appropriately changed as the case may be.
[0322] [Usage Example] FIG. 25 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed radio stations, are installed in each cell.
[0323] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above elements. The devices may be directly or indirectly connected to each other via a telephone network or short-range wireless, etc., without passing through the base stations ex106 to ex110, which are fixed radio stations. Further, the streaming server ex103 is connected to devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smartphone ex115 via the Internet ex101 or the like. Also, the streaming server ex103 is connected to terminals etc. within a hotspot in an airplane ex117 via a satellite ex116.
[0324] Note that, instead of the base stations ex106 to ex110, a wireless access point, a hotspot, or the like may be used. Further, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.
[0325] The camera ex113 is a device capable of taking still images and moving images such as a digital camera. Further, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handyphone System) or the like corresponding to the mobile communication system standards generally called 2G, 3G, 3.9G, 4G, and in the future 5G.
[0326] The home appliance ex118 is a device included in a refrigerator or a household fuel cell cogeneration system or the like.
[0327] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 through the base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal in the airplane ex117, etc.) performs the encoding process described in each of the above embodiments on the still image or moving image content photographed by the user using the terminal, multiplexes the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present disclosure.
[0328] On the one hand, the streaming server ex103 streams the transmitted content data to the requested client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc., which is capable of decrypting the encoded data. Each device that receives the distributed data decrypts and plays back the received data. That is, each device functions as an image decoding device according to one aspect of the present disclosure.
[0329] [Distributed processing] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world. In a CDN, an edge server physically close to the client is dynamically assigned according to the client. Then, by caching and distributing the content to the edge server, the delay can be reduced. Also, when some error occurs or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the distribution, so high-speed and stable distribution can be realized.
[0330] Furthermore, not only the distribution process itself can be decentralized, but the encoding process of the captured data can be performed on each terminal, on the server side, or shared between them. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image or the amount of code is detected in units of frames or scenes. In the second loop, processing is performed to improve the encoding efficiency while maintaining the image quality. For example, by having the terminal perform the first encoding process and the server that receives the content perform the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode in almost real time, the already encoded data from the first encoding performed by the terminal can also be received and played back by other terminals, enabling more flexible real-time distribution.
[0331] As another example, cameras such as ex113 perform feature extraction from the image, compress the data related to the features as metadata, and send it to the server. The server performs compression according to the meaning of the image, such as determining the importance of the object from the features and switching the quantization accuracy. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) can be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed on the server.
[0332] As yet another example, in a stadium, shopping mall, or factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using the plurality of terminals that performed the shooting, and other terminals and servers that did not perform the shooting as necessary, encoding processes are respectively assigned and decentralized processing is performed, for example, in units of GOP (Group of Picture), picture, or tiles obtained by dividing the picture. This can reduce the delay and achieve more real-time performance.
[0333] Also, since the plurality of video data is of substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can be referred to each other. Alternatively, the server may receive the encoded data from each terminal and change the reference relationship among the plurality of data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.
[0334] Also, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based one, or convert H.264 to H.265.
[0335] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" are used as the subject performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.
[0336] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and use different scenes captured by terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal are integrated based on the relative positional relationship between the terminals obtained separately, or the regions where the feature points included in the videos match.
[0337] The server may not only encode two-dimensional moving images, but also automatically encode still images based on scene analysis of the moving images, etc., or at a time specified by the user, and transmit them to the receiving terminal. When the server can further obtain the relative positional relationship between the shooting terminals, it can generate the three-dimensional shape of the scene based on not only two-dimensional moving images, but also videos shot from different angles of the same scene. Note that the server may separately encode three-dimensional data generated by point clouds, etc., or select or reconstruct the video transmitted to the receiving terminal from the videos shot by multiple terminals based on the results of recognizing or tracking a person or an object using the three-dimensional data.
[0338] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or enjoy the content obtained by cutting out the video from an arbitrary viewpoint from the three-dimensional data reconstructed using multiple images or videos. Further, similar to the video, sound is also collected from a plurality of different angles, and the server may multiplex and transmit the sound from a specific angle or space together with the video according to the video.
[0339] In recent years, content associating the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by Multi-View Coding (MVC), etc., or encode them as separate streams without referring to each other. At the time of decoding the separate streams, they should be synchronized and played back so that a virtual three-dimensional space is reproduced according to the user's viewpoint.
[0340] In the case of an AR image, the server superimposes virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold the virtual object information and the three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting the virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like, and encode it in a state where the portion is transparent. Alternatively, the server may set the RGB value of a predetermined value like a chroma key as the background, and generate data with the portion other than the object being the background color.
[0341] Similarly, the decoding process of the distributed data may be performed on each client terminal, on the server side, or shared between them. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request on another terminal, perform the decoding process, and the decoded signal may be transmitted to the device having a display. By dispersing the process and selecting appropriate content regardless of the performance of the communicable terminal itself, it is possible to reproduce data with good image quality. Also, as another example, while receiving large-size image data on a TV or the like, a part of the area such as a tile in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to check at hand the area of one's own field of responsibility or the area that one wants to check in more detail.
[0342] In the future, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used both indoors and outdoors, it is expected that content will be received seamlessly while switching appropriate data for the ongoing communication by utilizing a delivery system standard such as MPEG-DASH. As a result, the user can freely select not only their own terminal but also a decoding device or display device such as a display installed indoors and outdoors and switch in real time. Also, based on their own location information, etc., decoding can be performed while switching the terminal to be decoded and the terminal to be displayed. This enables, for example, while moving to a destination, to move while displaying map information on a part of the wall surface or ground of the neighboring building where a displayable device is embedded. Also, based on the ease of access to encoded data on the network, such as the encoded data being cached in a server that can be accessed from a receiving terminal in a short time, or being copied to an edge server in a content delivery service, etc., it is also possible to switch the bitrate of the received data.
[0343] [Scalable Encoding] Regarding content switching, it will be described using a scalable stream that is compression-encoded by applying the moving image encoding method shown in each of the above embodiments, as shown in FIG. 26. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of a temporally / spatially scalable stream realized by encoding in layers as shown in the figure, a configuration for switching content may be used. That is, by the decoding side determining up to which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on a smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0344] Furthermore, as described above, pictures are encoded layer by layer. In addition to the configuration that realizes scalability where an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of an image or the like, and the decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta information includes information for specifying linear or non-linear filter coefficients used in the super-resolution process, or information for specifying parameter values in filter processing, machine learning, or least-squares operations used in the super-resolution process, etc.
[0345] Alternatively, the picture may be divided into tiles or the like according to the meaning of an object or the like in the image, and the decoding side may decode only a part of the area by selecting the tile to be decoded. Also, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 27, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.
[0346] Also, the meta information may be stored in units composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining with the information per picture, can specify the picture in which the object exists and the position of the object in the picture.
[0347] [Optimization of Web Page] FIG. 28 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 29 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 28 and 29, a web page may include a plurality of link images that are links to image contents, and the appearance thereof may vary depending on the device for viewing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) displays a still image or an I picture that each content has as a link image, displays a video like a gif animation with a plurality of still images or I pictures, etc., or receives only the base layer and decodes and displays the video.
[0348] When a link image is selected by the user, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Also, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay (delay from the start of decoding of the content to the start of display) between the decoding time and the display time of the leading picture by decoding and displaying only the forward reference pictures (I pictures, P pictures, B pictures with only forward reference). Further, the display device may deliberately ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures with forward reference, and perform normal decoding as the received pictures increase over time.
[0349] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode these in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.
[0350] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal move, the receiving terminal can achieve seamless reception and decoding by transmitting the position information of the receiving terminal at the time of a reception request while switching between the base stations ex106 to ex110. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication band.
[0351] As described above, in the content supply system ex100, the client can receive, decode, and play back the encoded information transmitted by the user in real time.
[0352] [Delivery of Personal Content] Also, in the content supply system ex100, not only high-quality and long-duration content by video delivery providers but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. Also, it is considered that such personal content will increase in the future. In order to make personal content into better content, the server may perform an encoding process after performing an editing process. This can be realized, for example, with the following configuration.
[0353] During shooting in real time or by accumulation and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection on the original image or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone for editing. The server encodes the edited data based on the editing results. It is also known that if the shooting time is too long, the viewing rate will decrease. The server may automatically clip scenes with little movement as well as less important scenes as described above so that the content within a specific time range is obtained according to the shooting time, based on the image processing results. Or, the server may generate a digest based on the result of the semantic analysis of the scene, encode it, and output it.
[0354] Note that personal content may, as it is, contain elements that would infringe copyright, moral rights of the author, or the right of portrait, etc., and there may be inconvenient cases for individuals, such as the sharing scope exceeding the intended scope. Therefore, for example, the server may deliberately change the image of a person's face in the peripheral part of the screen or inside a house to an out-of-focus image and then encode it. Also, the server may recognize whether a face of a person different from a pre-registered person appears in the image to be encoded, and if it appears, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, the user designates a person or background area that the user wants to process the image from the perspective of copyright, etc., and the server can perform processing such as replacing the designated area with another video or blurring the focus. For a person, the video of the face part can be replaced while tracking the person in the moving image.
[0355] In addition, since the viewing of personal content with a small data volume has a strong requirement for real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and performs decoding and playback. During this time, the decoding device receives the enhancement layer. When playback is looped or played back two or more times, such as when the enhancement layer is also included, high-quality video may be played back. For a stream with scalable encoding like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played back for the first time and a second stream encoded with reference to the first video are configured as one stream.
[0356] [Other usage examples] In addition, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processing may be performed using the software. Further, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.
[0357] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays back the content.
[0358] Moreover, not limited to the content supply system ex100 via the Internet ex101, at least either the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can be incorporated into a digital broadcast system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by loading it on broadcast radio waves using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the configuration of the content supply system ex100 that is easy to perform unicast, but the same application is possible for encoding processing and decoding processing.
[0359] [Hardware Configuration] FIG. 30 is a diagram showing a smartphone ex115. FIG. 31 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by the camera unit ex465 and video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with the SIM ex468 for identifying the user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.
[0360] In addition, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a bus ex470.
[0361] When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack to activate the smartphone ex115 to an operable state.
[0362] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified, frequency conversion processing and analog-to-digital conversion processing are performed, spectral inverse spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. During the data communication mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body, etc., and similar transmission and reception processing is performed. When transmitting video, still images, or video and audio during the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while a video or still image, etc. is being captured by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.
[0363] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and a video or a still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming has become widespread, there may be a situation where it is not socially appropriate to play audio depending on the user's situation. Therefore, as an initial value, it is desirable to have a configuration that plays only the video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.
[0364] Also, although the smartphone ex115 has been described as an example here, as the terminal, in addition to the transceiver type terminal having both an encoder and a decoder, there are three implementation forms: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. Furthermore, in the digital broadcast system, although it has been described as receiving or transmitting multiplexed data in which audio data and the like are multiplexed in video data, in the multiplexed data, character data related to the video and the like may be multiplexed in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.
[0365] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, many terminals also have a GPU. Therefore, a configuration in which a wide area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so as to be commonly used may be adopted. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be realized. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures by the GPU instead of the CPU.
Industrial Applicability
[0366] The present disclosure can be applied to, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, or a digital video camera.
Explanation of Signs
[0367] 100 Encoding device 102 Splitting unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 200 Decoding device 202 Entropy decoding unit
Claims
1. 1. An encoding device for encoding an image block, comprising: The circuit, A memory, The circuit uses the memory to: performing a linear transformation of the residuals of the image block into linear coefficients in a linear space; determining whether to apply a secondary transform to the image block; (i) when the secondary transformation is not applied, calculating quantized primary coefficients by performing a first quantization on the primary coefficients; and (ii) when the secondary transformation is applied, calculating quantized secondary coefficients by performing a secondary transformation from the primary coefficients to secondary coefficients in a secondary space different from the primary space and performing a second quantization on the secondary coefficients different from the first quantization. Encoding device.
2. 1. A decoding device for decoding an image block, comprising: The circuit, A memory, The circuit uses the memory to: obtaining a quantization coefficient of the image block; determining whether to apply an inverse quadratic transform to the image block; (i) when the inverse secondary transform is not applied, calculating primary coefficients in a primary space by performing a first inverse quantization on the quantized coefficients; (ii) when the inverse secondary transform is applied, calculating secondary coefficients in a secondary space different from the primary space by performing a second inverse quantization different from the first inverse quantization on the quantized coefficients, performing an inverse secondary transform from the secondary coefficients to primary coefficients, and performing an inverse primary transform from the primary coefficients to residuals of the image block. Decryption device.
3. The circuit, a memory connected to the circuit; The circuit, in operation, generating information for causing a decoding device to execute an inverse secondary transform in the inverse transform process; including said information in a bitstream; The inverse transformation process is A quantization coefficient for the image block is obtained; if the information indicates that an inverse secondary transform is not applied to the image block, a first inverse quantization is performed on the quantized coefficients to calculate primary coefficients in a primary space; if the information indicates that the inverse secondary transform is applied to the image block, a second inverse quantization different from the first inverse quantization is performed on the quantized coefficients to calculate secondary coefficients in a secondary space different from the primary space, an inverse secondary transform is performed on the secondary coefficients to calculate primary coefficients, and an inverse primary transform is performed on the primary coefficients to calculate residuals of the image block. Bitstream generator.
Citation Information
Patent Citations
Image information converter and image information conversion method
JP2001148852A
Video encoding and decoding using transformations
JP2014523175A
Bit number prediction for VLC coded DCT coefficients and its application in DV encoding / transcoding
US20020150157A1
Signal processing apparatus and method including quantization or inverse-quantization process
US20160212428A1
Non-separable secondary transform for video coding
US20170094313A1