Image processing device and method, program, and recording medium
By performing matrix operations on transformed coefficients and controlling secondary transforms based on non-zero coefficients, the method addresses the efficiency loss in image coding by preventing coefficient diffusion and maintaining energy compaction.
Patent Information
- Application Number
- JP2024153508
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-06-08
- Filing Date
- 2024-09-05
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2037-04-28
AI Technical Summary
Existing image coding methods face a decrease in coding efficiency due to the application of secondary transforms on sub-blocks with sparse non-zero coefficients, leading to diffusion of coefficients from low-order to high-order components, and the risk of reduced energy compaction.
Implement a matrix operation on a one-dimensional vector obtained by transforming transform coefficients, using a matrix for inverse transform processing, to obtain a prediction residual, and control the application of secondary transforms based on the number of non-zero coefficients in each sub-block.
This approach prevents the diffusion of coefficients and maintains energy compaction, thereby suppressing a decrease in coding efficiency.
Smart Images

Figure 0007758116000013 
Figure 0007758116000014 
Figure 0007758116000015
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing device and method , program, and recording medium and an image processing device and method that can suppress a decrease in encoding efficiency. , program, and recording medium Regarding. [Background technology]
[0002] Conventionally, in image coding, after performing a primary transform on a prediction residual, which is the difference between an image and its predicted image, a secondary transform is applied to each sub-block in a transform block in order to increase energy compaction (concentrate transform coefficients in a low frequency range) (see, for example, Non-Patent Document 1). Non-Patent Document 1 also discloses that a secondary transform identifier, which indicates which secondary transform to apply, is signaled on a CU-by-CU basis.
[0003] Furthermore, in an encoder, determining which secondary transform to apply on a CU-by-CU basis as described in Non-Patent Document 1 based on RDO (Rate-Distortion Optimization) involves high computational complexity, and it is disclosed that a secondary transform flag indicating whether or not to apply a secondary transform on a transform block-by-transform block basis is signaled (see, for example, Non-Patent Document 2). Non-Patent Document 2 also discloses that a secondary transform identifier indicating which secondary transform to apply is derived based on a primary transform identifier and an intra prediction mode.
[0004] However, in both the methods described in Non-Patent Document 1 and Non-Patent Document 2, when a sub-block with sparse non-zero coefficients is input to the secondary transform coefficients, a secondary transform is applied, which causes the coefficients to diffuse from low-order to high-order components within the sub-block, resulting in a decrease in energy compaction and a risk of reduced coding efficiency.
[0005] Furthermore, Joint Exploration Test Model 1 (JEM1) discloses that the maximum size of a CTU (Coding Tree Unit) is extended to 256x256 and the maximum size of a transform block is accordingly extended to 64x64 in order to improve the coding efficiency of high-resolution images such as 4K (see, for example, Non-Patent Document 1). Non-Patent Document 1 also discloses that when the transform block size is 64x64, the encoder performs band limitation (discards the high-frequency components) so that the transform coefficients of high-frequency components other than the low-frequency components in the upper left 32x32 of the transform block are forced to be 0, and encodes only the non-zero coefficients of the low-frequency components.
[0006] In this case, the decoder only needs to decode the non-zero coefficients of the low-frequency components and perform inverse quantization and inverse transform on those non-zero coefficients, which reduces the computational complexity and implementation cost of the decoder compared to when the encoder does not perform band limiting.
[0007] However, when an encoder performs a transform skip or a transform and quantization skip (hereinafter referred to as a transform / quantization bypass) on a 64x64 transform block, the transform coefficients of the 64x64 transform block are prediction residuals before transform. Therefore, if bandwidth limitation is performed in this case, distortion increases. As a result, there is a risk of reducing coding efficiency. Furthermore, even though a transform / quantization bypass is performed for the purpose of lossless coding, lossless coding cannot be performed. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] Jianle Chen, Elena Alshina, Gary J. Sullivan, Jens-Rainer Ohm, Jill Boyce, "Algorithm Description of Joint Exploration Test Model 2", JVET-B1001_v3, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 2nd Meeting: San Diego, USA, 20-26 February 2016 [Non-patent document 2] X.Zhao, A.Said, V.Seregin, M.Karczewicz, J.Chen, R.Joshi, "TU-level non-separable secondary transform", JVET-B0059, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 2nd Meeting: San Diego, USA, 20-26 February 2016 Summary of the Invention [Problem to be solved by the invention]
[0009] As described above, there is a risk that the coding efficiency will decrease.
[0010] The present disclosure has been made in light of such circumstances, and makes it possible to suppress a decrease in coding efficiency. [Means for solving the problem]
[0011] The image processing device according to one aspect of the present technology is Scanning Methodand a matrix calculation unit that performs a matrix operation on a one-dimensional vector obtained by transforming a transform coefficient, using a matrix for inverse transform processing on the transform coefficients set based on the above, to obtain a prediction residual, which is the difference between an image and a predicted image of the image, and a matrix generation unit that generates a matrix of the one-dimensional vector scaled to the one-dimensional vector on which the matrix operation has been performed.
[0012] An image processing method according to one aspect of the present technology includes: Scanning Method and performing a matrix operation on a one-dimensional vector obtained by transforming transform coefficients, which is a one-dimensional vector obtained by performing an inverse transform process using a matrix for inverse transform processing on transform coefficients set based on the above formula, to obtain a prediction residual, which is the difference between an image and a predicted image of the image; and matrixizing the one-dimensional vector scaled for the one-dimensional vector obtained by the matrix operation. A program according to one aspect of the present technology includes a secondary conversion identifier and Scanning Method and performing a matrix operation on a one-dimensional vector obtained by transforming a transform coefficient, which is a transformation coefficient used to obtain a prediction residual, which is the difference between an image and a predicted image of the image, by performing an inverse transform process using a matrix for an inverse transform process on the transform coefficients set based on the above; and matrixing the one-dimensional vector scaled for the one-dimensional vector on which the matrix operation has been performed. A recording medium according to one aspect of the present technology includes a secondary conversion identifier and Scanning Method and performing a matrix operation on a one-dimensional vector obtained by transforming a transform coefficient, which is obtained by performing an inverse transform process using a matrix for inverse transform processing on the transform coefficients set based on the above, to obtain a prediction residual, which is the difference between an image and a predicted image of the image; and matrixing the one-dimensional vector scaled for the one-dimensional vector on which the matrix operation has been performed.
[0013] In an image processing device, method, program, and recording medium according to one aspect of the present technology, a secondary conversion identifier and Scanning Method A matrix for inverse transform processing on the transform coefficients, which is set based on the above, is used, and a matrix operation is performed on a one-dimensional vector into which the transform coefficients have been transformed, and a prediction residual, which is the difference between an image and a predicted image of that image, is obtained by performing the inverse transform processing.A one-dimensional vector scaled to the one-dimensional vector on which the matrix operation has been performed is then put into a matrix. [Effects of the Invention]
[0014] According to the present disclosure, it is possible to process images, and in particular, to suppress a decrease in coding efficiency. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 10 is an explanatory diagram for explaining an outline of recursive block division for a CU. [Figure 2] FIG. 2 is an explanatory diagram for explaining setting of a PU to a CU shown in FIG. [Figure 3] FIG. 2 is an explanatory diagram for explaining setting of a TU to a CU shown in FIG. [Figure 4] FIG. 10 is an explanatory diagram for explaining the scanning order of CU / PU. [Figure 5] FIG. 2 is a block diagram illustrating an example of the main configuration of a secondary conversion unit. [Figure 6] FIG. 10 is a diagram illustrating an example of a secondary conversion. [Figure 7] FIG. 1 is a block diagram illustrating an example of the main configuration of an image encoding device. [Figure 8] FIG. 2 is a block diagram illustrating an example of the main configuration of a conversion unit. [Figure 9] FIG. 10 is a diagram illustrating examples of scan methods corresponding to scan identifiers. [Figure 10] FIG. 10 is a diagram illustrating an example of a matrix of a secondary transformation. [Figure 11] 10 is a flowchart illustrating an example of the flow of an image encoding process. [Figure 12] 10 is a flowchart illustrating an example of the flow of a conversion process. [Figure 13] FIG. 1 is a block diagram illustrating an example of the main configuration of an image decoding device. [Figure 14] FIG. 2 is a block diagram illustrating an example of the main configuration of an inverse conversion unit. [Figure 15] 10 is a flowchart illustrating an example of the flow of an image decoding process. [Figure 16] 10 is a flowchart illustrating an example of the flow of an inverse transformation process. [Figure 17] FIG. 10 is a diagram illustrating an example of the relationship between intra prediction modes and scanning methods. [Figure 18] FIG. 2 is a block diagram illustrating an example of the main configuration of a conversion unit. [Figure 19] FIG. 10 is a block diagram illustrating an example of the main configuration of a secondary conversion selection unit. [Figure 20] 10 is a flowchart illustrating an example of the flow of a conversion process. [Figure 21] FIG. 2 is a block diagram illustrating an example of the main configuration of an inverse conversion unit. [Figure 22] FIG. 10 is a block diagram illustrating an example of the main configuration of an inverse secondary transformation selection unit. [Figure 23] 10 is a flowchart illustrating an example of the flow of an inverse transformation process. [Figure 24] FIG. 2 is a block diagram illustrating an example of the main configuration of a conversion unit. [Figure 25] FIG. 10 is a diagram illustrating a bandwidth limit. [Figure 26] FIG. 2 is a block diagram illustrating an example of the main configuration of a conversion unit. [Figure 27] 10 is a flowchart illustrating an example of the flow of a conversion process. [Figure 28] FIG. 2 is a block diagram illustrating an example of the main configuration of an inverse conversion unit. [Figure 29] 10 is a flowchart illustrating an example of the flow of an inverse transformation process. [Figure 30] FIG. 2 is a diagram illustrating the shapes of CU, PU, and TU. [Figure 31] FIG. 2 is a block diagram illustrating an example of the main configuration of a conversion unit. [Figure 32] FIG. 10 illustrates an example of a band-limiting filter. [Figure 33] 10 is a flowchart illustrating an example of the flow of a conversion process. [Figure 34] FIG. 2 is a block diagram illustrating an example of the main configuration of an inverse conversion unit. [Figure 35] 10 is a flowchart illustrating an example of the flow of an inverse transformation process. [Figure 36] FIG. 10 is a diagram illustrating another example of a band-limiting filter. [Figure 37] FIG. 10 is a diagram illustrating another example of a band-limiting filter. [Figure 38] FIG. 10 is a diagram illustrating another example of a band-limiting filter. [Figure 39] 10 is a flowchart illustrating another example of the flow of the conversion process. [Figure 40] 10 is a flowchart illustrating another example of the flow of the inverse transformation process. [Figure 41] FIG. 1 is a block diagram illustrating an example of the main configuration of a computer. [Figure 42] FIG. 1 is a block diagram illustrating an example of a schematic configuration of a television device. [Figure 43] FIG. 1 is a block diagram showing an example of a schematic configuration of a mobile phone. [Figure 44] FIG. 1 is a block diagram showing an example of a schematic configuration of a recording / reproducing device. [Figure 45] FIG. 1 is a block diagram illustrating an example of a schematic configuration of an imaging device. [Figure 46] FIG. 1 is a block diagram showing an example of a schematic configuration of a video set. [Figure 47] FIG. 2 is a block diagram showing an example of a schematic configuration of a video processor. [Figure 48] FIG. 10 is a block diagram showing another example of the schematic configuration of the video processor. [Figure 49] FIG. 1 is a block diagram illustrating an example of a schematic configuration of a network system. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described in the following order. 1. First embodiment (skipping secondary transform for each sub-block) 2. Second embodiment (selection of secondary transformation using scanning method) 3. Third embodiment (Skip of bandwidth limitation when blocks are square) 4. Fourth embodiment (Skipping bandwidth limitation when blocks are square or rectangular) 5. Fifth embodiment (other)
[0017] <1. First embodiment> <Block division> In conventional image coding methods such as MPEG2 (Moving Picture Experts Group 2 (ISO / IEC 13818-2)) and MPEG-4 Part 10 (Advanced Video Coding, hereinafter referred to as AVC), coding is performed in processing units called macroblocks. A macroblock is a block having a uniform size of 16x16 pixels. In contrast, in HEVC (High Efficiency Video Coding), coding is performed in processing units (coding units) called CUs (Coding Units). A CU is a block of variable size formed by recursively dividing an LCU (Largest Coding Unit), which is the largest coding unit. The maximum size of a selectable CU is 64x64 pixels. The minimum size of a selectable CU is 8x8 pixels. A CU with the smallest size is called an SCU (Smallest Coding Unit). Note that the maximum size of a CU is not limited to 64x64 pixels, and larger block sizes such as 128x128 pixels or 256x256 pixels may be used.
[0018] As a result of employing CUs with variable sizes in this way, HEVC makes it possible to adaptively adjust image quality and coding efficiency according to the content of an image. Prediction processing for predictive coding is performed in processing units (prediction units) called PUs (Prediction Units). PUs are formed by dividing a CU using one of several division patterns. A PU is composed of processing units (prediction blocks) called PBs (Prediction Blocks) for each of luma (Y) and chroma (Cb, Cr). Orthogonal transform processing is performed in processing units (transform units) called TUs (Transform Units). TUs are formed by dividing a CU or PU to a certain depth. A TU is composed of processing units (transform blocks) called TBs (Transform Blocks) for each of luma (Y) and chroma (Cb, Cr).
[0019] <Recursive block division> Fig. 1 is an explanatory diagram for explaining an overview of recursive block division of a CU in HEVC. Block division of a CU is performed by recursively repeating the division of one block into four (=2x2) sub-blocks, resulting in the formation of a quad-tree-like tree structure. The entire quad-tree is called a CTB (Coding Tree Block), and a logical unit corresponding to a CTB is called a CTU.
[0020] At the top of FIG. 1, as an example, C01, which is a CU having a size of 64x64 pixels, is shown. The depth of the division of C01 is equal to zero. This means that C01 is the root of the CTU and corresponds to the LCU. The LCU size can be specified by parameters encoded in the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set). C02, which is a CU, is one of the four CUs divided from C01 and has a size of 32x32 pixels. The depth of the division of C02 is equal to 1. C03, which is a CU, is one of the four CUs divided from C02 and has a size of 16x16 pixels. The depth of the division of C03 is equal to 2. C04, which is a CU, is one of the four CUs divided from C03 and has a size of 8x8 pixels. The depth of the division of C04 is equal to 3. Thus, the CU is formed by recursively dividing the encoded image. The depth of the division is variable. For example, in a flat image region such as a blue sky, a CU of a larger size (i.e., a smaller depth) can be set. On the other hand, in a steep image region containing many edges, a CU of a smaller size (i.e., a larger depth) can be set. And each of the set CUs becomes a processing unit for the encoding process.
[0021] <Setting of PU to CU> The PU is a processing unit for prediction processing including intra prediction and inter prediction. The PU is formed by dividing the CU by one of several division patterns. FIG. 2 is an explanatory diagram for explaining the setting of the PU to the CU shown in FIG. 1. On the right side of FIG. 2, eight types of division patterns, namely 2Nx2N, 2NxN, Nx2N, NxN, 2NxnU, 2NxnD, nLx2N, and nRx2N, are shown. Among these division patterns, for intra prediction, two types, 2Nx2N and NxN, are selectable (NxN is only selectable in the SCU). In contrast, for inter prediction, when asymmetric motion division is enabled, all eight types of division patterns are selectable.
[0022] <Setting of TU to CU> A TU is a processing unit for orthogonal transformation processing. A TU is formed by dividing a CU (for an intra-CU, each PU within the CU) to a certain depth. FIG. 3 is an explanatory diagram for explaining the setting of TUs to the CU shown in FIG. 2. On the right side of FIG. 3, one or more TUs that can be set to C02 are shown. For example, T01, which is a TU, has a size of 32x32 pixels, and the depth of its TU division is equal to zero. T02, which is a TU, has a size of 16x16 pixels, and the depth of its TU division is equal to 1. T03, which is a TU, has a size of 8x8 pixels, and the depth of its TU division is equal to 2.
[0023] How to perform block division to set the above-mentioned blocks such as CU, PU, and TU in an image is typically determined based on a comparison of costs that affect encoding efficiency. The encoder compares the costs, for example, between one 2Mx2M pixel CU and four MxM pixel CUs, and if the encoding efficiency is higher when setting four MxM pixel CUs, it determines to divide the 2Mx2M pixel CU into four MxM pixel CUs.
[0024] <Scanning Order of CU and PU> When encoding an image, CTBs (or LCUs) set in a grid pattern within the image (or slice, tile) are scanned in raster scan order. Within one CTB, CUs are scanned by tracing the quadtree from left to right and top to bottom. When processing a current block, information on neighboring blocks above and to the left is used as input information. FIG. 4 is an explanatory diagram for explaining the scanning order of CUs and PUs. The upper left of FIG. 4 shows four CUs, C10, C11, C12, and C13, that may be included in one CTB. The numbers in the boxes of each CU indicate the processing order. The encoding process is performed in the order of the upper left CU C10, the upper right CU C11, the lower left CU C12, and the lower right CU C13. The right side of FIG. 4 shows one or more PUs for inter prediction that may be set to CU C11. The bottom of FIG. 4 shows one or more PUs for intra prediction that may be set to CU C12. As indicated by the numbers in the boxes of these PUs, the PUs are also scanned from left to right and from top to bottom.
[0025] In the following description, a "block" may be used to refer to a partial region or processing unit of an image (picture) (not a block of a processing unit). In this case, a "block" refers to any partial region within a picture, and its size, shape, characteristics, etc. are not limited. In other words, a "block" in this case includes any partial region (processing unit), such as a TB, TU, PB, PU, SCU, CU, LCU (CTB), sub-block, macroblock, tile, or slice.
[0026] <Secondary Transformation> Non-patent documents 1 and 2 describe that after a primary transform is performed on a prediction residual, which is the difference between an image and its predicted image, a secondary transform is applied to each sub-block within the transform block in order to increase energy compaction (concentrate transform coefficients in the low range).
[0027] However, when a prediction residual with sparse non-zero coefficients is subjected to secondary transformation, the coefficients are diffused from low-order to high-order components, which may reduce energy compaction and decrease coding efficiency.
[0028] Fig. 5 is a block diagram showing an example of the main configuration of a secondary transform unit that performs secondary transform. The secondary transform unit 11 shown in Fig. 5 is a processing unit that performs secondary transform on primary transform coefficients obtained by primary transforming prediction residuals using the methods described in Non-Patent Document 1 and Non-Patent Document 2. As shown in Fig. 5, the secondary transform unit 11 has a rasterization unit 21, a matrix operation unit 22, a scaling unit 23, and a matrix generation unit 24.
[0029] The rasterization unit 21 scans the input primary transform coefficients Coeff_P in accordance with the scan method indicated by the scan identifier scanIdx, and converts the input primary transform coefficients Coeff_P into a one-dimensional vector X 1d For example, if the primary transform coefficients Coeff_P are a 4×4 matrix with sparse non-zero coefficients as shown in the following formula (1) and the scan identifier scanIdx indicates a horizontal scan (hor), the rasterization unit 21 converts the primary transform coefficients Coeff_P into a one-dimensional vector X 1d Convert to.
[0030]
number
number
[0031] The matrix calculation unit 22 calculates the one-dimensional vector X obtained as above. 1d For example, a matrix operation such as the following equation (3) is performed on the one-dimensional vector X shown in the above equation (2) using the matrix R of the secondary transformation. 1d By performing this matrix operation on 1dis obtained.
[0032] Y 1d T = R X 1d T ···(3)
number
[0033] The scaling unit 23 scales the one-dimensional vector Y 1d To normalize the norm, a bit shift operation of N bits (N is a natural number) is performed as shown in the following equation (5). For example, the one-dimensional vector Y 1d By performing this bit shift operation on the one-dimensional vector Z as shown in the following equation (6), 1d is obtained.
[0034] Z 1d = ( Y 1d )>>N ···(5)
number
[0035] The matrix generator 24 converts the one-dimensional vector Z obtained as above into 1d is converted into a matrix based on the scan method specified by the scan identifier scanIdx. This matrix is supplied to a subsequent processing unit (e.g., a quantization unit) as a secondary transform coefficient Coeff obtained by secondary transforming the primary transform coefficient Coeff_P. For example, the one-dimensional vector Z 1d By performing this matrix transformation on the above, secondary transform coefficients Coeff of a 4×4 matrix as shown in the following equation (7) are obtained.
[0036]
number
[0037] As described above, when the primary transform coefficients Coeff_P having sparse non-zero coefficients are subjected to secondary transform, the coefficients are diffused from low-order to high-order components, which may reduce energy compaction and decrease coding efficiency.
[0038] In addition, Non-Patent Document 2 discloses that in order to reduce the overhead of the secondary transform flag, if the number of non-zero coefficients in a transform block is equal to or less than a predetermined threshold, the secondary transform is not applied and the flag signal is omitted. For example, if the threshold TH = 2, and the primary transform coefficient Coeff_P is a matrix as shown in the above equation (1), the secondary transform is skipped (omitted).
[0039] However, in the methods described in Non-Patent Document 1 and Non-Patent Document 2, the transform block is subjected to secondary transform for each sub-block. Therefore, for example, as shown in FIG. 6, if the transform block of the primary transform coefficient Coeff_P is composed of 2×2 sub-blocks, and each sub-block is a 4×4 matrix as shown in the above formula (1), the number of non-zero coefficients in the transform block is 4, which is greater than the threshold TH (= 2), and the secondary transform is applied. As described above, the secondary transform is performed for each sub-block (each 4×4 matrix), and the secondary transform coefficient Coeff of each sub-block is as shown in the above formula (7). In other words, there is a risk that the coefficients will be diffused from low-order to high-order components, reducing energy compaction and reducing coding efficiency.
[0040] <Sub-block-level secondary conversion skip control> Therefore, the skipping of the transform process for the transform coefficients obtained from the prediction residual, which is the difference between an image and a predicted image of that image, is controlled for each sub-block based on the number of non-zero coefficients of the transform coefficients for each sub-block.Furthermore, the skipping of the inverse transform process for the transform coefficients from which the prediction residual, which is the difference between an image and a predicted image of that image, is obtained by performing the inverse transform process is controlled for each sub-block based on the number of non-zero coefficients of the transform coefficients for each sub-block.For example, if the number of non-zero coefficients in a sub-block is equal to or less than a threshold, the secondary transform and the inverse secondary transform are skipped (omitted).
[0041] By doing this, it is possible to prevent the application of transform processing (inverse transform processing) to transform coefficients of sub-blocks with sparse non-zero coefficients, thereby preventing a decrease in energy compaction and a decrease in coding efficiency.
[0042] <Image encoding device> Fig. 7 is a block diagram showing an example of the configuration of an image coding device, which is one aspect of an image processing device to which the present technology is applied. The image coding device 100 shown in Fig. 7 is a device that codes a prediction residual between an image and its predicted image, like AVC or HEVC. For example, the image coding device 100 implements a technology proposed in HEVC or a technology proposed by JVET (Joint Video Exploration Team).
[0043] Note that Fig. 7 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 7. In other words, in the image encoding device 100, there may be processing units that are not shown as blocks in Fig. 7, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 7.
[0044] As shown in FIG. 7, the image encoding device 100 includes a control unit 101, a calculation unit 111, a transformation unit 112, a quantization unit 113, an encoding unit 114, an inverse quantization unit 115, an inverse transformation unit 116, a calculation unit 117, a frame memory 118, and a prediction unit 119.
[0045] The control unit 101 divides a video image input to the image encoding device 100 into blocks of processing units (CU, PU, transform block (TB), etc.) based on an externally or pre-specified block size of the processing unit, and supplies an image I corresponding to the divided blocks to the calculation unit 111. The control unit 101 also determines encoding parameters (header information Hinfo, prediction mode information Pinfo, transform information Tinfo, etc.) to be supplied to each block based on, for example, RDO (Rate-Distortion Optimization). The determined encoding parameters are supplied to each block.
[0046] The header information Hinfo includes information such as a video parameter set (VPS (Video Parameter Set)), a sequence parameter set (SPS (Sequence Parameter Set)), a picture parameter set (PPS (Picture Parameter Set)), a slice header (SH), etc. For example, the header information Hinfo includes information that specifies an image size (horizontal width PicWidth, vertical width PicHeight), bit depth (luminance bitDepthY, chrominance bitDepthC), a maximum value MaxCUSize / minimum value MinCUSize of CU size, a maximum value MaxTBSize / minimum value MinTBSize of transform block size, a maximum value MaxTSSize of transform skip blocks (also referred to as maximum transform skip block size), an on / off flag (also referred to as a valid flag) of each encoding tool, etc. Of course, the contents of the header information Hinfo are arbitrary, and any information other than the above-mentioned examples may be included in the header information Hinfo.
[0047] The prediction mode information Pinfo includes, for example, PU size PUSize, which is information indicating the PU size (prediction block size) of the PU to be processed, intra prediction mode information IPinfo, which is information relating to the intra prediction mode of the block to be processed (for example, prev_intra_luma_pred_flag, mpm_idx, rem_intra_pred_mode, etc. in JCTVC-W1005, 7.3.8.5 Coding Unit syntax), motion prediction information MVinfo, which is information relating to the motion prediction of the block to be processed (for example, merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X={0,1}, mvd, etc. in JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax). Of course, the contents of the prediction mode information Pinfo are arbitrary, and any information other than the above-mentioned examples may be included in this prediction mode information Pinfo.
[0048] The conversion information Tinfo includes, for example, the following information:
[0049] The block size TBSize (or the logarithm of TBSize with base 2 as the base, log2TBSize, also referred to as the transform block size) is information indicating the block size of the transform block to be processed.
[0050] The secondary transformation identifier (st_idx) is an identifier that indicates which secondary transformation or inverse secondary transformation (also called (inverse) secondary transformation) is to be applied to the target data unit (see, for example, JVET-B1001, 2.5.2 Secondary Transforms. In JEM2, it is also called nsst_idx, rot_idx). In other words, this secondary transformation identifier is information about the content of the (inverse) secondary transformation of the target data unit.
[0051] For example, the secondary transform identifier st_idx is an identifier that specifies a matrix of an (inverse) secondary transform when its value is greater than 0. In other words, in this case, the secondary transform identifier st_idx indicates execution of an (inverse) secondary transform. Also, for example, the secondary transform identifier st_idx indicates skipping of the (inverse) secondary transform when its value is 0.
[0052] The scan identifier (scanIdx) is information about the scanning method. The quantization parameter (qp) is information indicating the quantization parameter used for (de)quantization of the target data unit. The quantization matrix (scaling_matrix) is information indicating the quantization matrix used for (de)quantization of the target data unit (for example, JCTVC-W1005, 7.3.4 Scaling list data syntax).
[0053] Of course, the content of the conversion information Tinfo is arbitrary, and any information other than the above-mentioned examples may be included in this conversion information Tinfo.
[0054] The header information Hinfo is supplied to, for example, each block. The prediction mode information Pinfo is supplied to, for example, the encoding unit 114 and the prediction unit 119. The transformation information Tinfo is supplied to, for example, the transformation unit 112, the quantization unit 113, the encoding unit 114, the inverse quantization unit 115, and the inverse transformation unit 116.
[0055] The calculation unit 111 subtracts the predicted image P supplied from the prediction unit 119 from the image I corresponding to the block of the input processing unit as shown in equation (8) to obtain a prediction residual D, and supplies it to the conversion unit 112.
[0056] D=IP ···(8)
[0057] The transform unit 112 performs a transform process on the prediction residual D supplied from the calculation unit 111 based on the transform information Tinfo supplied from the control unit 101, and derives a transform coefficient Coeff. The transform unit 112 supplies the transform coefficient Coeff to the quantization unit 113.
[0058] The quantization unit 113 scales (quantizes) the transform coefficients Coeff supplied from the transform unit 112, based on the transform information Tinfo supplied from the control unit 101. That is, the quantization unit 113 quantizes the transform coefficients Coeff that have been subjected to the transform process. The quantization unit 113 supplies the quantized transform coefficients obtained by the quantization, i.e., the quantized transform coefficient levels level, to the encoding unit 114 and the inverse quantization unit 115.
[0059] The encoding unit 114 uses a predetermined method to encode the quantized transform coefficient level etc. supplied from the quantization unit 113. For example, the encoding unit 114 converts the encoding parameters (header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, etc.) supplied from the control unit 101 and the quantized transform coefficient level supplied from the quantization unit 113 into syntax values of each syntax element in accordance with the definition of a syntax table, encodes each syntax value (for example, arithmetic coding), and generates a bit string (encoded data).
[0060] Furthermore, the encoding unit 114 derives residual information RInfo from the quantized transform coefficient level level, encodes the residual information RInfo, and generates a bit string (encoded data).
[0061] The residual information RInfo includes, for example, a last non-zero coefficient X coordinate (last_sig_coeff_x_pos), a last non-zero coefficient Y coordinate (last_sig_coeff_y_pos), a sub-block non-zero coefficient presence / absence flag (coded_sub_block_flag), a non-zero coefficient presence / absence flag (sig_coeff_flag), a GR1 flag (gr1_flag) that is flag information indicating whether the level of the non-zero coefficient is greater than 1, a GR2 flag (gr2_flag) that is flag information indicating whether the level of the non-zero coefficient is greater than 2, a sign sign (sign_flag) that is a sign indicating the positive or negative of the non-zero coefficient, and a non-zero coefficient residual level (coeff_abs_level_remaining) that is information indicating the residual level of the non-zero coefficient (see, for example, 7.3.8.11 Residual Coding syntax of JCTVC-W1005). Of course, the contents of the residual information RInfo are arbitrary, and any information other than the examples described above may be included in the residual information RInfo.
[0062] For example, the encoding unit 114 multiplexes the bit strings (encoded data) of the encoded syntax elements and outputs the multiplexed data as a bit stream.
[0063] The inverse quantization unit 115 scales (inverse quantizes) the value of the quantized transform coefficient level level supplied from the quantization unit 113 based on the transformation information Tinfo supplied from the control unit 101, and derives the transform coefficient Coeff_IQ after inverse quantization. The inverse quantization unit 115 supplies the transform coefficient Coeff_IQ to the inverse transform unit 116. The inverse quantization performed by this inverse quantization unit 115 is the inverse process of the quantization performed by the quantization unit 113, and is the same process as the inverse quantization performed in the image decoding device described later. Therefore, this inverse quantization will be described later in the description of the image decoding device.
[0064] The inverse transform unit 116 performs inverse transform on the transform coefficients Coeff_IQ supplied from the inverse quantization unit 115, based on the transform information Tinfo supplied from the control unit 101, to derive a prediction residual D'. The inverse transform unit 116 supplies the prediction residual D' to the calculation unit 117. The inverse transform performed by this inverse transform unit 116 is the inverse process of the transform performed by the transform unit 112, and is the same process as the inverse transform performed in an image decoding device, which will be described later. Therefore, this inverse transform will be described later in the description of the image decoding device.
[0065] The calculation unit 117 adds the prediction residual D' supplied from the inverse transform unit 116 and the predicted image P (prediction signal) corresponding to the prediction residual D' supplied from the prediction unit 119, as shown in the following equation (9), to derive a local decoded image Rec. The calculation unit 117 supplies the local decoded image Rec to the frame memory 118.
[0066] Rec=D'+P ···(9)
[0067] The frame memory 118 reconstructs a decoded image for each picture unit using the local decoded image Rec supplied from the calculation unit 117, and stores the reconstructed image in a buffer within the frame memory 118. The frame memory 118 reads a decoded image specified by the prediction unit 119 from the buffer as a reference image, and supplies the reference image to the prediction unit 119. The frame memory 118 may also store header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, and the like, related to generation of the decoded image, in a buffer within the frame memory 118.
[0068] The prediction unit 119 obtains a decoded image stored in the frame memory 118 as a reference image, which is specified by the prediction mode information PInfo, and generates a predicted image P using the reference image by the prediction method specified by the prediction mode information Pinfo. The prediction unit 119 supplies the generated predicted image P to the calculation unit 111 and the calculation unit 117.
[0069] The image coding device 100 includes a control unit that controls, for each sub-block, skipping of transform processing for transform coefficients obtained from prediction residuals, which are the difference between an image and a predicted image of that image, based on the number of non-zero coefficients of the transform coefficients for each sub-block. That is, the transform unit 112 controls, for each sub-block, skipping of transform processing for transform coefficients obtained from prediction residuals, which are the difference between an image and a predicted image of that image, based on the number of non-zero coefficients of the transform coefficients for each sub-block.
[0070] <Conversion section> 8 is a block diagram showing an example of the main configuration of conversion unit 112. In FIG. 8, conversion unit 112 includes primary conversion unit 131 and secondary conversion unit 132.
[0071] The primary transform unit 131 performs a primary transform such as an orthogonal transform on the prediction residual D supplied from the calculation unit 111, and derives transform coefficients Coeff_P (also referred to as primary transform coefficients) after the primary transform corresponding to the prediction residual D. That is, the primary transform unit 131 transforms the prediction residual D into primary transform coefficients Coeff_P. The primary transform unit 131 supplies the derived primary transform coefficients Coeff_P to the secondary transform unit 132 (the rasterization unit 141 and the switch 148, which will be described later).
[0072] The secondary conversion unit 132 converts the primary conversion coefficient Coeff_P supplied from the primary conversion unit 131 into a one-dimensional vector (also called a row vector), performs a matrix operation on the one-dimensional vector, scales the one-dimensional vector on which the matrix operation has been performed, and performs secondary conversion, which is a conversion process that converts the scaled one-dimensional vector into a matrix.
[0073] The secondary transform unit 132 performs a secondary transform on the primary transform coefficient Coeff_P based on a secondary transform identifier st_idx, which is information about the content of the secondary transform, and a scan identifier scanIdx, which is information about a scan method for the transform coefficients, to derive transform coefficients Coeff after the secondary transform (also referred to as secondary transform coefficients). That is, the secondary transform unit 132 transforms the primary transform coefficients Coeff_P into secondary transform coefficients Coeff. The secondary transform unit 132 supplies the secondary transform coefficients Coeff to the quantization unit 113.
[0074] It should be noted that the secondary transform unit 132 can also skip (omit) the secondary transform and supply the primary transform coefficients Coeff_P to the quantization unit 113 as the secondary transform coefficients Coeff.
[0075] As shown in FIG. 8, the secondary transform unit 132 includes a rasterizer 141, a matrix calculator 142, a scaler 143, a matrix generator 144, a secondary transform selector 145, a quantizer 146, a non-zero coefficient number determiner 147, and a switch 148.
[0076] The rasterization unit 141 converts the primary transform coefficients Coeff_P supplied from the primary transform unit 131 into a one-dimensional vector X 1d The rasterization unit 141 converts the obtained one-dimensional vector X 1d is supplied to the matrix calculation unit 142.
[0077] 9A shows the scan type scanType specified by each value of the scan identifier scanIdx. As shown in FIG. 9A, when the scan identifier scanIdx is 0, an up-right diagonal scan is specified, when the scan identifier scanIdx is 1, a horizontal fast scan is specified, and when the scan identifier scanIdx is 2, a vertical fast scan is specified. FIGS. 9B to 9D show the scan order of coefficients in each scan of a 4×4 sub-block. In FIGS. 9B to 9D, the numbers assigned to each coefficient position indicate the order in which the coefficient position is scanned. FIG. 9B shows an example of the scan order of a horizontal fast scan, FIG. 9C shows an example of the scan order of a vertical fast scan, and FIG. 9D shows an example of the scan order of a diagonal scan.
[0078] The secondary transform selection unit 145 reads out the matrix R of the secondary transform specified by the secondary transform identifier st_idx from an internal memory (not shown) of the secondary transform selection unit 145, and supplies it to the matrix calculation unit 142. For example, when the secondary transform identifier st_idx has a certain value, the secondary transform selection unit 145 reads out the 16×16 matrix R shown in FIG.
[0079] Note that the secondary transform selection unit 145 may select the matrix R for secondary transform according to the secondary transform identifier st_idx and intra prediction mode information IPinfo (for example, a prediction mode number). Alternatively, the secondary transform selection unit 145 may select the matrix R for secondary transform according to the motion prediction information MVinfo and the secondary transform identifier st_idx instead of the intra prediction mode information IPinfo.
[0080] The matrix calculation unit 142 calculates a one-dimensional vector X1d and the secondary transformation matrix R, perform the matrix operation shown in the following equation (10), and the resulting one-dimensional vector Y 1d is supplied to the scaling unit 143. In equation (10), the operator "·" represents an operation for performing an inner product (matrix multiplication) between matrices, and the operator "T" represents an operation for performing a transpose of a matrix.
[0081] Y 1d T =R·X 1d T ···(10)
[0082] The scaling unit 143 scales the signal Y 1d (i.e., one-dimensional vector Y 1d ) is normalized by performing a bit shift operation of N bits (N is a natural number) as shown in the following equation (11), and the signal Z after the bit shift is 1d (i.e., one-dimensional vector Z 1d ) is calculated. As shown in the following equation (12), before the N-bit shift operation, the value of 1<<(N-1) is added as an offset to the one-dimensional vector Z 1d It is also possible to add to each element of
[0083] Z 1d =(Y 1d )>>N ···(11) Z 1d =(Y 1d +((N-1)<<1)·E)>>N ···(12)
[0084] In equation (12), E is a 1×16-dimensional vector in which all elements have a value of 1. For example, since the matrix R of the secondary transformation shown in FIG. 10 is an 8-bit scaled matrix, the value of N used for norm normalization in the scaling unit 143 is 8. Generally, when the matrix R of the secondary transformation is scaled by N bits, the bit shift amount for norm normalization is N bits. The scaling unit 143 scales the one-dimensional vector Z obtained as above. 1d is supplied to the matrix generator 144.
[0085] The matrix generator 144 generates a norm-normalized 1×16-dimensional vector Z 1d into a 4 × 4 matrix. The matrix generator 144 supplies the obtained transform coefficients Coeff to the quantizer 146 and the switch 148.
[0086] The quantization unit 146 performs basically the same processing (quantization) as the quantization unit 113. However, the quantization unit 146 receives as input the primary transform coefficient Coeff_P, the secondary transform coefficient Coeff (after the processing from the rasterization unit 141 to the matrix generation unit 144 has been executed), a quantization parameter qp that is part of the transformation information Tinfo, and a quantization matrix scaling_matrix. The quantization unit 146 quantizes the primary transform coefficient Coeff_P supplied from the primary transformation unit 131 and the secondary transform coefficient Coeff, which is a transformation coefficient supplied from the matrix generation unit 144, by referring to the quantization parameter qp and the quantization matrix scaling_matrix, as shown in, for example, the following equations (13) and (14), and derives the quantized primary transform coefficient level_P (referred to as the quantized primary transform coefficient) and the quantized secondary transform coefficient level_S (referred to as the quantized secondary transform coefficient).
[0087] level_P(i,j) = sign (coeff_P(i,j)) × (( abs ( coeff_P(i,j)) × f[qp%6] × (16 / w(i,j)) + offsetQ ) >> qp / 6 ) >> shift1 ···(13) level_S(i,j) = sign (coeff (i,j)) × (( abs ( coeff (i,j)) × f[qp%6] × (16 / w(i,j)) + offsetQ ) >> qp / 6 ) >> shift1 ···(14)
[0088] In equations (13) and (14), the operator sign(X) is an operator that returns the positive or negative sign of the input value X. For example, if X >= 0, it returns +1, and if X < 0, it returns -1. Also, f[] is a scaling factor that depends on the quantization parameter, and takes a value as shown in the following equation (15).
[0089] F[qp%6] = [26214, 23302, 20560, 18396, 16384, 16384, 14564] ···(15)
[0090] In addition, in equations (13) and (14), w(i,j) is the value of the quantization matrix scaling_matrix corresponding to the coefficient position (i,j) of Coeff(i,j) or (Coeff_P(i,j)). That is, this w(i,j) is calculated as shown in equation (16) below. Furthermore, shift1 and offsetQ are calculated as shown in equations (17) and (18) below.
[0091] w(i,j) = scaling_matrix(i,j) ···(16) shift1 = 29 - M - B ···(17) offsetQ = 28 - M - B ···(18) In equations (17) and (18), M is the logarithm of the block size TBSize of the transform block in base 2, and B is the bit depth bitDepth of the input signal. M = log2(TBSize) B = bitDepth
[0092] It should be noted that the quantization method can be changed within a practicable range, regardless of the quantization shown in equations (13) and (14).
[0093] The quantization unit 146 supplies the quantized primary transform coefficients level_P and the quantized secondary transform coefficients level_S obtained as described above to the non-zero coefficient number determination unit 147.
[0094] The non-zero coefficient number determination unit 147 receives the quantized primary transform coefficients level_P and the quantized secondary transform coefficients level_S for each sub-block. The non-zero coefficient number determination unit 147 references the quantized primary transform coefficients level_P and the quantized secondary transform coefficients level_S supplied from the quantization unit 146, and derives the numbers of non-zero coefficients in the sub-block, numSigInSBK_P (also referred to as the number of non-zero coefficients of the quantized primary transform coefficients) and numSingInSBK_S (also referred to as the number of non-zero coefficients of the quantized secondary transform coefficients), respectively, using, for example, the following equations (19) and (20). In equations (19) and (20), (i, j) represent coordinates within the sub-block, where i = 0...3, j = 0...3. The operator abs(X) returns the absolute value of the input value X. By using equations (19) and (20), it is possible to derive the number of transform coefficients (non-zero coefficients) whose level values after quantization are greater than 0. In equations (19) and (20), the determination condition for a non-zero coefficient (abs(level_X(i,j))>0) (X=P, S) may be replaced with the determination condition (level_X(i,j)!=0).
[0095] numSigInSBK_P = Σ{abs(level_P(i,j))>0 ? 1 : 0} ···(19) numSigInSBK_S = Σ{abs(level_S(i,j))>0 ? 1 : 0} ···(20)
[0096] The non-zero coefficient number determination unit 147 derives a secondary transform skip flag StSkipFlag, which is information regarding the skipping of the secondary transform, using the following equation (21) by referring to the number of non-zero coefficients numSigInSBK_P of the quantized primary transform coefficients, the number of non-zero coefficients numSigInSBK_S of the quantized secondary transform coefficients, and a predetermined threshold value TH.
[0097] StSkipFlag = ( numSigInSBK_P <= numSigInSBK_S && numSigInSBK_P <= TH ) ? 1 : 0 ···(twenty one)
[0098] The threshold value TH is set to, for example, 2, but is not limited to this and can be set to a value between 0 and 16. The threshold value TH may also be notified in header information such as a VPS / SPS / PPS / slice header SH. The threshold value TH may also be determined in advance between the encoding side (e.g., image encoding device 100) and the decoding side (e.g., image decoding device 200 described later), and notification from the encoding side to the decoding side (transmission of the threshold value TH from the encoding side to the decoding side) may be omitted.
[0099] In equation (21), if the number of non-zero coefficients of the quantized primary transform coefficients, numSigInSBK_P, is less than or equal to the number of non-zero coefficients of the quantized secondary transform coefficients, numSigInSBK_S, and if the number of non-zero coefficients of the quantized primary transform coefficients, numSigInSBK_P, is less than or equal to the threshold value, TH, the value of the secondary transform skip, StSkipFlag, is set to 1. That is, the secondary transform skip, StSkipFlag, indicates that the secondary transform is skipped. Otherwise (numSigInSBK_P>numSigInSBK_S || numSigInSBK_P>TH), the value of the secondary transform skip, StSkipFlag, is set to 0. That is, the secondary transform skip, StSkipFlag, indicates that the secondary transform is performed.
[0100] Note that the following equation (22) or equation (23) may be used instead of the above equation (21): When equation (23) is used, the quantization process of the secondary transform coefficients may be omitted.
[0101] StSkipFlag = ( numSigInSBK_P <= numSigInSBK_S && numSigInSBK_S <= TH ) ? 1 : 0 ···(twenty two) StSkipFlag = ( numSigInSBK_P <= TH ) ? 1 : 0 ···(twenty three)
[0102] The non-zero coefficient number determination unit 147 supplies the derived secondary transform skip flag StSkipFlag to the switch 148 .
[0103] The switch 148 receives the primary transform coefficient Coeff_P, the secondary transform coefficient Coeff, and the secondary transform skip flag StSkipFlag for each sub-block. The switch 148 controls skipping of the secondary transform in accordance with the secondary transform skip flag StSkipFlag supplied from the non-zero coefficient number determination unit 147.
[0104] For example, when the value of the secondary transform skip flag StSkipFlag is 0, that is, when the secondary transform skip StSkipFlag indicates that the secondary transform is to be performed, the switch 148 performs the secondary transform. That is, the switch 148 supplies the secondary transform coefficient Coeff supplied from the matrixing unit 144 to the quantization unit 113. Also, for example, when the value of the secondary transform skip flag StSkipFlag is 1, that is, when the secondary transform skip StSkipFlag indicates that the secondary transform is to be skipped, the switch 148 skips the secondary transform. That is, the switch 148 supplies the primary transform coefficient Coeff_P supplied from the primary transform unit 131 to the quantization unit 113 as the secondary transform coefficient Coeff.
[0105] The switch 148 can also control the skipping of the secondary transform based on the secondary transform identifier st_idx. For example, when the secondary transform identifier st_idx is 0 (indicating the skipping of the secondary transform), the switch 148 skips the secondary transform regardless of the value of the secondary transform skip flag StSkipFlag. That is, the switch 148 supplies the primary transform coefficients Coeff_P supplied from the primary transform unit 131 to the quantization unit 113 as the secondary transform coefficients Coeff. Also, for example, when the secondary transform identifier st_idx is greater than 0 (indicating the execution of the secondary transform), the switch 148 controls the secondary transform as described above by referring to the secondary transform skip flag StSkipFlag.
[0106] As described above, the nonzero coefficient number determination unit 147 sets the secondary transform skip flag StSkipFlag based on the number of nonzero coefficients for each sub-block, and the switch 148 controls the skipping of the secondary transform based on the secondary transform skip flag StSkipFlag. In this way, it becomes possible to skip the secondary transform for sub-blocks with sparse nonzero coefficients, thereby suppressing a decrease in energy compaction and a decrease in coding efficiency.
[0107] <Image encoding process flow> Next, an example of the flow of each process executed by the image encoding device 100 will be described. First, an example of the flow of the image encoding process will be described with reference to the flowchart in FIG.
[0108] When the image encoding process is started, in step S101, the control unit 101 performs an encoding control process, performs block division, sets encoding parameters, and so on.
[0109] In step S102, the prediction unit 119 performs a prediction process to generate a predicted image etc. in an optimal prediction mode. For example, in this prediction process, the prediction unit 119 performs intra prediction to generate a predicted image etc. in an optimal intra prediction mode, performs inter prediction to generate a predicted image etc. in an optimal inter prediction mode, and selects an optimal prediction mode from among them based on a cost function value etc.
[0110] In step S103, the calculation unit 111 calculates the difference between the input image and the predicted image of the optimal mode selected by the prediction process in step S102. That is, the calculation unit 111 generates a prediction residual D between the input image and the predicted image. The prediction residual D calculated in this way has a reduced data amount compared to the original image data. Therefore, the data amount can be compressed compared to when the image is encoded as is.
[0111] In step S104, the transform unit 112 performs a transform process on the prediction residuals D generated in the process of step S103 to derive transform coefficients Coeff. Details of the process of step S104 will be described later.
[0112] In step S105, the quantization unit 113 quantizes the transform coefficient Coeff obtained by the processing in step S104, for example, by using the quantization parameter calculated by the control unit 101, and derives the quantized transform coefficient level level.
[0113] In step S106, the inverse quantization unit 115 inverse quantizes the quantized transform coefficient level generated by the process of step S105 with characteristics corresponding to the quantization characteristics of step S105, to derive the transform coefficient Coeff_IQ.
[0114] In step S107, the inverse transform unit 116 inversely transforms the transform coefficient Coeff_IQ obtained by the process of step S106 using a method corresponding to the transform process of step S104, to derive a prediction residual D'. Note that this inverse transform process is the inverse process of the transform process of step S104, and is executed in the same manner as the inverse transform process executed in the image decoding process described later. Therefore, this inverse transform process will be described in the description of the decoding side.
[0115] In step S108, the calculation unit 117 generates a locally decoded image by adding the prediction image obtained by the prediction process in step S102 to the prediction residual D' derived in the process of step S107.
[0116] In step S109, the frame memory 118 stores the locally decoded image obtained by the process of step S108.
[0117] In step S110, the encoding unit 114 encodes the quantized transform coefficient level LEVEL obtained by the process of step S105. For example, the encoding unit 114 encodes the quantized transform coefficient level LEVEL, which is information related to the image, by arithmetic coding or the like to generate encoded data. At this time, the encoding unit 114 also encodes various encoding parameters (header information HInfo, prediction mode information PInfo, and transform information TInfo). Furthermore, the encoding unit 114 derives residual information RInfo from the quantized transform coefficient level LEVEL and encodes the residual information RInfo. The encoding unit 114 compiles the encoded data of the various information thus generated and outputs it as a bitstream to the outside of the image encoding device 100. This bitstream is transmitted to the decoding side, for example, via a transmission path or a recording medium.
[0118] When the process of step S110 ends, the image encoding process ends.
[0119] The processing units of these processes are arbitrary and do not have to be the same. Therefore, the processing of each step can be executed in parallel with the processing of other steps, or the processing order can be changed as appropriate.
[0120] <Conversion process flow> Next, an example of the flow of the conversion process executed in step S104 in FIG. 11 will be described with reference to the flowchart in FIG.
[0121] When the transform process starts, in step S121, the primary transform unit 131 performs primary transform on the prediction residual D based on the primary transform identifier pt_idx, and derives the primary transform coefficient Coeff_P.
[0122] In step S122, the secondary transform unit 132 (switch 148) determines whether the secondary transform identifier st_idx applies the secondary transform (st_idx>0). If it is determined that the secondary transform identifier st_idx is 0 (indicating skipping of the secondary transform), the secondary transform (processing of steps S123 to S134) is skipped, the transform process ends, and the process returns to FIG. 11. That is, the secondary transform unit 132 (switch 148) supplies the primary transform coefficient Coeff_P to the quantization unit 113 as the transform coefficient Coeff.
[0123] Also, if it is determined in step S122 that the secondary transformation identifier st_idx is greater than 0 (indicating that secondary transformation has been performed), the process proceeds to step S123.
[0124] In step S123, the secondary transform selection unit 145 selects the matrix R of the secondary transform specified by the secondary transform identifier st_idx.
[0125] In step S124, the secondary transform unit 132 divides the transform block to be processed into sub-blocks, and selects unprocessed sub-blocks.
[0126] In step S125, the rasterization unit 141 converts the primary transform coefficients Coeff_P into a one-dimensional vector X based on the scan method specified by the scan identifier scanIdx. 1d Convert to.
[0127] In step S126, the matrix calculation unit 142 calculates the one-dimensional vector X 1d and the secondary transformation matrix R to obtain the one-dimensional vector Y 1d Ask for.
[0128] In step S127, the scaling unit 143 calculates the one-dimensional vector Y 1d Normalize the norm of the 1-dimensional vector Z 1d Ask for.
[0129] In step S128, the matrix generator 144 generates a one-dimensional vector Z 1d is converted into a 4×4 matrix to obtain the secondary transform coefficient Coeff of the sub-block to be processed.
[0130] In step S129, the quantization unit 146 quantizes each of the primary transform coefficient Coeff_P and the secondary transform coefficient Coeff by referring to the quantization parameter qp and the quantization matrix scaling_matrix, and derives the quantized primary transform coefficient level_P and the quantized secondary transform coefficient level_S.
[0131] In step S130, the non-zero coefficient number determination unit 147 derives the secondary transform skip flag StSkipFlag for each sub-block based on the quantized primary transform coefficients level_P, the quantized secondary transform coefficients level_S, and the threshold value TH as described above.
[0132] In step S131, the switch 148 determines whether the secondary conversion skip flag StSkipfFlag derived in step S130 indicates skipping of the secondary conversion. If it indicates execution of the secondary conversion, that is, if the value of the secondary conversion skip flag StSkipFlag is 0, the process proceeds to step S132.
[0133] In step S132, the switch 148 outputs the secondary transform coefficient Coeff obtained by the process of step S128 (supplies it to the quantization unit 113). When the process of step S132 ends, the process proceeds to step S134.
[0134] Also, in step S131, if the secondary transformation skip flag StSkipFlag is set to 1, indicating that the secondary transformation is to be skipped, the process proceeds to step S133.
[0135] In step S133, the switch 148 outputs the primary transform coefficient Coeff_P obtained by the processing in step S121 as the secondary transform coefficient Coeff (supplies it to the quantization unit 113). When the processing in step S133 ends, the process proceeds to step S134.
[0136] In step S134, the secondary transform unit 132 determines whether all sub-blocks of the transform block being processed have been processed. If it is determined that unprocessed sub-blocks exist, the process returns to step S124, and the subsequent processes are repeated. That is, the processes (secondary transform) of steps S124 to S134 are executed for each sub-block of the transform block being processed. If it is determined in step S134 that all sub-blocks have been processed (secondary transforms have been executed or skipped for all sub-blocks), the transform process ends, and the process returns to FIG. 11.
[0137] The order of the steps of the transformation process may be changed or the content of the process may be modified to the extent possible. For example, if it is determined in step S122 that the secondary transformation identifier st_idx=0, a 16×16 identity matrix may be selected as the matrix R of the secondary transformation, and the processes of steps S124 to S134 may be executed.
[0138] By performing each process as described above, it is possible to control the skipping (execution) of secondary transform in units of sub-blocks. Therefore, it is possible to suppress a decrease in energy compaction for residual signals with sparse non-zero coefficients. That is, it is possible to suppress a decrease in coding efficiency. In other words, it is possible to suppress an increase in the load of coding (secondary transform and inverse secondary transform) while suppressing a decrease in coding efficiency.
[0139] <Image decoding device> Next, decoding of coded data coded as above will be described. Fig. 13 is a block diagram showing an example of the configuration of an image decoding device, which is one aspect of an image processing device to which the present technology is applied. The image decoding device 200 shown in Fig. 13 is an image decoding device corresponding to the image coding device 100 of Fig. 7, and decodes coded data (bitstream) generated by the image coding device 100 using a decoding method corresponding to the coding method used by the image coding device 100. For example, the image decoding device 200 implements a technology proposed in HEVC or a technology proposed in JVET.
[0140] Note that Fig. 13 shows the main processing units, data flows, etc., and does not necessarily show everything. That is, in the image decoding device 200, there may be processing units that are not shown as blocks in Fig. 13, and there may be processing or data flows that are not shown as arrows or the like in Fig. 13.
[0141] 13, the image decoding device 200 includes a decoding unit 211, an inverse quantization unit 212, an inverse transform unit 213, a calculation unit 214, a frame memory 215, and a prediction unit 216. Encoded data generated by the image encoding device 100 or the like is supplied to the image decoding device 200 as, for example, a bit stream via, for example, a transmission medium or a recording medium.
[0142] The decoding unit 211 decodes the supplied coded data using a predetermined decoding method corresponding to the coding method. For example, the decoding unit 211 decodes the syntax value of each syntax element from the bit string of the supplied coded data (bit stream) according to the definition of the syntax table. The syntax elements include, for example, header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, residual information Rinfo, and the like.
[0143] The decoding unit 211 derives the quantized transform coefficient level "level" at each coefficient position in each transform block by referring to the residual information Rinfo. The decoding unit 211 supplies the prediction mode information Pinfo, the quantized transform coefficient level "level", and the transform information Tinfo obtained by decoding to each block. For example, the decoding unit 211 supplies the prediction mode information Pinfo to the prediction unit 216, supplies the quantized transform coefficient level "level" to the inverse quantization unit 212, and supplies the transform information Tinfo to the inverse quantization unit 212 and the inverse transform unit 213.
[0144] The inverse quantization unit 212 scales (inverse quantizes) the value of the quantized transform coefficient level level supplied from the decoding unit 211 based on the transform information Tinfo supplied from the decoding unit 211, and derives the transform coefficient Coeff_IQ after inverse quantization. This inverse quantization is the inverse process of the quantization performed by the quantization unit 113 (FIG. 7) of the image encoding device 100. Note that the inverse quantization unit 115 (FIG. 7) performs the same inverse quantization as the inverse quantization unit 212. The inverse quantization unit 212 supplies the obtained transform coefficient Coeff_IQ to the inverse transform unit 213.
[0145] The inverse transform unit 213 inversely transforms the transform coefficients Coeff_IQ supplied from the inverse quantization unit 212, based on the transform information Tinfo supplied from the decoding unit 211, to derive a prediction residual D'. This inverse transform is the inverse process of the transform process performed by the transform unit 112 (FIG. 7) of the image encoding device 100. It should be noted that the inverse transform unit 116 performs the same inverse transform as this inverse transform unit 213. Details of this inverse transform will be described later. The inverse transform unit 213 supplies the obtained prediction residual D' to the calculation unit 214.
[0146] The calculation unit 214 adds the prediction residual D' supplied from the inverse transform unit 213 to a predicted image P (prediction signal) corresponding to the prediction residual D', as shown in the following equation (24), to derive a local decoded image Rec. The calculation unit 214 reconstructs a decoded image for each picture using the obtained local decoded image Rec, and outputs the obtained decoded image to the outside of the image decoding device 200. The calculation unit 214 also supplies the local decoded image Rec to the frame memory 215.
[0147] Rec=D'+P ···(twenty four)
[0148] The frame memory 215 reconstructs a decoded image for each picture unit using the local decoded image Rec supplied from the calculation unit 214, and stores the reconstructed image in a buffer within the frame memory 215. The frame memory 215 reads a decoded image specified by prediction mode information Pinfo of the prediction unit 216 from the buffer as a reference image, and supplies the read image to the prediction unit 216. The frame memory 215 may also store header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, and the like related to the generation of the decoded image in a buffer within the frame memory 215.
[0149] The prediction unit 216 obtains, as a reference image, a decoded image stored in the frame memory 215, which is specified by the prediction mode information PInfo supplied from the decoding unit 211, and generates a predicted image P using the reference image by the prediction method specified by the prediction mode information Pinfo. The prediction unit 216 supplies the generated predicted image P to the calculation unit 214.
[0150] The image decoding device 200 includes a control unit that controls, for each sub-block, skipping of inverse transform processing for transform coefficients that, when inverse transformed, result in prediction residuals, which are the differences between an image and its predicted image, based on the number of non-zero coefficients of the transform coefficients for each sub-block. That is, the inverse transform unit 213 controls, for each sub-block, skipping of inverse transform processing for transform coefficients that, when inverse transformed, result in prediction residuals, which are the differences between an image and its predicted image, based on the number of non-zero coefficients of the transform coefficients for each sub-block.
[0151] <Inverse conversion section> Fig. 14 is a block diagram showing an example of the main configuration of inverse conversion unit 213 in Fig. 13. As shown in Fig. 14, inverse conversion unit 213 has an inverse secondary conversion unit 231 and an inverse primary conversion unit 232.
[0152] The inverse secondary transform unit 231 converts the transform coefficients Coeff_IQ supplied from the inverse quantization unit 212, i.e., the transform coefficients Coeff_IQ (also referred to as secondary transform coefficients) obtained by decoding and inverse quantizing the encoded data, into a one-dimensional vector, performs a matrix operation on the one-dimensional vector, scales the one-dimensional vector on which the matrix operation has been performed, and performs an inverse secondary transform, which is a conversion process that converts the scaled one-dimensional vector into a matrix.
[0153] The inverse secondary transform unit 231 performs an inverse secondary transform on the secondary transform coefficients Coeff_IQ based on a secondary transform identifier st_idx, which is information about the content of the secondary transform, and a scan identifier scanIdx, which is information about a method for scanning the transform coefficients, to derive transform coefficients Coeff_IS (also referred to as primary transform coefficients) after the inverse secondary transform. That is, the inverse secondary transform unit 231 converts the secondary transform coefficients Coeff_IQ into primary transform coefficients Coeff_IS. The inverse secondary transform unit 231 supplies the primary transform coefficients Coeff_IS to the inverse primary transform unit 232.
[0154] Note that the inverse secondary transform unit 231 can also skip (omit) the inverse secondary transform and supply the secondary transform coefficients Coeff_IQ as the primary transform coefficients Coeff_IS to the inverse primary transform unit 232. Details of the inverse secondary transform unit 231 will be described later.
[0155] The inverse primary transform unit 232 performs an inverse primary transform such as an inverse orthogonal transform on the primary transform coefficients Coeff_IS supplied from the inverse secondary transform unit 231, and derives prediction residuals D'. That is, the inverse primary transform unit 232 transforms the primary transform coefficients Coeff_IS into prediction residuals D'. The inverse primary transform unit 232 supplies the derived prediction residuals D' to the calculation unit 214.
[0156] Next, a description will be given of the inverse secondary transform unit 231. As shown in Fig. 14, the inverse secondary transform unit 231 has a non-zero coefficient number determination unit 241, a switch 242, a rasterization unit 243, a matrix operation unit 244, a scaling unit 245, a matrix generation unit 246, and an inverse secondary transform selection unit 247.
[0157] The non-zero coefficient number determination unit 241 receives the secondary transform coefficient Coeff_IQ for each sub-block as input. The non-zero coefficient number determination unit 241 derives the number of non-zero coefficients numSigInSBK in the sub-block (also referred to as the number of non-zero coefficients of the transform coefficients) by referring to the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212, for example, as shown in the following equation (25). Note that in equation (25), (i, j) in the secondary transform coefficient Coeff_IQ(i, j) represents coordinates in the sub-block, where i = 0...3, j = 0...3. The operator abs(X) returns the absolute value of the input value X.
[0158] numSigInSBK = Σ{abs(Coeff_IQ(i,j))>0 ? 1 : 0} ···(twenty five)
[0159] Note that the number of non-zero coefficients numSigInSBK of the transform coefficients may be derived based on the non-zero coefficient presence / absence flag sig_coeff_flag without referring to the secondary transform coefficient Coeff_IQ.
[0160] Then, the non-zero coefficient number determination unit 241 determines whether the number of non-zero coefficients numSigInSBK of the transform coefficients is less than or equal to a predetermined threshold TH, and derives the secondary transform skip flag StSkipFlag based on the determination result, as shown in the following equation (26).
[0161] StSkipFlag = numSigInSBK <= TH ? 1 : 0 ···(26)
[0162] This threshold value TH may be, for example, 2 or any value from 0 to 16. Furthermore, the threshold value TH may be notified from outside (for example, the encoding side or the control side) in header information such as a VPS / SPS / PPS / slice header SH. Furthermore, the threshold value TH may be agreed upon in advance between the encoding side (for example, the image encoding device 100) and the decoding side (for example, the image decoding device 200 described later), and notification from the encoding side to the decoding side (transmission of the threshold value TH from the encoding side to the decoding side) may be omitted.
[0163] In equation (26), for example, if the value of the number of non-zero coefficients of the transform coefficients numSigInSBK is equal to or less than the threshold value TH, the value of the secondary transform skip flag StSkipFlag is set to 1. Also, for example, if the value of the number of non-zero coefficients of the transform coefficients numSigInSBK is greater than the threshold value TH, the value of the secondary transform skip flag StSkipFlag is set to 0.
[0164] The non-zero coefficient number determination unit 241 supplies the derived StSkipFlag to the switch 242 .
[0165] The switch 242 receives the secondary transform coefficients Coeff_IQ for each sub-block and the secondary transform skip flag StSkipFlag as inputs. The switch 242 controls skipping of the inverse secondary transform in accordance with the secondary transform skip flag StSkipFlag supplied from the non-zero coefficient number determination unit 241.
[0166] For example, when the value of the secondary transform skip flag StSkipFlag is 0, that is, when the secondary transform skip StSkipFlag indicates that the secondary transform is to be performed, the switch 242 performs the secondary transform. That is, the switch 242 supplies the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212 to the rasterization unit 243. Also, for example, when the value of the secondary transform skip flag StSkipFlag is 1, that is, when the secondary transform skip StSkipFlag indicates that the inverse secondary transform is to be skipped, the switch 242 skips the inverse secondary transform. That is, the switch 242 supplies the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212 to the inverse primary transform unit 232 as the primary transform coefficient Coeff_IS.
[0167] The rasterizing unit 243 converts the transform coefficients Coeff_IQ supplied from the switch 242 into a one-dimensional vector X for each sub-block (4×4 sub-block) based on the transform coefficient scanning method specified by the scan identifier scanIdx supplied from the decoding unit 211. 1d The rasterization unit 243 converts the obtained one-dimensional vector X 1d is supplied to the matrix calculation unit 244.
[0168] The inverse secondary transform selection unit 247 selects the matrix IR (=R T ) from an internal memory (not shown) of the inverse secondary transform selection unit 247 and supplies it to the matrix operation unit 244. For example, when the value of a certain secondary transform identifier st_idx is given, the inverse secondary transform selection unit 247 selects the transposed matrix R of the 16×16 matrix R shown in FIG. 10 as the matrix IR of the inverse secondary transform. T and supplies it to the matrix calculation unit 244.
[0169] Note that the inverse secondary transform selection unit 247 selects the matrix IR (=R T ) may be selected. Alternatively, the inverse transform IR may be selected according to the motion prediction information MVinfo and the secondary transform identifier st_idx instead of the intra prediction mode information IPinfo.
[0170] The matrix calculation unit 244 calculates a one-dimensional vector X for each sub-block (4×4 sub-block). 1d and the inverse secondary transformation matrix IR (=R T ) matrix, perform the matrix operation shown in the following equation (27), and as a result, obtain a one-dimensional vector Y 1d Here, the operator "·" represents an operation of performing an inner product (matrix multiplication) between matrices, and the operator "T" represents an operation of transposing a matrix. The matrix operation unit 244 derives the derived one-dimensional vector Y 1d is supplied to the scaling unit 245.
[0171] Y 1d T =IR·X 1d T =R T X 1d T ···(27)
[0172] The scaling unit 245 calculates the one-dimensional vector Y supplied from the matrix calculation unit 244 for each sub-block (4×4 sub-block). 1d In order to normalize the norm of , we perform a bit shift operation of N (N is a natural number) bits as shown in the following equation (28) on the one-dimensional vector Y 1d and then shift all elements of the bit-shifted one-dimensional vector Z 1d Ask for.
[0173] Z 1d =(Y 1d )>>N ···(28)
[0174] As shown in the following equation (29), before the N-bit shift operation, the value 1<<(N-1) is added to the one-dimensional vector Z 1d In equation (29), vector E is a one-dimensional vector in which all elements have a value of 1.
[0175] Z 1d =(Y 1d +((N-1)<<1)·E)>>N ···(29)
[0176] For example, the matrix IR of the inverse secondary transformation (=R T ) matrix is the transposed matrix of the matrix R of the secondary transformation shown in FIG. 10 and is an 8-bit scaled matrix, so the value of N used for normalizing the norm in the scaling unit 245 is 8. Generally, the matrix IR (=R T ) is scaled by N bits, the bit shift amount for norm normalization is N bits. The scaling unit 245 scales the one-dimensional vector Z after norm normalization obtained as above. 1d is supplied to the matrix generator 246.
[0177] The matrix generator 246 generates a norm-normalized one-dimensional vector Z 1d and the scan identifier scanIdx are input, and a one-dimensional vector Z 1d into a 4×4 matrix of primary transform coefficients Coeff_IS. The matrix generator 246 supplies the obtained primary transform coefficients Coeff_IS to the inverse primary transform unit 232.
[0178] As described above, the nonzero coefficient number determination unit 241 sets the secondary transform skip flag StSkipFlag based on the number of nonzero coefficients for each sub-block, and the switch 242 controls the skipping of the secondary transform based on the secondary transform skip flag StSkipFlag. In this way, it becomes possible to skip the secondary transform for sub-blocks with sparse nonzero coefficients, thereby suppressing a decrease in energy compaction and a decrease in coding efficiency.
[0179] <Flow of image decoding process> Next, a description will be given of the flow of each process executed by the above-described image decoding device 200. First, an example of the flow of image decoding process will be described with reference to the flowchart in FIG.
[0180] When the image decoding process starts, in step S201, the decoding unit 211 decodes the bitstream (encoded data) supplied to the image decoding device 200, and obtains information such as header information Hinfo, prediction mode information Pinfo, transformation information Tinfo, residual information Rinfo, and quantized transformation coefficient level level.
[0181] In step S202, the inverse quantization unit 212 inverse quantizes the quantized transform coefficient level obtained by the process of step S201 to derive the transform coefficient Coeff_IQ. This inverse quantization is the inverse process of the quantization performed in step S105 (FIG. 11) of the image encoding process, and is the same process as the inverse quantization performed in step S106 (FIG. 11) of the image encoding process.
[0182] In step S203, the inverse transform unit 213 inversely transforms the transform coefficient Coeff_IQ obtained by the process of step S202 to derive a prediction residual D'. This inverse transform is the inverse process of the transform process performed in step S104 (FIG. 11) of the image encoding process, and is the same process as the inverse transform performed in step S107 (FIG. 11) of the image encoding process.
[0183] In step S204, the prediction unit 216 performs prediction in the same prediction mode as that used in encoding, based on the prediction mode information PInfo, and generates a predicted image.
[0184] In step S205, the calculation unit 214 adds the prediction image obtained in the process of step S204 to the prediction residual D' obtained in the process of step S203 to obtain a decoded image.
[0185] When the process of step S205 ends, the image decoding process ends.
[0186] <Flow of the inverse conversion process> Next, an example of the flow of the inverse conversion process executed in step S203 in FIG. 15 will be described with reference to the flowchart in FIG.
[0187] When the inverse transform process starts, in step S221, the inverse secondary transform unit 231 (switch 242) determines whether the secondary transform identifier st_idx applies the inverse secondary transform (st_idx>0). If it is determined that the secondary transform identifier st_idx is 0 (the secondary transform identifier st_idx indicates that the inverse secondary transform is to be skipped), the inverse secondary transform (the processes of steps S222 to S230) is skipped, and the process proceeds to step S231. That is, the inverse secondary transform unit 231 (switch 242) supplies the secondary transform coefficients Coeff_IQ obtained by the process of step S202 in FIG. 15 to the inverse primary transform unit 232 as the primary transform coefficients Coeff_IS.
[0188] Also, if it is determined in step S221 that the secondary transformation identifier st_idx is greater than 0 (the secondary transformation identifier st_idx indicates that an inverse secondary transformation is to be performed), the process proceeds to step S222.
[0189] In step S222, the inverse secondary transform selection unit 247 selects the matrix IR of the inverse secondary transform specified by the secondary transform identifier st_idx.
[0190] In step S223, the inverse secondary transform unit 231 selects an unprocessed sub-block included in the transform block to be processed.
[0191] In step S224, the non-zero coefficient number determination unit 241 derives the number of non-zero coefficients numSigInSBK of the transform coefficients based on the sub-block-based secondary transform coefficients Coeff_IQ obtained by processing step S202 of Figure 15, as described above, and further derives the secondary transform skip flag StSkipFlag using the number of non-zero coefficients numSigInSBK of the transform coefficients and the threshold value TH.
[0192] In step S225, switch 242 determines whether the secondary transform skip flag StSkipFlag obtained by the process of step S224 indicates skipping of the inverse secondary transform. If it is determined that the secondary transform skip flag StSkipFlag indicates execution of the secondary transform, that is, if it is determined that the value of the secondary transform skip flag StSkipFlag is 0, the process proceeds to step S226.
[0193] In step S226, the rasterizing unit 243 converts the secondary transform coefficients Coeff_IQ obtained by the process of step S202 in FIG. 15 into a one-dimensional vector X 1d Convert to.
[0194] In step S227, the matrix calculation unit 244 calculates the one-dimensional vector X 1d and the matrix IR of the inverse secondary transformation obtained by the processing in step S222, and a one-dimensional vector Y 1d Ask for.
[0195] In step S228, the scaling unit 245 calculates the one-dimensional vector Y 1d Normalize the norm of the 1-dimensional vector Z 1d Ask for.
[0196] In step S229, the matrix generator 246 generates the one-dimensional vector Z 1d is transformed into a 4x4 matrix to determine the primary transform coefficient Coeff_IS of the sub-block to be processed. When the processing of step S229 ends, the process proceeds to step S230. Also, in step S225, if the inverse secondary transform is to be skipped, that is, if the value of the secondary transform skip flag StSkipFlag is 1, the process proceeds to step S230.
[0197] In step S230, the inverse secondary transform unit 231 determines whether all sub-blocks of the transform block being processed have been processed. If it is determined that unprocessed sub-blocks exist, the process returns to step S223, and the subsequent processes are repeated. That is, the processes (inverse secondary transform) of steps S223 to S230 are executed for each sub-block of the transform block being processed. If it is determined in step S230 that all sub-blocks have been processed (the inverse secondary transform of all sub-blocks has been executed or skipped), the process proceeds to step S231.
[0198] In step S231, the inverse primary transform unit 232 performs inverse primary transform on the primary transform coefficient Coeff_IS based on the primary transform identifier pt_idx to derive a prediction residual D′. This prediction residual D′ is supplied to the calculation unit 214.
[0199] When the process of step S231 ends, the inverse conversion process ends and the process returns to FIG.
[0200] Note that the above inverse transformation process may be performed by rearranging the processing order of each step or changing the processing content within a feasible range. For example, if it is determined in step S221 that the secondary transformation identifier st_idx is 0, a 16 × 16 identity matrix may be selected as the matrix IR of the inverse secondary transformation, and the processes of steps S222 to S230 may be performed.
[0201] By performing each process as described above, it is possible to control the skipping (execution) of the inverse secondary transform in units of sub-blocks. Therefore, it is possible to suppress a decrease in energy compaction for a residual signal with sparse non-zero coefficients. That is, it is possible to suppress a decrease in coding efficiency. In other words, it is possible to suppress an increase in the load of decoding (inverse secondary transform) while suppressing a decrease in coding efficiency.
[0202] In the above, we have explained how the skip of the (inverse) secondary transform is controlled for each sub-block, but the control of the skip for each sub-block can be applied to any transform process, not just the (inverse) secondary transform.
[0203] <2. Second embodiment> <Selecting a secondary transformation using a scanning method> In both the methods disclosed in Non-Patent Document 1 and Non-Patent Document 2, the secondary transform had matrices equal to the number of intra prediction mode classes and the number of secondary transforms corresponding to each class. Therefore, a huge memory size was required to store the secondary transform matrices. For example, in the method disclosed in Non-Patent Document 1, the number of intra prediction mode classes was 12, and the number of secondary transforms for each class was 3, resulting in 12*3 = 36 matrices. Furthermore, in the method disclosed in Non-Patent Document 2, the number of intra prediction mode classes was 35, and the number of secondary transforms for each class was 5, resulting in 35*5 = 175 matrices.
[0204] Therefore, for example, when each matrix element is held with 9-bit precision, the method described in Non-Patent Document 1 requires a memory size as shown in the following formula (30): Furthermore, the method described in Non-Patent Document 2 requires a memory size as shown in the following formula (31):
[0205] Memory size = 9bit * 16*16 * 36 = 829944 (bits) = 10368 (bytes) = 10.125 (KB) ···(30) Memory size = 9bit * 16*16 * 175 = 403200 (bits) = 50400 (bytes) = 49.21875 (KB) ···(31)
[0206] In this way, if the amount of data of the (inverse) secondary transform matrix to be stored increases, there is a risk that the load of encoding and decoding will increase, and the required memory size will also increase, which may result in an increase in costs.
[0207] Therefore, the matrix R of the secondary transform and the matrix IR of the inverse secondary transform are set based on the secondary transform identifier and the scan identifier. That is, focusing on the correspondence between the direction of the intra prediction mode and the scan method, the classification of the intra prediction mode is replaced with a scan identifier (scanIdx), which is information related to the scan method. Figure 17 shows an example of the correspondence relationship between the intra prediction mode and the scan identifier (scanIdx).
[0208] There are five types of secondary transforms corresponding to each value of the scan identifier (scanIdx). This is because they are assigned to each secondary transform identifier (st_idx). Therefore, in this case, the total number of secondary transforms is 3 x 5 = 15. In other words, the number of secondary transforms can be reduced compared to the methods described in Non-Patent Document 1 and Non-Patent Document 2 above. In the case of 9-bit precision, the memory size required to store all secondary transforms is given by the following equation (32).
[0209] Memory size = 9bit * 16 * 16 = 15 = 34560 (bits) = 4320 (bytes) = 4.21875 (KB) ···(32)
[0210] Therefore, compared to the above-mentioned formulas (30) and (31) (the methods described in Non-Patent Documents 1 and 2), the amount of data in the secondary transform matrix can be significantly reduced. This makes it possible to suppress an increase in the encoding / decoding load and an increase in the memory size required to store the (inverse) secondary transform matrix.
[0211] <Conversion section> In this case, the image coding device 100 also has a configuration basically similar to that of the first embodiment. However, in this case, the image coding device 100 includes a setting unit that sets a matrix for a transform process on transform coefficients based on the content of the transform process and a scanning method, a rasterization unit that converts transform coefficients obtained by transforming a prediction residual, which is the difference between an image and a predicted image of the image, into a one-dimensional vector, a matrix operation unit that performs a matrix operation on the one-dimensional vector using the matrix set by the setting unit, a scaling unit that scales the one-dimensional vector on which the matrix operation has been performed, and a matrix conversion unit that converts the scaled one-dimensional vector into a matrix. That is, the conversion unit 112 in this case sets a matrix for a transform process on transform coefficients based on the content of the transform process and a scanning method, converts transform coefficients obtained by transforming a prediction residual, which is the difference between an image and a predicted image of the image, into a one-dimensional vector, performs a matrix operation on the one-dimensional vector using the set matrix, scales the one-dimensional vector on which the matrix operation has been performed, and matrix converts the scaled one-dimensional vector.
[0212] Fig. 18 is a block diagram showing an example of the main configuration of the conversion unit 112 in this case. As shown in Fig. 18, the conversion unit 112 in this case also has basically the same configuration as in the first embodiment (Fig. 8). However, the secondary conversion unit 132 in this case can omit the quantization unit 146 to the switch 148, and also has a secondary conversion selection unit 301 instead of the secondary conversion selection unit 145.
[0213] The secondary transform selection unit 301 receives the secondary transform identifier st_idx and the scan identifier scanIdx as input. The secondary transform selection unit 301 selects a matrix R of the secondary transform based on the input secondary transform identifier st_idx and the scan identifier scanIdx, and supplies the matrix R to the matrix calculation unit 142.
[0214] <Secondary conversion selection section> 19 is a block diagram showing an example of the main configuration of the secondary transform selection unit 301. As shown in FIG.
[0215] The secondary transform derivation unit 311 receives the secondary transform identifier st_idx and the scan identifier scanIdx as input. Based on the input secondary transform identifier st_idx and scan identifier scanIdx, the secondary transform derivation unit 311 reads out the corresponding secondary transform matrix R from the secondary transform matrix table LIST_FwdST[][] stored in the secondary transform holding unit 312, as shown in the following equation (33), and outputs it to the outside. Here, the inverse secondary transform matrix table LIST_FwdST[][] stores the secondary transform matrix R corresponding to each scan identifier scanIdx and secondary transform identifier st_idx.
[0216] R = LIST_FwdST[ scanIdx ][ st_idx ] ···(33)
[0217] The secondary transformation holding unit 312 holds a secondary transformation matrix table LIST_FwdST[][] that stores a secondary transformation matrix R corresponding to each scan identifier scanIdx and secondary transformation identifier st_idx. Based on an instruction from the secondary transformation derivation unit 311, the secondary transformation holding unit 312 supplies the corresponding secondary transformation matrix R to the secondary transformation derivation unit 311.
[0218] <Conversion process flow> Next, an example of the flow of each process executed by the image coding device 100 will be described. In this case, the image coding device 100 performs the image coding process basically in the same way as in the first embodiment (FIG. 11). An example of the flow of the conversion process in this case will be described with reference to the flowchart in FIG.
[0219] When the transform process starts, the processes of steps S301 and S302 are executed in the same manner as the processes of steps S121 and S122 in Fig. 12. That is, when it is determined that the secondary transform identifier st_idx is 0 (indicating that the secondary transform is skipped), the secondary transform (the processes of steps S303 to S309) is skipped, the transform process ends, and the process returns to Fig. 11. That is, the secondary transform unit 132 supplies the primary transform coefficient Coeff_P to the quantization unit 113 as the transform coefficient Coeff.
[0220] Also, if it is determined in step S302 that the secondary transformation identifier st_idx is greater than 0 (indicating that secondary transformation has been performed), the process proceeds to step S303.
[0221] In step S303, the secondary transform selection unit 301 selects a matrix R of a secondary transform corresponding to the secondary transform identifier st_idx and the scan identifier scanIdx. That is, the secondary transform derivation unit 311 reads out and selects the matrix R of a secondary transform corresponding to the secondary transform identifier st_idx and the scan identifier scanIdx from the secondary transform matrix table held in the secondary transform holding unit 312.
[0222] The processes of steps S304 to S309 are executed in the same manner as the processes of steps S124 to S128 and step S134 in Fig. 12. That is, the processes of steps S304 to S309 are performed for each sub-block, thereby performing secondary transformation for each sub-block. Then, if it is determined in step S309 that all sub-blocks have been processed, the transformation process ends and the process returns to Fig. 11.
[0223] The order of the steps of the transformation process may be changed or the content of the process may be modified to the extent possible. For example, if it is determined in step S302 that the secondary transformation identifier st_idx=0, a 16×16 identity matrix may be selected as the matrix R of the secondary transformation, and the processes of steps S304 to S309 may be executed.
[0224] By performing each process as described above, the secondary transform matrix R can be selected based on the secondary transform identifier st_idx and the scan identifier scanIdx. Therefore, the amount of data for the secondary transform matrix can be significantly reduced. This prevents an increase in the encoding load and an increase in the memory size required to store the secondary transform matrix.
[0225] <Inverse conversion section> Next, the image decoding device 200 will be described. The image decoding device 200 in this case also has a configuration basically similar to that in the first embodiment. However, in this case, the image decoding device 200 includes: a setting unit that sets a matrix for inverse transform processing on transform coefficients based on the content of the inverse transform processing and a scanning method; a rasterization unit that converts the transform coefficients, which are obtained by inverse transform processing to obtain prediction residuals, which are the differences between an image and a predicted image of that image, into one-dimensional vectors; a matrix operation unit that performs a matrix operation on the one-dimensional vector using the matrix set by the setting unit; a scaling unit that scales the one-dimensional vector on which the matrix operation has been performed; and a matrix generation unit that transforms the scaled one-dimensional vector into a matrix. That is, the inverse transform unit 213 sets a matrix for inverse transform processing on transform coefficients based on the content of the inverse transform processing and a scanning method, transforms the transform coefficients, which are obtained by inverse transform processing to obtain prediction residuals, which are the differences between an image and a predicted image of that image, into one-dimensional vectors, performs a matrix operation on the one-dimensional vector using the set matrix, scales the one-dimensional vector on which the matrix operation has been performed, and matrix generates the scaled one-dimensional vector.
[0226] Fig. 21 is a block diagram showing an example of the main configuration of inverse transform unit 213 in this case. As shown in Fig. 21, in this case, inverse transform unit 213 basically has the same configuration as in the first embodiment (Fig. 14). However, inverse secondary transform unit 231 in this case can omit non-zero coefficient number determination unit 241 and switch 242, and also has an inverse secondary transform selection unit 321 instead of inverse secondary transform selection unit 247.
[0227] The inverse secondary transform selection unit 321 receives the secondary transform identifier st_idx and the scan identifier scanIdx as input. The inverse secondary transform selection unit 321 selects a matrix IR for the inverse secondary transform based on the input secondary transform identifier st_idx and scan identifier scanIdx, and supplies the matrix IR to the matrix calculation unit 244.
[0228] <Inverse secondary transformation selection section> 22 is a block diagram showing an example of the main configuration of the inverse secondary transform selection unit 321. As shown in FIG.
[0229] The inverse secondary transform derivation unit 331 receives the secondary transform identifier st_idx and the scan identifier scanIdx as input. Based on the input secondary transform identifier st_idx and the scan identifier scanIdx, the inverse secondary transform derivation unit 331 retrieves the corresponding inverse secondary transform matrix IR (=R T ) is read out as shown in the following equation (34) and output to the outside. Here, the inverse secondary transform matrix table LIST_InvST[][] stores the matrix IR of the inverse secondary transform corresponding to each scan identifier scanIdx and secondary transform identifier st_idx.
[0230] IR = LIST_InvST[ scanIdx ][ st_idx ] ···(34)
[0231] The inverse secondary transform holding unit 332 holds an inverse secondary transform matrix table LIST_InvST[][] in which the matrix IR of the inverse secondary transform corresponding to each scan identifier scanIdx and secondary transform identifier st_idx is stored. Based on an instruction from the inverse secondary transform derivation unit 331, the inverse secondary transform holding unit 332 holds the matrix IR of the corresponding inverse secondary transform (=R T ) is supplied to the inverse secondary transformation derivation unit 331.
[0232] <Flow of the inverse conversion process> Next, an example of the flow of each process executed by the image decoding device 200 will be described. In this case, the image decoding device 200 performs the image decoding process basically in the same way as in the first embodiment (FIG. 15). An example of the flow of the inverse transform process in this case will be described with reference to the flowchart in FIG.
[0233] When the inverse transform process starts, in step S321, the inverse secondary transform unit 231 determines whether the secondary transform identifier st_idx applies the inverse secondary transform (st_idx>0). If it is determined that the secondary transform identifier st_idx is 0 (the secondary transform identifier st_idx indicates that the inverse secondary transform is to be skipped), the inverse secondary transform (the processes of steps S322 to S328) is skipped, and the process proceeds to step S329. That is, the inverse secondary transform unit 231 supplies the secondary transform coefficients Coeff_IQ obtained by the process of step S202 in FIG. 15 to the inverse primary transform unit 232 as the primary transform coefficients Coeff_IS.
[0234] Also, if it is determined in step S321 that the secondary transformation identifier st_idx is greater than 0 (the secondary transformation identifier st_idx indicates that an inverse secondary transformation is to be performed), the process proceeds to step S322.
[0235] In step S322, the inverse secondary transform selection unit 321 selects the matrix IR of the inverse secondary transform that corresponds to the secondary transform identifier st_idx and the scan identifier scanIdx. That is, the inverse secondary transform derivation unit 331 reads out and selects the matrix IR of the inverse secondary transform that corresponds to the secondary transform identifier st_idx and the scan identifier scanIdx from the inverse secondary transform matrix table held in the inverse secondary transform holding unit 332.
[0236] In step S323, the inverse secondary transform unit 231 selects an unprocessed sub-block included in the transform block to be processed.
[0237] The processes of steps S324 to S328 are executed in the same manner as the processes of steps S226 to S230 in Fig. 16. That is, the processes of steps S323 to S328 are performed for each sub-block, thereby performing secondary transformation for each sub-block. Then, if it is determined in step S328 that all sub-blocks have been processed, the process proceeds to step S329.
[0238] In step S329, the inverse primary transform unit 232 performs inverse primary transform on the primary transform coefficient Coeff_IS based on the primary transform identifier pt_idx to derive a prediction residual D′. This prediction residual D′ is supplied to the calculation unit 214.
[0239] When the process of step S231 ends, the inverse conversion process ends and the process returns to FIG.
[0240] Note that the above inverse transformation process may be performed by rearranging the order of the steps or changing the content of the processes within a feasible range. For example, if it is determined in step S321 that the secondary transformation identifier st_idx is 0, a 16 × 16 identity matrix may be selected as the matrix IR for the inverse secondary transformation, and the processes of steps S323 to S328 may be performed.
[0241] By performing each process as described above, the matrix IR of the inverse secondary transform can be selected based on the secondary transform identifier st_idx and the scan identifier scanIdx. Therefore, the amount of data for the matrix of the inverse secondary transform can be significantly reduced. This prevents an increase in the encoding and decoding load and prevents an increase in the memory size required to store the matrix of the (inverse) secondary transform.
[0242] <3. Third Embodiment> <Bandwidth Limit> Non-Patent Document 1 describes that when the transform block size is 64x64, after performing an orthogonal transform of 1 or more on the prediction residual, band limitation is performed to forcibly set high frequency components other than the low frequency components of the upper left 32x32 to 0, thereby reducing the computational complexity and implementation cost of the decoder.
[0243] 24 is a block diagram showing an example of the main configuration of a conversion unit that performs primary conversion, secondary conversion, and band limitation on prediction residuals D in the method described in Non-Patent Document 1. As shown in FIG. 24, the conversion unit 400 includes a switch 401, a primary conversion unit 402, a secondary conversion unit 403, and a band limitation unit 404.
[0244] The switch 401 receives as input a transform skip flag ts_flag, a transform quantization bypass flag transquant_bypass_flag, and a prediction residual D. The transform skip flag ts_flag is information indicating whether or not to skip the (inverse) primary transform and the (inverse) secondary transform for a target data unit. For example, when the transform skip flag ts_flag is 1 (true), the (inverse) primary transform and the (inverse) secondary transform are skipped. On the other hand, when the transform skip flag ts_flag is 0 (false), the (inverse) primary transform and the (inverse) secondary transform are executed.
[0245] Furthermore, the transform quantization bypass flag transquant_bypass_flag is information indicating whether or not to skip (bypass) the (inverse) primary transform, the (inverse) secondary transform, and the (inverse) quantization in the target data unit. For example, when the transform quantization bypass flag transquant_bypass_flag is 1 (true), the (inverse) primary transform, the (inverse) secondary transform, and the (inverse) quantization are bypassed. When the transform quantization bypass flag transquant_bypass_flag is 0 (false), the (inverse) primary transform, the (inverse) secondary transform, and the (inverse) quantization are not bypassed.
[0246] The switch 401 controls skipping of the primary transform and secondary transform for the prediction residual D based on the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag.
[0247] Specifically, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 0, the switch 401 supplies the prediction residual D to the primary transform unit 402, thereby causing the primary transform and secondary transform to be executed. On the other hand, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 401 supplies the prediction residual D to the band limiting unit 404 as a secondary transform coefficient Coeff, thereby causing the primary transform and secondary transform to be skipped.
[0248] Similar to the primary transform unit 131 in FIG. 8, the primary transform unit 402 performs a primary transform on the prediction residuals D based on the primary transform identifier pt_idx to derive primary transform coefficients Coeff_P. The primary transform identifier pt_idx is an identifier indicating which (inverse) primary transform to apply to the vertical and horizontal (inverse) primary transforms for the target data unit (see, for example, JVET-B1001, 2.5.1 Adaptive multiple Core transform; in JEM2, it is also referred to as emt_idx). The primary transform unit 402 supplies the primary transform coefficients Coeff_P to the secondary transform unit 403.
[0249] The secondary transform unit 403 performs secondary transform on the primary transform coefficients Coeff_P supplied from the primary transform unit 402 based on the secondary transform identifier st_idx using the method described in Non-Patent Document 1, to derive secondary transform coefficients Coeff. The secondary transform unit 403 supplies the secondary transform coefficients Coeff to the band limiting unit 404.
[0250] When the block size TBSize of the transform block to be processed is 64×64, the band limiting unit 404 performs band limiting on the secondary transform coefficients Coeff supplied from the switch 401 or the secondary transform unit 403 so as to set high frequency components to zero, and derives secondary transform coefficients Coeff' after band limiting. On the other hand, when the block size TBSize of the transform block to be processed is not 64×64, the band limiting unit 404 uses the secondary transform coefficients Coeff as they are as the secondary transform coefficients Coeff' after band limiting. The band limiting unit 404 outputs the secondary transform coefficients Coeff' after band limiting.
[0251] As described above, in the transform unit 400 of Fig. 24, even when the primary transform and secondary transform are skipped, band limitation (low-pass filter processing) is performed when the block size TBSize is 64 × 64. However, when the primary transform and secondary transform are skipped, the secondary transform coefficient Coeff input to the band limiting unit 404 is the prediction residual D. Therefore, the band limiting unit 404 performs band limitation on the prediction residual D, and the band limiting increases distortion.
[0252] For example, if the prediction residual D of a 64×64 transform block is a 64×64 image shown in A of FIG. 25, when the primary transform and secondary transform are skipped, the 64×64 image shown in A of FIG. 25 is input to the band-limiting unit 404 as the secondary transform coefficients Coeff. Therefore, when the band-limiting unit 404 performs band-limiting so as to set all transform coefficients other than the 32×32 transform coefficient in the upper left corner of the transform block as high-frequency components to 0, the secondary transform coefficients Coeff′ after band-limiting become a 64×64 image shown in B of FIG. 25. Therefore, distortion increases in the secondary transform coefficients Coeff′ after band-limiting. As a result, coding efficiency decreases. Furthermore, despite the transform quantization bypass being performed for the purpose of lossless coding, lossless coding cannot be performed.
[0253] In FIG. 25, the white area is an area where the pixel value is 0, and the black area is an area where the pixel value is 255.
[0254] <Prohibit bandwidth limiting when using transform skip or transform quantization bypass> In the third embodiment, the band limiting for the secondary transform coefficient Coeff is controlled based on the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag. Also, the band limiting for the secondary transform coefficient Coeff_IQ obtained by dequantization is controlled based on the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag. Specifically, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the (inverse) primary transform, the (inverse) secondary transform, and the band limiting are skipped.
[0255] By doing so, application of band limitation to the secondary transform coefficient Coeff (Coeff_IQ) which is the prediction residual D, i.e., the secondary transform coefficient in the pixel domain, is prohibited, thereby suppressing an increase in distortion. As a result, it is possible to suppress a decrease in coding efficiency. Furthermore, by performing transform quantization bypass, it is possible to perform lossless coding.
[0256] <Image encoding device> The configuration of the third embodiment of the image coding device as an image processing device to which the present technology is applied is the same as the configuration in Fig. 7 except for the configurations of the transformation information Tinfo, the transformation unit, and the inverse transformation unit. Therefore, hereinafter, only the transformation information Tinfo, the transformation unit, and the inverse transformation unit will be described.
[0257] In the third embodiment, the transform information Tinfo includes a transform skip flag ts_flag, a transform quantization bypass flag transquant_bypass_flag, a primary transform identifier pt_idx, a secondary transform identifier st_idx, and a block size TBSize.
[0258] Furthermore, the inverse transformation performed by the inverse transformation unit in the third embodiment is the inverse process of the transformation performed by the transformation unit, and is the same process as the inverse transformation performed in the image decoding device described later. Therefore, this inverse transformation will be described later in the description of the image decoding device.
[0259] <Conversion section> FIG. 26 is a block diagram showing an example of the main configuration of a conversion unit in the third embodiment of an image coding device as an image processing device to which the present technology is applied.
[0260] Of the components shown in Fig. 26, the same components as those in Fig. 24 are denoted by the same reference numerals, and duplicated explanations will be omitted where appropriate.
[0261] The configuration of the conversion unit 420 in FIG. 26 differs from the configuration of the conversion unit 400 in FIG. 24 in that a switch 421 and a band-pass filter 424 are provided instead of the switch 401 and the band-pass filter 404.
[0262] The switch 421 of the transform unit 420 receives the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag from the control unit 101, and receives the prediction residual D from the calculation unit 111. The switch 421 controls the primary transform, secondary transform, and skip of band limitation for the prediction residual D, based on the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag.
[0263] Specifically, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 0, the switch 421 supplies the prediction residual D to the primary transform unit 402, thereby causing the primary transform, secondary transform, and band limitation to be performed. On the other hand, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 421 (control unit) supplies the prediction residual D to the quantization unit 113 as a band-limited secondary transform coefficient Coeff', thereby skipping the primary transform, secondary transform, and band limitation.
[0264] The band limiting unit 424 performs band limiting on the secondary transform coefficients Coeff output from the secondary transform unit 403 based on the block size TBSize supplied from the control unit 101 so as to set high frequency components to zero, and derives secondary transform coefficients Coeff' after band limiting. The band limiting unit 424 supplies the secondary transform coefficients Coeff' after band limiting to the quantization unit 113.
[0265] Specifically, when the block size TBSize is smaller than a predetermined block size TH_TBSize, the band limiting unit 424 does not perform band limiting on the secondary transform coefficients Coeff, and supplies the secondary transform coefficients Coeff' after band limiting to the quantization unit 113.
[0266] On the other hand, when the block size TBSize is equal to or larger than a predetermined block size TH_TBSize, the band limiting unit 424 derives the secondary transform coefficient Coeff(i,j)' after band limitation for each coordinate (i,j) in the transform block using the band limiting filter H(i,j) and the secondary transform coefficient Coeff(i,j) according to the following equation (35).
[0267]
number
[0268] The band-limiting filter H(i,j) is defined by the following equation (36).
[0269]
number
[0270] Furthermore, TBXSize and TBYSize are the horizontal and vertical sizes of the transform block to be processed, respectively. According to equations (35) and (36), the high-frequency components of the secondary transform coefficient Coeff(i,j)' are 0. The band limiting unit 424 supplies the secondary transform coefficient Coeff' after band limiting derived as described above to the quantization unit 113.
[0271] <Image encoding process flow> The image coding process executed by the third embodiment of the image coding device as an image processing device to which the present technology is applied is the same as the image coding process in Fig. 11 except for the transform process in step S104 and the inverse transform process in S107. The inverse transform process is the inverse process of the transform process, and is executed in the same manner as the inverse transform process executed in the image decoding process described later, so only the transform process will be described here.
[0272] <Conversion process flow> FIG. 27 is a flowchart illustrating an example of the flow of conversion processing executed by the third embodiment of the image coding device serving as the image processing device to which the present technology is applied.
[0273] When the conversion process starts, in step S401, the switch 421 determines whether the conversion skip flag ts_flag or the conversion and quantization bypass flag transquant_bypass_flag supplied from the control unit 101 is 1 or not.
[0274] If it is determined in step S401 that the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 421 determines to skip the primary transform, secondary transform, and band limitation. Then, the switch 421 does not perform the processes of steps S402 to S406, but supplies the prediction residual D supplied from the calculation unit 111 to the quantization unit 113 as the band-limited secondary transform coefficient Coeff', and ends the process.
[0275] On the other hand, if it is determined in step S401 that the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0, the switch 421 determines to perform primary transform, secondary transform, and band limitation. Then, the switch 421 supplies the prediction residual D to the primary transform unit 402, and the process proceeds to S402.
[0276] In step S402, the primary transform unit 402 performs primary transform on the prediction residuals D supplied from the switch 421, based on the primary transform identifier pt_idx supplied from the control unit 101, to derive primary transform coefficients Coeff_P. The primary transform unit 402 supplies the primary transform coefficients Coeff_P to the secondary transform unit 403.
[0277] In step S403, the secondary conversion unit 403 determines whether the secondary conversion identifier st_idx is greater than 0. If it is determined in step S403 that the secondary conversion identifier st_idx is greater than 0, that is, if the secondary conversion identifier st_idx indicates that a secondary conversion is to be performed, the process proceeds to step S404.
[0278] In step S404, the secondary transform unit 403 performs a secondary transform corresponding to the secondary transform identifier st_idx on the primary transform coefficient Coeff_P to derive the secondary transform coefficient Coeff. The secondary transform unit 403 supplies the secondary transform coefficient Coeff to the band limiting unit 424, and the process proceeds to step S405.
[0279] On the other hand, if it is determined in step S403 that the secondary transformation identifier st_idx is not greater than 0, that is, if the secondary transformation identifier st_idx indicates skipping of the secondary transformation, the secondary transformation unit 403 skips the processing of step S404. Then, the secondary transformation unit 403 supplies the primary transformation coefficient Coeff_P obtained by the processing of step S402 to the band limiting unit 424 as the secondary transformation coefficient Coeff, and the processing proceeds to step S405.
[0280] In step S405, the bandwidth limiting unit 424 determines whether the block size TBSize of the transform block to be processed is equal to or larger than the predetermined block size TH_TBSize. If it is determined in step S405 that the block size TBSize is equal to or larger than the predetermined block size TH_TBSize, the process proceeds to step S406.
[0281] In step S406, the band-limiting unit 424 applies the band-limiting filter H defined by the above-mentioned equation (36) to the secondary transform coefficient Coeff according to the above-mentioned equation (35) to perform band-limiting, and derives the band-limited secondary transform coefficient Coeff'. Then, the band-limiting unit 424 supplies the band-limited secondary transform coefficient Coeff' to the quantization unit 113, and the process ends.
[0282] On the other hand, if it is determined in step S405 that the block size TBSize is smaller than the predetermined block size TH_TBSize, the band-limiting unit 424 skips the process of step S406. Then, the band-limiting unit 424 supplies the secondary transform coefficients Coeff as they are to the quantization unit 113 as the secondary transform coefficients Coeff' after band-limiting, and ends the process.
[0283] The order of the steps of the transformation process may be changed or the content of the process may be modified to the extent possible. For example, if it is determined in step S403 that the secondary transformation identifier st_idx is not greater than 0, a unit matrix may be selected as the matrix R of the secondary transformation, and the process of step S404 may be performed.
[0284] As described above, when transform skip or transform quantization bypass is not performed, the switch 421 supplies the prediction residual D to the band limiting unit 424 via the primary transform unit 402 and the secondary transform unit 403. As a result, when the block size TBSize is equal to or larger than a predetermined block size TH_TBSize, the band limiting unit 424 performs band limiting so as to set the high-frequency components of the secondary transform coefficients Coeff to 0. Therefore, in this case, the image coding device needs to code only the low-frequency components of the secondary transform coefficients Coeff, and the coding process of the secondary transform coefficients Coeff can be reduced.
[0285] Furthermore, when transform skip or transform quantization bypass is performed, the switch 421 does not supply the prediction residual D to the band limiting unit 424. As a result, the band limiting unit 424 does not perform band limiting on the secondary transform coefficient Coeff, which is the prediction residual D, even if the block size TBSize is equal to or larger than the predetermined block size TH_TBSize. In other words, when transform skip or transform quantization bypass is performed, band limiting is also skipped. Therefore, even when transform skip or transform quantization bypass is performed, it is possible to prevent an increase in distortion when transform skip or transform quantization bypass is performed and the block size TBSize is equal to or larger than the predetermined block size TH_TBSize, compared to when band limiting is performed. As a result, a decrease in coding efficiency is suppressed, that is, coding efficiency can be improved. Furthermore, by performing transform quantization bypass, lossless coding can be performed.
[0286] <Image decoding device> Next, decoding of the coded data coded as above will be described. The configuration of the third embodiment of the image decoding device as an image processing device to which the present technology is applied is the same as the configuration in Fig. 13 except for the configurations of the transformation information Tinfo and the inverse transformation unit. Since the transformation information Tinfo has been described above, only the inverse transformation unit will be described below.
[0287] <Inverse conversion section> FIG. 28 is a block diagram showing an example of the main configuration of an inverse transform unit in the third embodiment of an image decoding device as an image processing device to which the present technology is applied.
[0288] As shown in FIG. 28, the inverse conversion unit 440 includes a switch 441, a band limiting unit 442, an inverse secondary conversion unit 443, and an inverse primary conversion unit 444.
[0289] The switch 441 is supplied with the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag from the decoding unit 211, and is supplied with the secondary transform coefficient Coeff_IQ from the inverse quantization unit 212. The switch 441 controls the band limiting of the secondary transform coefficient Coeff_IQ, the inverse secondary transform, and the skipping of the inverse primary transform, based on the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag.
[0290] Specifically, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 0, the switch 441 executes band limiting, inverse secondary transform, and inverse primary transform by supplying the secondary transform coefficient Coeff_IQ to the band limiting unit 442. On the other hand, when the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 441 (control unit) supplies the secondary transform coefficient Coeff_IQ to the calculation unit 214 as a prediction residual D', thereby skipping band limiting, inverse secondary transform, and inverse primary transform.
[0291] The band limiting unit 442 performs band limiting on the secondary transform coefficients Coeff_IQ supplied from the switch 441 based on the block size TBSize of the transform block to be processed supplied from the decoding unit 211 so as to set high frequency components to zero, and derives secondary transform coefficients Coeff_IQ' after band limiting. The band limiting unit 442 supplies the secondary transform coefficients Coeff_IQ' after band limiting to the inverse secondary transform unit 443.
[0292] Specifically, when the block size TBSize is less than a predetermined block size TH_TBSize, the band limiting unit 442 does not perform band limiting on the secondary transform coefficient Coeff_IQ, and supplies the secondary transform coefficient Coeff_IQ to the inverse secondary transform unit 443 as the band-limited secondary transform coefficient Coeff_IQ'.
[0293] On the other hand, when the block size TBSize is equal to or larger than the predetermined block size TH_TBSize, the band limiting unit 442 applies the band limiting filter H shown in the above-mentioned equation (35) to the secondary transform coefficient Coeff_IQ to derive the secondary transform coefficient Coeff_IQ(i,j)'.
[0294] That is, for each coordinate (i, j) in the transform block, the band-limiting unit 442 derives the secondary transform coefficient Coeff_IQ(i, j)' after band-limiting using the band-limiting filter H(i, j) and the secondary transform coefficient Coeff_IQ(i, j) according to the following equation (37).
[0295]
number
[0296] According to equation (37), the high frequency components of the secondary transform coefficient Coeff_IQ(i,j)′ are 0. The band limiting unit 442 supplies the band-limited secondary transform coefficient Coeff_IQ′ derived as described above to the inverse secondary transform unit 443.
[0297] The inverse secondary transform unit 443 performs an inverse secondary transform on the band-limited secondary transform coefficients Coeff_IQ′ supplied from the band limiting unit 442, based on the secondary transform identifier st_idx supplied from the decoding unit 211, to derive primary transform coefficients Coeff_IS. The inverse secondary transform unit 443 supplies the primary transform coefficients Coeff_IS to the inverse primary transform unit 444.
[0298] The inverse primary transform unit 444 performs inverse primary transform on the primary transform coefficients Coeff_IS supplied from the inverse secondary transform unit 443 based on the primary transform identifier pt_idx, in the same manner as the inverse primary transform unit 232 in Fig. 14 , to derive prediction residuals D'. The inverse primary transform unit 444 supplies the derived prediction residuals D' to the calculation unit 214.
[0299] <Flow of image decoding process> The image decoding process performed by the third embodiment of the image decoding device as an image processing device to which the present technology is applied is the same as the image decoding process of Figure 15 except for the inverse transform process of step S203, so below we will only describe the inverse transform process.
[0300] <Flow of the inverse conversion process> FIG. 29 is a flowchart illustrating an example of the flow of inverse transform processing executed by the third embodiment of the image decoding device as the image processing device to which the present technology is applied.
[0301] When the inverse transform process starts, in step S421, the switch 441 determines whether the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag supplied from the decoding unit 211 is 1 or not.
[0302] If it is determined in step S421 that the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 441 determines to skip the band limiting, the inverse secondary transform, and the inverse primary transform. Then, the switch 441 does not perform the processes of steps S422 to S426, but supplies the secondary transform coefficients Coeff_IQ supplied from the inverse quantization unit 212 to the calculation unit 214 as the prediction residuals D', and ends the process.
[0303] On the other hand, if it is determined in step S421 that the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0, the switch 441 determines to perform band limitation, inverse secondary transform, and inverse primary transform. Then, the switch 441 supplies the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212 to the band limitation unit 442, and the process proceeds to step S422.
[0304] In step S422, the bandwidth limiting unit 442 determines whether the block size TBSize of the transform block to be processed is equal to or larger than the predetermined block size TH_TBSize. If it is determined in step S422 that the block size TBSize is equal to or larger than the predetermined block size TH_TBSize, the process proceeds to step S423.
[0305] In step S423, the band limiting unit 442 applies a band limiting filter H to the secondary transform coefficient Coeff_IQ to limit the band, and derives the band-limited secondary transform coefficient Coeff_IQ'. Then, the band limiting unit 424 supplies the band-limited secondary transform coefficient Coeff_IQ' to the inverse secondary transform unit 443, and the process proceeds to step S424.
[0306] On the other hand, if it is determined in step S422 that the block size TBSize is smaller than the predetermined block size TH_TBSize, the band-limiting unit 442 skips the process of step S423. Then, the band-limiting unit 442 supplies the secondary transform coefficients Coeff_IQ as they are to the inverse secondary transform unit 443 as secondary transform coefficients Coeff_IQ′ after band-limiting, and proceeds to the process of step S424.
[0307] In step S424, the inverse secondary transformation unit 443 determines whether the secondary transformation identifier st_idx is greater than 0. If it is determined in step S424 that the secondary transformation identifier st_idx is greater than 0, that is, if the secondary transformation identifier st_idx indicates that an inverse secondary transformation is to be performed, the process proceeds to step S425.
[0308] In step S425, the inverse secondary transform unit 443 performs an inverse secondary transform corresponding to the secondary transform identifier st_idx on the band-limited secondary transform coefficient Coeff_IQ′ supplied from the band limiting unit 442, to derive the primary transform coefficient Coeff_P. The inverse secondary transform unit 443 supplies the primary transform coefficient Coeff_P to the inverse primary transform unit 444, and the process proceeds to step S426.
[0309] On the other hand, if it is determined in step S424 that the secondary transform identifier st_idx is not greater than 0, that is, if the secondary transform identifier st_idx indicates that the inverse secondary transform should be skipped, the inverse secondary transform unit 443 skips the processing of step S425. Then, the inverse secondary transform unit 443 supplies the band-limited secondary transform coefficients Coeff_IQ' to the inverse primary transform unit 444 as the primary transform coefficients Coeff_IS, and the processing proceeds to step S426.
[0310] In step S426, the inverse primary transform unit 444 performs inverse primary transform on the primary transform coefficients Coeff_IS supplied from the inverse secondary transform unit 443, derives prediction residuals D', and supplies them to the calculation unit 214. Then, the processing ends.
[0311] Note that the above inverse transformation process may be performed by rearranging the processing order of each step or changing the content of the process to the extent possible. For example, if it is determined in step S424 that the secondary transformation identifier st_idx is not greater than 0, the identity matrix may be selected as the matrix IR of the inverse secondary transformation, and the process of step S425 may be performed.
[0312] As described above, when an inverse transform skip or a skip of inverse quantization and inverse transform (hereinafter referred to as an inverse quantization / inverse transform bypass) is not performed, the switch 441 supplies the secondary transform coefficient Coeff_IQ to the band limiting unit 442. As a result, when the block size TBSize is equal to or larger than a predetermined block size TH_TBSize, the band limiting unit 442 performs band limiting so as to set the high frequency components of the secondary transform coefficient Coeff_IQ to 0. Therefore, in this case, the image decoding device needs to inverse transform only the low frequency components of the secondary transform coefficient Coeff_IQ, and the inverse transform processing of the secondary transform coefficient Coeff_IQ can be reduced.
[0313] Furthermore, when transform skip or inverse quantization / inverse transform bypass is performed, the switch 441 does not supply the secondary transform coefficient Coeff_IQ to the band limiting unit 442. As a result, the band limiting unit 442 does not perform band limiting on the secondary transform coefficient Coeff_IQ, which is a prediction residual, even if the block size TBSize is equal to or larger than the predetermined block size TH_TBSize. That is, when inverse transform skip or inverse quantization / inverse transform bypass is performed, band limiting is also skipped. Therefore, when transform skip or transform quantization bypass is performed, by also skipping band limiting, coded data can be decoded with reduced distortion. As a result, coded data with improved coding efficiency can be decoded. Furthermore, coded data losslessly coded by transform quantization bypass can be losslessly decoded.
[0314] In the above description, bandwidth limitation was performed in the third embodiment of the image decoding apparatus. However, when a bitstream constraint of "when the block size TBSize is greater than or equal to a predetermined block size TH_TBSize, non-zero coefficients are restricted within the upper left low-frequency region ((TH_TBSize>>1)×(TH_TBSize>>1))" is provided, bandwidth limitation may not be performed. This bitstream constraint can also be expressed as "when the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0, and the block size TBSize is the predetermined block size TH_TBSize, the values of the last non-zero coefficient X coordinate (last_sig_coeff_x_pos) and the last non-zero coefficient Y coordinate (last_sig_coeff_y_pos) indicating the position of the last non-zero coefficient in the transform block are within the range from 0 to (TH_TBSize / 2)-1".
[0315] In addition to 64×64, the predetermined block size TH_TBSize can be set to any value.
[0316] Furthermore, the method of the secondary transform in the third embodiment is not limited to the method described in Non-Patent Document 1. For example, the methods in the first and second embodiments may also be used. Also, the secondary transform coefficients may be clipped and restricted within a predetermined range.
[0317] <4. Fourth Embodiment> <Shapes of CU, PU, and TU> FIG. 30 is a diagram for explaining the shapes of CU, PU, and TU in the fourth embodiment.
[0318] The CU, PU, and TU in the fourth embodiment are CU, PU, and TU of QTBT (Quad tree plus binary tree) described in JVET-C0024, "EE2.1: Quadtree plus binary tree structure integration with JEM tools."
[0319] Specifically, in the block division of a CU in the fourth embodiment, one block can be divided not only into four (=2x2) sub-blocks but also into two (=1x2, 2x1) sub-blocks. That is, in the fourth embodiment, the block division of a CU is performed by recursively repeating the division of one block into four or two sub-blocks, resulting in the formation of a quad-tree or a horizontal or vertical binary-tree tree structure.
[0320] As a result, the shape of the CU can be rectangular as well as square. For example, if the LCU size is 128x128, the CU size (horizontal size w × vertical size h) can be any square size such as 128x128, 64x64, 32x32, 16x16, 8x8, or 4x4, as shown in Figure 30. The following sizes are also available: 128x64, 128x32, 128x16, 128x8, 128x4, 64x128, 32x128, 16x128, 8x128, 4x128, 64x32, 64x16, 64x8, 64x4, 32x64, 16x64, 8x64, 4x64, 32x16, 32x8, 32x4, 16x32, 8x32, 4x32, 16x8, 16x4, 8x16, 4x16, It may have a rectangular size of 8x4, 4x8, etc. In the fourth embodiment, PU and TU are the same as CU.
[0321] <Image encoding device> The configuration of the fourth embodiment of the image coding device as an image processing device to which the present technology is applied is the same as the configuration of the third embodiment, except for the configurations of the transformation information Tinfo, the transformation unit, and the inverse transformation unit. Therefore, hereinafter, only the transformation information Tinfo, the transformation unit, and the inverse transformation unit will be described.
[0322] The conversion information Tinfo in the fourth embodiment is identical to the conversion information Tinfo in the third embodiment, except that instead of the block size TBSize, it includes the width size (horizontal size) TBXSize and the height size (vertical size) TBYSize of the conversion block.
[0323] Furthermore, the inverse transformation performed by the inverse transformation unit in the fourth embodiment is the inverse process of the transformation performed by the transformation unit, and is the same process as the inverse transformation performed in the image decoding device described later. Therefore, this inverse transformation will be described later in the description of the image decoding device.
[0324] <Conversion section> FIG. 31 is a block diagram showing an example of the main configuration of a conversion unit in a fourth embodiment of an image coding device as an image processing device to which the present technology is applied.
[0325] Of the components shown in Fig. 31, the same components as those in Fig. 26 are denoted by the same reference numerals, and duplicated explanations will be omitted where appropriate.
[0326] The configuration of the conversion section 460 in FIG. 31 differs from the configuration of the conversion section 420 in FIG. 26 in that a band limiting section 464 is provided instead of the band limiting section 404.
[0327] The band limitation unit 464 of the conversion unit 460 performs band limitation on the secondary conversion coefficient Coeff output from the secondary conversion unit 403 so that the high-frequency components are set to 0 based on the horizontal size TBXSize and the vertical size TBYSize supplied from the control unit 101, and derives the secondary conversion coefficient Coeff' after band limitation. The band limitation unit 464 supplies the secondary conversion coefficient Coeff' after band limitation to the quantization unit 113.
[0328] Specifically, when the larger of the horizontal size TBXSize and the vertical size TBYSize, max(TBXSize, TBYSize), is less than a predetermined size THSize (max(TBXSize, TBYSize) < THSize), the band limitation unit 464 does not perform band limitation on the secondary conversion coefficient Coeff, and supplies the secondary conversion coefficient Coeff to the quantization unit 113 as the secondary conversion coefficient Coeff' after band limitation.
[0329] On the other hand, when max(TBXSize, TBYSize) is greater than or equal to a predetermined size THSize (max(TBXSize, TBYSize) >= THSize), the band limitation unit 464 derives the secondary conversion coefficient Coeff(i, j)' after band limitation for each pixel unit coordinate (i, j) in the conversion block according to the above-described formula (35). That is, when the logical value of the conditional expression max(TBXSize, TBYSize) >= THSize is 1 (true), the band limitation unit 464 performs band limitation, and when it is 0 (false), the band limitation is not performed.
[0330] Note that the band limitation filter H(i, j) is defined by the following formula (38).
[0331]
Equation
[0332] TH is a threshold value. According to formulas (35) and (38), 2, which is the processing unit in the conversion block in the secondary conversion including the coordinate (i, j) Nx2 M If the sub-block i is a sub-block other than the square area consisting of TH × TH sub-blocks in the upper left corner, the secondary transform coefficient Coeff(i,j)' becomes 0. That is, the high frequency component of the secondary transform coefficient Coeff(i,j)' becomes 0. The band limiting unit 464 supplies the secondary transform coefficient Coeff' after band limiting derived as described above to the quantization unit 113.
[0333] <Example of a band-limiting filter> FIG. 32 is a diagram illustrating an example of the band-limiting filter H(i,j).
[0334] In Fig. 32, the solid line rectangles indicate transform blocks, and the dotted line rectangles indicate sub-blocks. This also applies to Figs. 36 to 38, which will be described later. In the example of Fig. 32, the size of the sub-blocks is 2 2 x2 2 where the size THSize is the size of 8 sub-blocks, and the threshold TH in the band-limiting filter H(i,j) is 4. Also, A of Fig. 32 shows an example where the transform block size TBXSize×TBYSize is a size consisting of 8×8 sub-blocks (a transform block size of 32×32), B of Fig. 32 shows an example where the transform block size TBXSize×TBYSize is a size consisting of 8×4 sub-blocks (a transform block size of 32×16), and C of Fig. 32 shows an example where the transform block size TBXSize×TBYSize is a size consisting of 4×8 sub-blocks (a transform block size of 16×32).
[0335] In this case, as shown in A of Fig. 32, when the horizontal size TBXSize and the vertical size TBYSize are both 8 sub-blocks, a band-limiting filter H(i, j) is set. In this case, the band-limiting filter H(i, j) at the coordinates (i, j) included in the 4x4 sub-block group 471A marked with diagonal lines in the upper left of the drawing is 1, and the band-limiting filter H(i, j) at the coordinates (i, j) in sub-blocks other than the sub-block group 471A is 0.
[0336] 32B, when the horizontal size TBXSize is a size of eight sub-blocks and the vertical size TBYSize is a size of four sub-blocks smaller than the horizontal size TBXSize, a band-limiting filter H(i,j) is also set. In this case, the band-limiting filter H(i,j) at the coordinates (i,j) included in the 4×4 sub-block group 471B marked with diagonal lines in the diagram on the left is 1, and the band-limiting filter H(i,j) at the coordinates (i,j) in sub-blocks other than the sub-block group 471B is 0.
[0337] Furthermore, as shown in C of Fig. 32, when the horizontal size TBXSize is the size of four sub-blocks and the vertical size TBYSize is the size of eight sub-blocks larger than the horizontal size TBXSize, a band-limiting filter H(i, j) is also set. In this case, the band-limiting filter H(i, j) at the coordinates (i, j) included in the 4x4 sub-block group 471C marked with diagonal lines in the upper diagram is 1, and the band-limiting filter H(i, j) at the coordinates (i, j) in sub-blocks other than the sub-block group 471C is 0.
[0338] As described above, when max(TBXSize, TBYSize) is equal to or greater than the size THSize, the band-limiting filter H(i,j) is set to 0 or 1 depending on the position of the sub-block containing the coordinates (i,j).
[0339] <Image encoding process flow> The image coding process executed by the fourth embodiment of the image coding device as an image processing device to which the present technology is applied is the same as the image coding process in the third embodiment, except for the transform process and the inverse transform process. The inverse transform process is the inverse process of the transform process, and is executed in the same way as the inverse transform process executed in the image decoding process described later, so only the transform process will be described here.
[0340] <Conversion process flow> FIG. 33 is a flowchart illustrating an example of the flow of conversion processing executed by the fourth embodiment of the image coding device serving as the image processing device to which the present technology is applied.
[0341] The processing in steps S461 to S464 is the same as the processing in steps S401 to S404 in FIG. 27, and therefore a description thereof will be omitted.
[0342] In step S465, the bandwidth limiting unit 464 determines whether max(TBXSize, TBYSize) of the transform block to be processed is equal to or greater than the predetermined size THSize. If it is determined in step S465 that max(TBXSize, TBYSize) is equal to or greater than the predetermined size THSize, the process proceeds to step S466.
[0343] In step S466, the band-limiting unit 464 applies the band-limiting filter H defined by the above-mentioned equation (38) to the secondary transform coefficient Coeff according to the above-mentioned equation (35) to perform band-limiting, and derives the band-limited secondary transform coefficient Coeff'. Then, the band-limiting unit 464 supplies the band-limited secondary transform coefficient Coeff' to the quantization unit 113, and the process ends.
[0344] On the other hand, if it is determined in step S465 that max(TBXSize, TBYSize) is less than the predetermined size THSize, the band-limiting unit 464 skips the process of step S466. Then, the band-limiting unit 464 supplies the secondary transform coefficients Coeff as they are to the quantization unit 113 as the secondary transform coefficients Coeff' after band-limiting, and ends the process.
[0345] Note that the order of the steps of the transformation process may be changed or the details of the process may be modified to the extent possible. For example, if it is determined in step S463 that the secondary transformation identifier st_idx is not greater than 0, a unit matrix may be selected as the matrix R of the secondary transformation, and the process of step S464 may be performed.
[0346] As described above, the band limiting unit 464 performs band limiting based on the horizontal size TBXSize and vertical size TBYSize of the transform block. Therefore, even if the shape of the transform block is rectangular, it is possible to perform band limiting appropriately.
[0347] Furthermore, the band limiting unit 464 sets the region where the secondary transform coefficient Coeff is set to 0 by band limiting to a region other than a square region made up of TH×TH sub-blocks. Therefore, even if the shape of the transform block is rectangular, the secondary transform coefficient Coeff can be set to 0 in a region other than a square region of a predetermined size.
[0348] <Image decoding device> Next, decoding of the coded data coded as above will be described. The configuration of the fourth embodiment of the image decoding device as an image processing device to which the present technology is applied is the same as the configuration of the third embodiment, except for the configurations of the transformation information Tinfo and the inverse transformation unit. Since the transformation information Tinfo has been described above, only the inverse transformation unit will be described below.
[0349] <Inverse conversion section> FIG. 34 is a block diagram showing an example of the main configuration of an inverse transform unit in the fourth embodiment of an image decoding device as an image processing device to which the present technology is applied.
[0350] Of the components shown in Fig. 34, the same components as those in Fig. 28 are denoted by the same reference numerals, and duplicated explanations will be omitted where appropriate.
[0351] The configuration of the inverse transform unit 480 in FIG. 34 differs from the configuration of the inverse transform unit 440 in FIG. 28 in that a band limiting unit 482 is provided instead of the band limiting unit 442.
[0352] The band limiting unit 482 of the inverse transform unit 480 performs band limiting on the secondary transform coefficients Coeff_IQ supplied from the switch 441 so as to set high-frequency components to zero, based on the horizontal size TBXSize and vertical size TBYSize of the transform block to be processed supplied from the decoding unit 211. The band limiting unit 482 supplies the band-limited secondary transform coefficients Coeff_IQ' to the inverse secondary transform unit 443.
[0353] Specifically, if max(TBXSize, TBYSize) is less than a predetermined size THSize, the band limiting unit 482 does not perform band limiting on the secondary transform coefficient Coeff_IQ, and supplies the secondary transform coefficient Coeff_IQ to the inverse secondary transform unit 443 as the band-limited secondary transform coefficient Coeff_IQ'.
[0354] On the other hand, if max(TBXSize, TBYSize) is equal to or larger than the predetermined size THSize, the band-limiting unit 482 applies the band-limiting filter H defined by the above-mentioned equation (38) to the secondary transform coefficients Coeff_IQ according to the above-mentioned equation (37). The band-limiting unit 482 supplies the secondary transform coefficients Coeff_IQ′ after the band limitation derived as a result to the inverse secondary transform unit 443.
[0355] <Flow of image decoding process> The image decoding process performed by the fourth embodiment of the image decoding device as an image processing device to which the present technology is applied is the same as the image decoding process of Figure 15 except for the inverse transform process of step S203, so below we will only describe the inverse transform process.
[0356] <Flow of the inverse conversion process> FIG. 35 is a flowchart illustrating an example of the flow of inverse transform processing executed by the fourth embodiment of the image decoding device as the image processing device to which the present technology is applied.
[0357] When the inverse transform process starts, in step S481, the switch 441 determines whether the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag supplied from the decoding unit 211 is 1 or not.
[0358] If it is determined in step S481 that the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 441 determines to skip the band limiting, inverse secondary transform, and inverse primary transform. Then, the switch 441 does not perform the processes of steps S482 to S486, but supplies the secondary transform coefficients Coeff_IQ supplied from the inverse quantization unit 212 to the calculation unit 214 as prediction residuals D', and ends the process.
[0359] On the other hand, if it is determined in step S481 that the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0, the switch 441 determines to perform band limitation, inverse secondary transform, and inverse primary transform. Then, the switch 441 supplies the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212 to the band limitation unit 482, and the process proceeds to step S482.
[0360] In step S482, the bandwidth limiting unit 482 determines whether max(TBXSize, TBYSize) is equal to or greater than the predetermined size THSize. If it is determined in step S482 that max(TBXSize, TBYSize) is equal to or greater than the predetermined size THSize, the process proceeds to step S483.
[0361] In step S483, the band-limiting unit 482 applies the band-limiting filter H to the secondary transform coefficient Coeff_IQ defined by the above-mentioned equation (38) to perform band-limiting using the above-mentioned equation (37), and derives the band-limited secondary transform coefficient Coeff_IQ'. Then, the band-limiting unit 464 supplies the band-limited secondary transform coefficient Coeff_IQ' to the inverse secondary transform unit 443, and the process proceeds to step S484.
[0362] On the other hand, if it is determined in step S482 that max(TBXSize, TBYSize) is less than the predetermined size THSize, the band-limiting unit 482 skips the process of step S483. Then, the band-limiting unit 482 supplies the secondary transform coefficients Coeff_IQ as they are to the inverse secondary transform unit 443 as the band-limited secondary transform coefficients Coeff_IQ′, and proceeds to the process of step S484.
[0363] The processing in steps S484 to S486 is the same as the processing in steps S424 to S426 in FIG. 29, and therefore a description thereof will be omitted.
[0364] As described above, the band limiting unit 482 performs band limiting based on the horizontal size TBXSize and vertical size TBYSize of the transform block. Therefore, even if the shape of the transform block is rectangular, the band limiting unit 482 can perform band limiting appropriately.
[0365] Furthermore, the band limiting unit 482 sets the region where the secondary transform coefficient Coeff is set to 0 by band limiting to a region other than a square region made up of TH×TH sub-blocks. Therefore, even if the shape of the transform block is rectangular, the secondary transform coefficient Coeff can be set to 0 in a region other than a square region of a predetermined size.
[0366] In the above description, bandwidth limitation was performed in the fourth embodiment of the image decoding device. However, if a bitstream constraint is set such that "when max(TBXSize, TBYSize) is equal to or greater than a predetermined size THSize, non-zero coefficients are restricted to the upper left low-frequency region (e.g., the region of (THSize>>1)×(THSize>>1)", bandwidth limitation need not be performed. This bitstream constraint can also be expressed as "when the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0 and max(TBXSize, TBYSize) of the transform block is equal to or greater than a predetermined size THSize, the values of the last non-zero coefficient X coordinate (last_sig_coeff_x_pos) and the last non-zero coefficient Y coordinate (last_sig_coeff_y_pos) are within the range from 0 to a predetermined value (e.g., (THSize / 2)-1)".
[0367] Furthermore, the secondary transform method in the fourth embodiment is not limited to the method described in Non-Patent Document 1. For example, it may be the method in the first or second embodiment. Furthermore, the secondary transform coefficients may be clipped to be limited to a predetermined range.
[0368] Furthermore, the method of determining whether to perform bandwidth limitation is not limited to the above-mentioned method, as long as it is a method of determining whether the size of the transform block is greater than or equal to a predetermined size based on the horizontal size TBXSize (or its logarithmic value log2TBXSize) and the vertical size TBYSize (or its logarithmic value log2TBYSize).
[0369] For example, the bandwidth limiting unit 464 (482) may perform bandwidth limiting when max(log2TBXSize, log2TBYSize), which is the larger of the logarithms of the horizontal size TBXSize and the vertical size TBYSize, is equal to or greater than a predetermined threshold value log2THSize (when the logical value of max(log2TBXSize, log2TBYSize)>=log2THSize) is 1 (true)).Furthermore, the bandwidth limiting unit 464 (482) may perform bandwidth limiting when the sum or product of the horizontal size TBXSize and the vertical size TBYSize is equal to or greater than a predetermined threshold value TH_Size (when the logical value of TBXSize+TBYSize>=TH_Size or TBXSize*TBYSize>=TH_Size is 1 (true)). Furthermore, the bandwidth limiting unit 464 (482) may perform bandwidth limiting when the sum of the logarithms of the horizontal size TBXSize and the vertical size TBYSize is equal to or greater than a predetermined threshold value log2THSize (when the logical value of log2TBXSize+log2TBYSize>=log2THSize is 1 (true)). In these cases, when the logical value is 0 (false), no bandwidth limiting is performed.
[0370] Furthermore, the band-limiting unit 464 (482) may set the band-limiting filter H(i,j) to 0 or 1 depending on the position of the coordinate (i,j) rather than the position of the sub-block containing the coordinate (i,j). In this case, the band-limiting filter H(i,j) is defined by the following equation (39).
[0371]
number
[0372] <Another example of a band-limiting filter> 36 to 38 are diagrams showing other examples of the band-limiting filter H(i,j).
[0373] 36 to 38, the band-limiting filter H(i,j) is defined by the following equation (40) based on the processing order (scan order) sbk_scanorder(i>>N,j>>M) of the sub-block including the coordinates (i,j). Note that the sub-block scan order sbk_scanorder(i>>N,j>>M) is determined based on the scan identifier scanIdx and the size of the transform block.
[0374]
number
[0375] THscanorder is a threshold value. In equation (40), the threshold value THscanorder can be set to, for example, alpha (alpha is a value between 0 and 1) times the total number of sub-blocks in the transform block.
[0376] According to equation (35) and equation (40), the coordinates (i, j) are N x2 M If the scan order sbk_scanorder(i>>N,j>>M) of the sub-blocks is equal to or greater than the threshold THscanorder, the secondary transform coefficient Coeff(i,j)' becomes 0. In other words, the secondary transform coefficient Coeff(i,j)' of the sub-blocks of high frequency components other than the THscanorder number of sub-blocks with the earliest scan order becomes 0.
[0377] In the examples of FIGS. 36 to 38, the size of the sub-block is 2 2 x2 2and the size THSize is a size of 8 sub-blocks. Furthermore, in the example of Fig. 36, the threshold THscanorder in the bandlimiting filter H(i,j) is 10, and in the examples of Figs. 37 and 38, it is 16. Furthermore, A of Fig. 36 to A of Fig. 38 show an example where the transform block size TBXSize×TBYSize is a size consisting of 8×8 sub-blocks (a transform block size of 32×32), B of Fig. 36 to B of Fig. 38 show an example where the transform block size TBXSize×TBYSize is a size consisting of 8×4 sub-blocks (a transform block size of 32×16), and C of Fig. 36 to C of Fig. 38 show an example where the transform block size TBXSize×TBYSize is a size consisting of 4×8 sub-blocks (a transform block size of 16×32).
[0378] In this case, as shown in A of Figure 36 to A of Figure 38, when both the horizontal size TBXSize and the vertical size TBYSize are the size of eight sub-blocks, a band-limiting filter H(i, j) is set.
[0379] As shown in A of Figure 36, when the sub-blocks are scanned diagonally (up-right diagonal scan), the band-limiting filter H(i,j) of the coordinates (i,j) included in the sub-block group 501A of the 10 shaded sub-blocks in the upper left of the figure is 1, and the band-limiting filter H(i,j) of the coordinates (i,j) in the sub-blocks other than that sub-block group 501A is 0.
[0380] Also, as shown in A of Figure 37, when the sub-blocks are scanned horizontally (horizontal fast scan), the band-limiting filter H(i,j) at the coordinates (i,j) included in the upper 16 shaded sub-blocks 502A in the figure is 1, and the band-limiting filter H(i,j) at the coordinates (i,j) in sub-blocks other than that sub-block group 502A is 0.
[0381] As shown in A of Figure 38, when sub-blocks are scanned vertically (vertical fast scan), the band-limiting filters H(i,j) at the coordinates (i,j) included in the 16 shaded sub-blocks 503A on the left side of the figure are 1, and the band-limiting filters H(i,j) at the coordinates (i,j) in sub-blocks other than the sub-block group 503A are 0.
[0382] Also, as shown in B of Figure 36 to B of Figure 38, the band-limiting filter H(i, j) is set when the horizontal size TBXSize is the size of 8 sub-blocks and the vertical size TBYSize is the size of 4 sub-blocks smaller than the horizontal size TBXSize.
[0383] As shown in B of Figure 36, when sub-blocks are scanned diagonally, the band-limiting filter H(i,j) at the coordinates (i,j) included in the top left 10 shaded sub-blocks 501B in the figure is 1, and the band-limiting filter H(i,j) at the coordinates (i,j) in sub-blocks other than that sub-block group 501B is 0.
[0384] Also, as shown in B of Figure 37, when sub-blocks are scanned horizontally, the band-limiting filters H(i,j) at the coordinates (i,j) included in the upper 16 shaded sub-block groups 502B in the figure are 1, and the band-limiting filters H(i,j) at the coordinates (i,j) in sub-blocks other than those sub-block group 502B are 0.
[0385] As shown in B of Figure 38, when sub-blocks are scanned vertically, the band-limiting filter H(i,j) at the coordinates (i,j) included in the 16 shaded sub-blocks 503B on the left side of the figure is 1, and the band-limiting filter H(i,j) at the coordinates (i,j) in sub-blocks other than that sub-block group 503B is 0.
[0386] Furthermore, as shown in C of Figure 36 to C of Figure 38, the band-limiting filter H(i, j) is also set when the horizontal size TBXSize is a size of four sub-blocks and the vertical size TBYSize is a size of eight sub-blocks larger than the horizontal size TBXSize.
[0387] As shown in C of Figure 36, when sub-blocks are scanned diagonally, the band-limiting filter H(i,j) at the coordinates (i,j) included in the top left 10 shaded sub-block group 501C in the figure is 1, and the band-limiting filter H(i,j) at the coordinates (i,j) in sub-blocks other than that sub-block group 501C is 0.
[0388] Also, as shown in C of Figure 37, when the sub-blocks are scanned horizontally, the band-limiting filters H(i,j) at the coordinates (i,j) included in the upper 16 shaded sub-block groups 502C in the figure are 1, and the band-limiting filters H(i,j) at the coordinates (i,j) in sub-blocks other than the sub-block group 502C are 0.
[0389] As shown in C of Figure 38, when sub-blocks are scanned vertically, the band-limiting filters H(i,j) at coordinates (i,j) included in the 16 shaded sub-blocks 503C on the left side of the figure are 1, and the band-limiting filters H(i,j) at coordinates (i,j) in sub-blocks other than the sub-block group 503C are 0.
[0390] As described above, in the examples of Figures 36 to 38, when max(TBXSize, TBYSize) is equal to or larger than the size THSize, the band-limiting filter H(i, j) is set to 0 or 1 based on the scan order of the sub-block including the coordinates (i, j). Note that in the examples of Figures 36 to 38, the threshold value THscanorder is different when the sub-block is scanned diagonally and when the sub-block is scanned horizontally or vertically, but it may also be the same.
[0391] The band-limiting filter H(i,j) may be defined by the following equation (41) based on the scan order coef_scanorder(i,j) of the coordinates (i,j). The scan order coef_scanorder(i,j) is determined based on the scan identifier scanIdx and the size of the transform block.
[0392]
number
[0393] In equation (41), the threshold THscanorder can be set to, for example, alpha (alpha is a value between 0 and 1) times the total number of pixels in the transform block (TBXSize×TBYSize).
[0394] <Another example of the flow of conversion and inverse conversion processing> In the fourth embodiment, the bandlimiting may be controlled not only based on the transform block size but also based on a bandlimiting filter enable flag bandpass_filter_enabled_flag indicating whether or not to apply a bandlimiting filter when the transform block size is equal to or larger than a predetermined block size. In this case, the control unit 101 includes the bandlimiting filter enable flag bandpass_filter_enabled_flag in the encoding parameters in units of SPS, PPS, SH, CU, etc. This bandlimiting filter enable flag bandpass_filter_enabled_flag is supplied to the encoding unit 114 for encoding, and is also supplied to the bandlimiting unit 464 of the transform unit 460.
[0395] FIG. 39 is a flowchart illustrating an example of the flow of the conversion process of the conversion unit 460 in this case.
[0396] The processing in steps S501 to S504 in FIG. 39 is similar to the processing in steps S461 to S464 in FIG. 33, and therefore a description thereof will be omitted.
[0397] In step S505, the bandpass unit 464 determines whether the bandpass filter enable flag bandpass_filter_enabled_flag is 1 (true), which indicates that a bandpass filter is applied when the transform block size is equal to or larger than a predetermined block size. If it is determined in step S505 that the bandpass filter enable flag bandpass_filter_enabled_flag is 1, the process proceeds to step S506. The processes of steps S506 and S507 are the same as the processes of steps S465 and S466 in FIG. 33, and therefore description thereof will be omitted.
[0398] On the other hand, if it is determined in step S505 that the bandpass filter enable flag bandpass_filter_enabled_flag is 0 (false), which indicates that the bandpass filter is not applied even if the transform block size is equal to or larger than the predetermined block size, the transform process ends.
[0399] As described above, the bandpass unit 464 controls whether to apply a bandpass filter that limits non-zero coefficients to the low frequency range when the block size of a rectangular transform block consisting of a square or rectangle is equal to or larger than a predetermined block size, based on the bandpass filter enable flag bandpass_filter_enabled_flag.
[0400] Specifically, when the bandpass filter enable flag bandpass_filter_enabled_flag is 1 (true), if the transform block size is equal to or larger than a predetermined block size, the bandpass unit 464 performs band limiting so that the transform coefficients of high frequency components are forcibly set to 0. Therefore, the encoding unit 114 encodes only the non-zero coefficients of the low frequency components of the transform block. This makes it possible to reduce the encoding process of non-zero coefficients while maintaining encoding efficiency.
[0401] Furthermore, when the bandpass filter enable flag bandpass_filter_enabled_flag is 0 (false), the bandpass unit 464 does not perform band limiting even when the transform block is equal to or larger than a predetermined block size. Therefore, the encoding unit 114 encodes non-zero coefficients present in the low-frequency to high-frequency components. Therefore, in large transform blocks, the reproducibility of high-frequency components is improved and encoding efficiency is improved compared to when a bandpass filter is applied.
[0402] Furthermore, the bandpass_filter_enabled_flag is set for each predetermined unit (SPS, PPS, SH, CU, etc.). Therefore, the bandpass unit 464 can use the bandpass_filter_enabled_flag to adaptively control, for each predetermined unit, whether to apply a bandpass filter when the transform block size is equal to or larger than a predetermined block size.
[0403] Furthermore, in the fourth embodiment, when controlling the band limiting based on not only the transform block size but also the band limiting filter enable flag bandpass_filter_enabled_flag, the decoding unit 211 decodes the coded data of the coding parameters including the band limiting filter enable flag bandpass_filter_enabled_flag in a predetermined unit (SPS, PPS, SH, CU, etc.). The band limiting filter enable flag bandpass_filter_enabled_flag obtained as a result of the decoding is supplied to the band limiting unit 482 of the inverse transform unit 480.
[0404] FIG. 40 is a flowchart illustrating an example of the flow of the inverse conversion process of the inverse conversion unit 480 in this case.
[0405] In step S521 of FIG. 40, the switch 441 determines whether the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag supplied from the decoding unit 211 is 1 or not.
[0406] If it is determined in step S521 that the transform skip flag ts_flag or the transform quantization bypass flag transquant_bypass_flag is 1, the switch 441 determines to skip the band limiting, the inverse secondary transform, and the inverse primary transform. Then, the switch 441 does not perform the processes of steps S522 to S527, but supplies the secondary transform coefficients Coeff_IQ supplied from the inverse quantization unit 212 to the calculation unit 214 as the prediction residuals D', and ends the process.
[0407] On the other hand, if it is determined in step S521 that the transform skip flag ts_flag and the transform quantization bypass flag transquant_bypass_flag are 0, the switch 441 supplies the secondary transform coefficient Coeff_IQ supplied from the inverse quantization unit 212 to the band limiting unit 482. Then, the processing proceeds to step S522.
[0408] In step S522, the bandpass unit 482 determines whether the value of the bandpass filter enable flag bandpass_filter_enabled_flag is 1 (true). If it is determined in step S522 that the value of the bandpass filter enable flag is 1 (true), the process proceeds to step S523. The processes of steps S523 to S527 are the same as the processes of steps S482 to S486 in Fig. 35, and therefore description thereof will be omitted.
[0409] If it is determined in step S522 that the value of the band-limiting filter valid flag is 0 (false), the process skips steps S523 and S524 and proceeds to step S525.
[0410] As described above, the bandpass unit 482 controls whether to apply a bandpass filter that limits non-zero coefficients to the low frequency range when the block size of a rectangular transform block consisting of a square or rectangle is equal to or larger than a predetermined block size, based on the bandpass filter enable flag bandpass_filter_enabled_flag.
[0411] Specifically, when the bandpass filter enable flag bandpass_filter_enabled_flag is 1 (true) and the transform block size is equal to or larger than a predetermined block size, the bandpass unit 482 performs band limitation so that the transform coefficients of high-frequency components are forcibly set to zero. Therefore, the calculation unit 214 decodes only the non-zero coefficients of the low-frequency components of the transform block. This reduces the amount of decoding processing for non-zero coefficients while maintaining coding efficiency. Furthermore, the inverse transform unit 213 only needs to perform inverse transform processing on the non-zero coefficients of the low-frequency components of the transform block, thereby reducing the amount of inverse transform processing. Therefore, the computational complexity and implementation costs of the image decoding device can be reduced compared to when the image encoding device does not perform band limitation.
[0412] Furthermore, when the bandpass filter enable flag bandpass_filter_enabled_flag is 0 (false), the bandpass unit 482 does not perform bandpass filtering, even when the transform block is equal to or larger than a predetermined block size. Therefore, the calculation unit 214 decodes non-zero coefficients present in the low- to high-frequency components. Therefore, in large transform blocks, the reproducibility of high-frequency components is improved compared to when a bandpass filter is applied, resulting in improved image quality of the decoded image.
[0413] Furthermore, the bandpass_filter_enabled_flag is set for each predetermined unit (SPS, PPS, SH, CU, etc.) Therefore, the bandpass unit 482 can use the bandpass_filter_enabled_flag to adaptively control, for each predetermined unit, whether to apply a bandpass filter when the transform block size is equal to or larger than a predetermined block size.
[0414] In the examples of Figures 39 and 40, whether the transform block size is equal to or larger than a predetermined block size is determined by whether max(TBXSize, TBYSize) is equal to or larger than a predetermined size THSize, but the determination may also be made using other of the methods described above.
[0415] Also in the third embodiment, similarly to the fourth embodiment, band limitation may be controlled based not only on the transform block size but also on the bandpass filter enable flag bandpass_filter_enabled_flag.
[0416] <5. Fifth Embodiment> <Data unit of information> The data units in which the information about the image and the information about encoding and decoding of the image are set (or the data that is the target) described above are each arbitrary and are not limited to the above examples. For example, this information may be set for each TU, TB, PU, PB, CU, LCU, sub-block, block, tile, slice, picture, sequence, or component, or may target data of these data units. Of course, this data unit is set for each piece of information. In other words, all of the information does not have to be set for (or target) the same data unit. Note that the storage location of this information is arbitrary and may be stored in the header or parameter set of the above-mentioned data unit, or the like. Also, it may be stored in multiple locations.
[0417] <Control information> Control information related to the present technology described in each of the above embodiments may be transmitted from the encoding side to the decoding side. For example, control information (e.g., enabled_flag) that controls whether or not to permit (or prohibit) application of the above-described present technology may be transmitted. Also, for example, control information that specifies an upper limit or a lower limit, or both, of a block size that permits (or prohibits) application of the above-described present technology may be transmitted.
[0418] <Encoding / Decoding> The present technology can be applied to any image encoding / decoding that performs primary transform and secondary transform (inverse secondary transform and inverse primary transform). In other words, the specifications of transform (inverse transform), quantization (inverse quantization), encoding (decoding), prediction, etc. are arbitrary and are not limited to the above-mentioned examples. For example, in the transform (inverse transform), (inverse) transforms other than (inverse) primary transform and (inverse) secondary transform (i.e., three or more (inverse) transforms) may be performed. Furthermore, encoding (decoding) may be a lossless method or a lossy method. Furthermore, quantization (inverse quantization), prediction, etc. may be omitted. Furthermore, processing not described above, such as filtering, may be performed.
[0419] <Application areas of this technology> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, such as transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring.
[0420] For example, the present technology can be applied to systems and devices that transmit images for viewing. Furthermore, for example, the present technology can be applied to systems and devices used for transportation. Furthermore, for example, the present technology can be applied to systems and devices used for security. Furthermore, for example, the present technology can be applied to systems and devices used for sports. Furthermore, for example, the present technology can be applied to systems and devices used for agriculture. Furthermore, for example, the present technology can be applied to systems and devices used for livestock farming. Furthermore, the present technology can be applied to systems and devices that monitor natural conditions such as volcanoes, forests, and oceans. Furthermore, the present technology can be applied to weather observation systems and weather observation devices that observe weather, temperature, humidity, wind speed, sunshine hours, and the like. Furthermore, the present technology can be applied to systems and devices that observe the ecology of wildlife such as birds, fish, reptiles, amphibians, mammals, insects, and plants.
[0421] <Application to multi-viewpoint image encoding and decoding systems> The above-described series of processes can be applied to a multi-viewpoint image encoding / decoding system that encodes and decodes multi-viewpoint images including images from multiple views. In this case, the present technology can be applied to the encoding and decoding of each view.
[0422] <Application to hierarchical image coding and decoding systems> Furthermore, the above-described series of processes can be applied to a scalable image coding / decoding system that encodes / decodes layered images that are layered to have a scalability function for a predetermined parameter. In this case, the present technology can be applied to the encoding / decoding of each layer.
[0423] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, etc., that can execute various functions by installing various programs.
[0424] FIG. 41 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0425] In a computer 800 shown in FIG. 41, a CPU (Central Processing Unit) 801, a ROM (Read Only Memory) 802, and a RAM (Random Access Memory) 803 are interconnected via a bus 804.
[0426] An input / output interface 810 is also connected to the bus 804. To the input / output interface 810, an input unit 811, an output unit 812, a storage unit 813, a communication unit 814, and a drive 815 are connected.
[0427] The input unit 811 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, etc. The output unit 812 includes, for example, a display, a speaker, an output terminal, etc. The storage unit 813 includes, for example, a hard disk, a RAM disk, a non-volatile memory, etc. The communication unit 814 includes, for example, a network interface. The drive 815 drives removable media 821 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0428] In a computer configured as above, the CPU 801 performs the above-described series of processes by, for example, loading a program stored in the storage unit 813 into the RAM 803 via the input / output interface 810 and the bus 804 and executing the program. The RAM 803 also stores data necessary for the CPU 801 to execute various processes as appropriate.
[0429] The program executed by the computer (CPU 801) can be applied by recording it on removable media 821 such as package media, for example. In this case, the program can be installed in the storage unit 813 via the input / output interface 810 by inserting the removable media 821 into the drive 815.
[0430] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 814 and installed in the storage unit 813.
[0431] Alternatively, this program can be installed in advance in the ROM 802 or the storage unit 813 .
[0432] <Applications of this technology> The image encoding device 100 and the image decoding device 200 according to the above-described embodiments can be applied to various electronic devices, such as transmitters and receivers for cable broadcasting such as satellite broadcasting and cable TV, distribution over the Internet, and distribution to terminals via cellular communication, or recording devices that record images on media such as optical disks, magnetic disks, and flash memories, and playback devices that play back images from these storage media.
[0433] <First application example: Television receiver> 42 shows an example of a schematic configuration of a television device to which the above-described embodiment is applied. The television device 900 includes an antenna 901, a tuner 902, a demultiplexer 903, a decoder 904, a video signal processing unit 905, a display unit 906, an audio signal processing unit 907, a speaker 908, an external interface (I / F) unit 909, a control unit 910, a user interface (I / F) unit 911, and a bus 912.
[0434] The tuner 902 extracts a signal of a desired channel from a broadcast signal received via the antenna 901 and demodulates the extracted signal. The tuner 902 then outputs the coded bit stream obtained by demodulation to the demultiplexer 903. In other words, the tuner 902 serves as a transmission unit in the television device 900 that receives a coded stream in which an image is coded.
[0435] The demultiplexer 903 separates the video stream and audio stream of the program to be viewed from the coded bitstream, and outputs each separated stream to the decoder 904. The demultiplexer 903 also extracts auxiliary data such as an EPG (Electronic Program Guide) from the coded bitstream, and supplies the extracted data to the control unit 910. Note that the demultiplexer 903 may descramble the coded bitstream if it is scrambled.
[0436] The decoder 904 decodes the video stream and audio stream input from the demultiplexer 903. Then, the decoder 904 outputs the video data generated by the decoding process to a video signal processing unit 905. The decoder 904 also outputs the audio data generated by the decoding process to an audio signal processing unit 907.
[0437] The video signal processing unit 905 reproduces the video data input from the decoder 904 and displays the video on the display unit 906. The video signal processing unit 905 may also display an application screen supplied via a network on the display unit 906. The video signal processing unit 905 may also perform additional processing on the video data, such as noise removal, depending on the settings. Furthermore, the video signal processing unit 905 may generate images of a GUI (Graphical User Interface), such as a menu, button, or cursor, and superimpose the generated image on the output image.
[0438] The display unit 906 is driven by a drive signal supplied from the video signal processing unit 905, and displays a video or image on the screen of a display device (e.g., a liquid crystal display, a plasma display, or an OLED (Organic ElectroLuminescence Display)).
[0439] The audio signal processing unit 907 performs playback processing such as D / A conversion and amplification on the audio data input from the decoder 904, and outputs audio from a speaker 908. The audio signal processing unit 907 may also perform additional processing such as noise removal on the audio data.
[0440] The external interface unit 909 is an interface for connecting the television device 900 to an external device or a network. For example, a video stream or an audio stream received via the external interface unit 909 may be decoded by the decoder 904. That is, the external interface unit 909 also serves as a transmission unit in the television device 900 that receives an encoded stream in which an image is encoded.
[0441] The control unit 910 has a processor such as a CPU, and memories such as RAM and ROM. The memory stores programs executed by the CPU, program data, EPG data, data acquired via a network, and the like. The programs stored in the memory are read and executed by the CPU, for example, when the television device 900 is started up. By executing the programs, the CPU controls the operation of the television device 900 in response to operation signals input from the user interface unit 911, for example.
[0442] The user interface unit 911 is connected to the control unit 910. The user interface unit 911 has, for example, buttons and switches for the user to operate the television device 900, a receiver for remote control signals, etc. The user interface unit 911 detects operations by the user via these components, generates an operation signal, and outputs the generated operation signal to the control unit 910.
[0443] The bus 912 interconnects the tuner 902, demultiplexer 903, decoder 904, video signal processing unit 905, audio signal processing unit 907, external interface unit 909, and control unit 910.
[0444] In the television device 900 configured in this manner, the decoder 904 may have the functions of the above-described image decoding device 200. That is, the decoder 904 may decode coded data using the methods described in the above-described embodiments. In this way, the television device 900 can obtain the same effects as those of the embodiments described above with reference to Figs. 1 to 23.
[0445] Furthermore, in the television device 900 configured as above, the video signal processing unit 905 may be configured to encode image data supplied from the decoder 904, for example, and output the resulting encoded data to the outside of the television device 900 via the external interface unit 909. The video signal processing unit 905 may then have the functions of the above-mentioned image encoding device 100. That is, the video signal processing unit 905 may be configured to encode the image data supplied from the decoder 904 using the methods described in each of the above-mentioned embodiments. In this way, the television device 900 can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23.
[0446] <Second application example: mobile phone> 43 shows an example of the schematic configuration of a mobile phone to which the above-described embodiment is applied. The mobile phone 920 includes an antenna 921, a communication unit 922, an audio codec 923, a speaker 924, a microphone 925, a camera unit 926, an image processing unit 927, a demultiplexing unit 928, a recording / playback unit 929, a display unit 930, a control unit 931, an operation unit 932, and a bus 933.
[0447] The antenna 921 is connected to the communication unit 922. The speaker 924 and microphone 925 are connected to the audio codec 923. The operation unit 932 is connected to the control unit 931. The bus 933 interconnects the communication unit 922, audio codec 923, camera unit 926, image processing unit 927, demultiplexing unit 928, recording / playback unit 929, display unit 930, and control unit 931.
[0448] The mobile phone 920 performs operations such as sending and receiving voice signals, sending and receiving e-mail or image data, taking images, and recording data in various operating modes including voice call mode, data communication mode, photography mode, and videophone mode.
[0449] In the voice call mode, an analog voice signal generated by the microphone 925 is supplied to the voice codec 923. The voice codec 923 converts the analog voice signal into voice data, A / D converts the converted voice data, and compresses it. The voice codec 923 then outputs the compressed voice data to the communication unit 922. The communication unit 922 encodes and modulates the voice data to generate a transmission signal. The communication unit 922 then transmits the generated transmission signal to a base station (not shown) via the antenna 921. The communication unit 922 also amplifies and frequency-converts a wireless signal received via the antenna 921 to acquire a received signal. The communication unit 922 then demodulates and decodes the received signal to generate voice data, and outputs the generated voice data to the voice codec 923. The voice codec 923 then expands and D / A-converts the voice data to generate an analog voice signal. The voice codec 923 then supplies the generated voice signal to the speaker 924 to output voice.
[0450] In the data communication mode, for example, the control unit 931 generates character data constituting an email in response to a user operation via the operation unit 932. The control unit 931 also displays the characters on the display unit 930. The control unit 931 also generates email data in response to a transmission instruction from the user via the operation unit 932 and outputs the generated email data to the communication unit 922. The communication unit 922 encodes and modulates the email data to generate a transmission signal. The communication unit 922 then transmits the generated transmission signal to a base station (not shown) via the antenna 921. The communication unit 922 also amplifies and frequency-converts a radio signal received via the antenna 921 to obtain a received signal. The communication unit 922 then demodulates and decodes the received signal to restore the email data and outputs the restored email data to the control unit 931. The control unit 931 also displays the contents of the email on the display unit 930 and supplies the email data to the recording / playback unit 929, where it is written to its storage medium.
[0451] The recording / playback unit 929 has any readable / writable storage medium. For example, the storage medium may be a built-in storage medium such as RAM or flash memory, or may be an external storage medium such as a hard disk, a magnetic disk, a magneto-optical disk, an optical disk, a USB (Universal Serial Bus) memory, or a memory card.
[0452] Also, in the shooting mode, for example, the camera unit 926 captures an image of a subject to generate image data and outputs the generated image data to the image processing unit 927. The image processing unit 927 encodes the image data input from the camera unit 926, supplies the encoded stream to the recording / playback unit 929, and causes it to be written into its storage medium.
[0453] Furthermore, in the image display mode, the recording / playback unit 929 reads out the coded stream recorded on the storage medium and outputs it to the image processing unit 927. The image processing unit 927 decodes the coded stream input from the recording / playback unit 929, supplies the image data to the display unit 930, and displays the image.
[0454] Also, in videophone mode, for example, the demultiplexing unit 928 multiplexes the video stream encoded by the image processing unit 927 with the audio stream input from the audio codec 923 and outputs the multiplexed stream to the communication unit 922. The communication unit 922 encodes and modulates the stream to generate a transmission signal. The communication unit 922 then transmits the generated transmission signal to a base station (not shown) via the antenna 921. The communication unit 922 also amplifies and frequency-converts a wireless signal received via the antenna 921 to obtain a received signal. These transmission signals and received signals may include coded bitstreams. The communication unit 922 then demodulates and decodes the received signal to restore the stream and outputs the restored stream to the demultiplexing unit 928. The demultiplexing unit 928 separates the input stream into a video stream and an audio stream and outputs the video stream to the image processing unit 927 and the audio stream to the audio codec 923. The image processing unit 927 decodes the video stream to generate video data. The video data is supplied to a display unit 930, and a series of images is displayed on the display unit 930. The audio codec 923 expands the audio stream and performs D / A conversion to generate an analog audio signal. The audio codec 923 then supplies the generated audio signal to a speaker 924 to output the audio.
[0455] In the mobile phone 920 configured in this manner, for example, the image processing unit 927 may have the functions of the image encoding device 100 described above. That is, the image processing unit 927 may encode image data using the method described in each of the above embodiments. By doing so, the mobile phone 920 can obtain the same effects as those of each of the above embodiments described with reference to Figures 1 to 23.
[0456] Furthermore, in the mobile phone 920 configured as above, for example, the image processing unit 927 may have the functions of the above-described image decoding device 200. That is, the image processing unit 927 may decode coded data using the methods described in the above-described embodiments. By doing so, the mobile phone 920 can obtain the same effects as those of the embodiments described above with reference to Figs. 1 to 23.
[0457] <Third application example: recording / playback device> 44 shows an example of a schematic configuration of a recording / playback device to which the above-described embodiment is applied. The recording / playback device 940, for example, encodes audio data and video data of a received broadcast program and records the encoded data on a recording medium. The recording / playback device 940 may also encode audio data and video data acquired from another device and record the encoded data on a recording medium. The recording / playback device 940 may also play data recorded on a recording medium on a monitor and a speaker in response to a user's instruction, for example. At this time, the recording / playback device 940 decodes the audio data and video data.
[0458] The recording / playback device 940 includes a tuner 941, an external interface (I / F) unit 942, an encoder 943, an HDD (Hard Disk Drive) unit 944, a disk drive 945, a selector 946, a decoder 947, an OSD (On-Screen Display) unit 948, a control unit 949, and a user interface (I / F) unit 950.
[0459] The tuner 941 extracts a signal of a desired channel from a broadcast signal received via an antenna (not shown) and demodulates the extracted signal. The tuner 941 then outputs the coded bit stream obtained by demodulation to the selector 946. That is, the tuner 941 serves as a transmission unit in the recording / playback device 940.
[0460] The external interface unit 942 is an interface for connecting the recording / playback device 940 to an external device or a network. The external interface unit 942 may be, for example, an IEEE (Institute of Electrical and Electronic Engineers) 1394 interface, a network interface, a USB interface, or a flash memory interface. For example, video data and audio data received via the external interface unit 942 are input to an encoder 943. In other words, the external interface unit 942 serves as a transmission unit in the recording / playback device 940.
[0461] The encoder 943 encodes the video data and audio data when the video data and audio data input from the external interface unit 942 have not been encoded. The encoder 943 then outputs the encoded bit stream to the selector 946.
[0462] The HDD unit 944 records coded bit streams in which content data such as video and audio is compressed, various programs, and other data on an internal hard disk, and also reads this data from the hard disk when playing back video and audio.
[0463] The disc drive 945 records and reads data on a recording medium that is loaded into the disc drive 945. The recording medium loaded into the disc drive 945 may be, for example, a DVD (Digital Versatile Disc) disc (DVD-Video, DVD-RAM (DVD - Random Access Memory), DVD-R (DVD - Recordable), DVD-RW (DVD - Rewritable), DVD+R (DVD + Recordable), DVD+RW (DVD + Rewritable), etc.) or a Blu-ray (registered trademark) disc.
[0464] When recording video and audio, the selector 946 selects the coded bit stream input from the tuner 941 or the encoder 943, and outputs the selected coded bit stream to the HDD 944 or the disk drive 945. When playing video and audio, the selector 946 outputs the coded bit stream input from the HDD 944 or the disk drive 945 to the decoder 947.
[0465] The decoder 947 decodes the coded bitstream to generate video data and audio data, and outputs the generated video data to an OSD unit 948. The decoder 947 also outputs the generated audio data to an external speaker.
[0466] The OSD unit 948 reproduces and displays the video data input from the decoder 947. The OSD unit 948 may also superimpose GUI images such as menus, buttons, or cursors on the video to be displayed.
[0467] The control unit 949 has a processor such as a CPU, and memories such as RAM and ROM. The memory stores programs executed by the CPU, program data, etc. The programs stored in the memory are read and executed by the CPU, for example, when the recording / playback device 940 is started up. By executing the programs, the CPU controls the operation of the recording / playback device 940 in response to operation signals input from the user interface unit 950, for example.
[0468] The user interface unit 950 is connected to the control unit 949. The user interface unit 950 has, for example, buttons and switches for the user to operate the recording / playback device 940, a unit for receiving remote control signals, etc. The user interface unit 950 detects operations by the user via these components, generates operation signals, and outputs the generated operation signals to the control unit 949.
[0469] In the recording / reproducing device 940 configured in this manner, for example, the encoder 943 may have the functions of the above-described image encoding device 100. That is, the encoder 943 may encode image data using the methods described in the above-described embodiments. In this way, the recording / reproducing device 940 can obtain the same effects as those of the embodiments described above with reference to Figures 1 to 23.
[0470] Furthermore, in the recording / reproducing device 940 configured as above, for example, the decoder 947 may have the functions of the above-mentioned image decoding device 200. That is, the decoder 947 may decode coded data using the methods described in the above-mentioned embodiments. In this way, the recording / reproducing device 940 can obtain the same effects as those of the above-mentioned embodiments with reference to Figures 1 to 23.
[0471] <Fourth application example: imaging device> 45 shows an example of a schematic configuration of an imaging device to which the above-described embodiments are applied. The imaging device 960 captures an image of a subject to generate an image, encodes the image data, and records it on a recording medium.
[0472] The imaging device 960 includes an optical block 961, an imaging unit 962, a signal processing unit 963, an image processing unit 964, a display unit 965, an external interface (I / F) unit 966, a memory unit 967, a media drive 968, an OSD unit 969, a control unit 970, a user interface (I / F) unit 971, and a bus 972.
[0473] The optical block 961 is connected to an imaging unit 962. The imaging unit 962 is connected to a signal processing unit 963. The display unit 965 is connected to an image processing unit 964. The user interface unit 971 is connected to a control unit 970. The bus 972 interconnects the image processing unit 964, the external interface unit 966, the memory unit 967, the media drive 968, the OSD unit 969, and the control unit 970.
[0474] The optical block 961 has a focus lens, an aperture mechanism, etc. The optical block 961 forms an optical image of a subject on an imaging surface of the imaging unit 962. The imaging unit 962 has an image sensor such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor), and converts the optical image formed on the imaging surface into an image signal as an electrical signal by photoelectric conversion. The imaging unit 962 then outputs the image signal to the signal processing unit 963.
[0475] The signal processing unit 963 performs various camera signal processing such as knee correction, gamma correction, and color correction on the image signal input from the imaging unit 962. The signal processing unit 963 outputs the image data after the camera signal processing to the image processing unit 964.
[0476] The image processing unit 964 encodes image data input from the signal processing unit 963 to generate encoded data. Then, the image processing unit 964 outputs the generated encoded data to the external interface unit 966 or the media drive 968. Furthermore, the image processing unit 964 decodes the encoded data input from the external interface unit 966 or the media drive 968 to generate image data. Then, the image processing unit 964 outputs the generated image data to the display unit 965. Furthermore, the image processing unit 964 may output the image data input from the signal processing unit 963 to the display unit 965 to display an image. Furthermore, the image processing unit 964 may superimpose display data acquired from the OSD unit 969 on the image to be output to the display unit 965.
[0477] The OSD unit 969 generates GUI images such as menus, buttons, or cursors, and outputs the generated images to the image processing unit 964.
[0478] The external interface unit 966 is configured as, for example, a USB input / output terminal. The external interface unit 966 connects the imaging device 960 to a printer, for example, when printing an image. A drive is also connected to the external interface unit 966 as needed. Removable media such as a magnetic disk or optical disk is loaded into the drive, and a program read from the removable media can be installed in the imaging device 960. Furthermore, the external interface unit 966 may be configured as a network interface connected to a network such as a LAN or the Internet. That is, the external interface unit 966 serves as a transmission unit in the imaging device 960.
[0479] The recording medium attached to the media drive 968 may be any removable readable / writable medium, such as a magnetic disk, a magneto-optical disk, an optical disk, or a semiconductor memory. Alternatively, a recording medium may be fixedly attached to the media drive 968, forming a non-portable storage unit such as an internal hard disk drive or an SSD (Solid State Drive).
[0480] The control unit 970 has a processor such as a CPU, and memories such as RAM and ROM. The memory stores programs executed by the CPU, program data, and the like. The programs stored in the memory are read and executed by the CPU, for example, when the imaging device 960 is started up. By executing the programs, the CPU controls the operation of the imaging device 960 in response to operation signals input from the user interface unit 971, for example.
[0481] The user interface unit 971 is connected to the control unit 970. The user interface unit 971 has, for example, buttons and switches that allow the user to operate the imaging device 960. The user interface unit 971 detects operations by the user via these components, generates an operation signal, and outputs the generated operation signal to the control unit 970.
[0482] In the imaging device 960 configured in this manner, for example, the image processing unit 964 may have the functions of the above-described image encoding device 100. That is, the image processing unit 964 may encode image data using the methods described in the above-described embodiments. In this manner, the imaging device 960 can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23.
[0483] Furthermore, in the imaging device 960 configured as above, for example, the image processing unit 964 may have the functions of the above-described image decoding device 200. That is, the image processing unit 964 may decode coded data using the methods described in the above-described embodiments. By doing so, the imaging device 960 can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23.
[0484] <5th Application Example: Video Set> In addition, the present technology can be implemented as any configuration mounted on an arbitrary device or device constituting a system, for example, a processor as a system LSI (Large Scale Integration) or the like, a module using multiple processors or the like, a unit using multiple modules or the like, a set in which other functions are further added to a unit, etc. (i.e., a configuration of a part of a device). Fig. 46 shows an example of a schematic configuration of a video set to which the present technology is applied.
[0485] In recent years, electronic devices have become increasingly multifunctional, and in their development and manufacture, when some of their components are sold or provided, it is not only the case that they are implemented as a component with one function, but also that multiple components with related functions are combined to be implemented as a set with multiple functions.
[0486] The video set 1300 shown in Figure 46 is such a multi-functional configuration, combining a device having functions related to image encoding and decoding (either one or both) with devices having other functions related to those functions.
[0487] As shown in FIG. 46, the video set 1300 includes a group of modules such as a video module 1311, an external memory 1312, a power management module 1313, and a front-end module 1314, as well as devices with associated functions such as connectivity 1321, a camera 1322, and a sensor 1323.
[0488] A module is a component that combines several interrelated component functions into a cohesive function. While the specific physical configuration is arbitrary, it may, for example, be an integrated circuit consisting of multiple processors, each with its own function, electronic circuit elements such as resistors and capacitors, and other devices, all arranged on a wiring board or similar. It is also conceivable to combine a module with other modules or processors to create a new module.
[0489] In the example of FIG. 46, the video module 1311 is a combination of components having functions related to image processing, and includes an application processor, a video processor, a broadband modem 1333, and an RF module 1334.
[0490] A processor is a device in which a configuration having a predetermined function is integrated on a semiconductor chip using SoC (System On a Chip), and is also known as, for example, System LSI (Large Scale Integration). This configuration having a predetermined function may be a logic circuit (hardware configuration), or may be a CPU, ROM, RAM, etc., and a program executed using these (software configuration), or may be a combination of both. For example, a processor may have a logic circuit, a CPU, ROM, RAM, etc., and realize some of its functions by the logic circuit (hardware configuration), and realize other functions by a program executed by the CPU (software configuration).
[0491] 46 is a processor that executes applications related to image processing. The applications executed by this application processor 1331 not only perform arithmetic processing to realize predetermined functions, but also control the configuration inside and outside the video module 1311, such as the video processor 1332, as necessary.
[0492] The video processor 1332 is a processor having functions related to image encoding and / or decoding.
[0493] The broadband modem 1333 digitally modulates or otherwise converts data (digital signals) transmitted via wired or wireless (or both) broadband communication over broadband lines such as the Internet or a public telephone network into analog signals, and demodulates analog signals received via such broadband communication and converts them into data (digital signals). The broadband modem 1333 processes any information, such as image data processed by the video processor 1332, streams in which image data is encoded, application programs, and setting data.
[0494] The RF module 1334 is a module that performs frequency conversion, modulation / demodulation, amplification, filtering, etc. on an RF (Radio Frequency) signal transmitted and received via an antenna. For example, the RF module 1334 performs frequency conversion, etc. on a baseband signal generated by the broadband modem 1333 to generate an RF signal. Also, for example, the RF module 1334 performs frequency conversion, etc. on an RF signal received via the front-end module 1314 to generate a baseband signal.
[0495] As indicated by a dotted line 1341 in FIG. 46, the application processor 1331 and the video processor 1332 may be integrated into one processor.
[0496] The external memory 1312 is a module provided outside the video module 1311 and having a storage device used by the video module 1311. The storage device of this external memory 1312 may be realized by any physical configuration, but since it is generally used to store large amounts of data such as image data in units of frames, it is desirable to realize it by a relatively inexpensive, large-capacity semiconductor memory such as a DRAM (Dynamic Random Access Memory).
[0497] The power management module 1313 manages and controls the power supply to the video module 1311 (each component within the video module 1311).
[0498] The front-end module 1314 is a module that provides a front-end function (circuitry at the transmitting and receiving end on the antenna side) for the RF module 1334. As shown in FIG. 46, the front-end module 1314 includes, for example, an antenna unit 1351, a filter 1352, and an amplifier unit 1353.
[0499] The antenna unit 1351 has an antenna for transmitting and receiving radio signals and its peripheral components. The antenna unit 1351 transmits signals supplied from the amplifier unit 1353 as radio signals, and supplies the received radio signals to the filter 1352 as electric signals (RF signals). The filter 1352 performs filtering and other processes on the RF signals received via the antenna unit 1351, and supplies the processed RF signals to the RF module 1334. The amplifier unit 1353 amplifies the RF signals supplied from the RF module 1334, and supplies the amplified signals to the antenna unit 1351.
[0500] The connectivity 1321 is a module having a function related to connection with the outside. The physical configuration of the connectivity 1321 is arbitrary. For example, the connectivity 1321 may have a configuration having a communication function other than the communication standard supported by the broadband modem 1333, an external input / output terminal, etc.
[0501] For example, the connectivity 1321 may include a module having a communication function conforming to a wireless communication standard such as Bluetooth (registered trademark), IEEE 802.11 (e.g., Wi-Fi (Wireless Fidelity, registered trademark)), NFC (Near Field Communication), or IrDA (InfraRed Data Association), or an antenna for transmitting and receiving signals conforming to the standard. Furthermore, for example, the connectivity 1321 may include a module having a communication function conforming to a wired communication standard such as USB (Universal Serial Bus) or HDMI (High-Definition Multimedia Interface, registered trademark), or a terminal conforming to the standard. Furthermore, for example, the connectivity 1321 may include other data (signal) transmission functions such as an analog input / output terminal.
[0502] The connectivity 1321 may include a device to which data (signals) are transmitted. For example, the connectivity 1321 may include a drive (including not only removable media drives but also hard disks, solid state drives (SSDs), network attached storage (NASs), etc.) that reads and writes data from and to recording media such as magnetic disks, optical disks, magneto-optical disks, or semiconductor memories. The connectivity 1321 may also include image and audio output devices (monitors, speakers, etc.).
[0503] The camera 1322 is a module having a function of capturing an image of a subject and obtaining image data of the subject. The image data obtained by capturing an image with the camera 1322 is supplied to, for example, a video processor 1332 and encoded there.
[0504] The sensor 1323 is a module having any sensor function, such as an audio sensor, ultrasonic sensor, light sensor, illuminance sensor, infrared sensor, image sensor, rotation sensor, angle sensor, angular velocity sensor, speed sensor, acceleration sensor, tilt sensor, magnetic identification sensor, impact sensor, temperature sensor, etc. Data detected by the sensor 1323 is supplied to the application processor 1331, for example, and used by an application or the like.
[0505] The configuration described above as a module may be realized as a processor, and conversely, the configuration described above as a processor may be realized as a module.
[0506] In the video set 1300 configured as above, as will be described later, the present technology can be applied to the video processor 1332. Therefore, the video set 1300 can be implemented as a set to which the present technology is applied.
[0507] <Video processor configuration example> FIG. 47 shows an example of a schematic configuration of the video processor 1332 (FIG. 46) to which the present technology is applied.
[0508] In the example of Figure 47, the video processor 1332 has the function of receiving input video and audio signals and encoding them in a predetermined format, and the function of decoding the encoded video and audio data and playing back and outputting the video and audio signals.
[0509] 47, the video processor 1332 includes a video input processing unit 1401, a first image scaling unit 1402, a second image scaling unit 1403, a video output processing unit 1404, a frame memory 1405, and a memory control unit 1406. The video processor 1332 also includes an encoding / decoding engine 1407, video ES (Elementary Stream) buffers 1408A and 1408B, and audio ES buffers 1409A and 1409B. The video processor 1332 also includes an audio encoder 1410, an audio decoder 1411, a multiplexer (MUX) 1412, a demultiplexer (DMUX) 1413, and a stream buffer 1414.
[0510] The video input processing unit 1401 acquires a video signal input from, for example, the connectivity 1321 (FIG. 46) and converts it into digital image data. The first image scaling unit 1402 performs format conversion, image scaling, and other processes on the image data. The second image scaling unit 1403 performs image scaling on the image data in accordance with the format of the output destination via the video output processing unit 1404, and performs format conversion and image scaling similar to those performed by the first image scaling unit 1402. The video output processing unit 1404 performs format conversion and conversion to an analog signal on the image data, and outputs it as a reproduced video signal to, for example, the connectivity 1321.
[0511] The frame memory 1405 is a memory for storing image data that is shared by the video input processing unit 1401, the first image scaling unit 1402, the second image scaling unit 1403, the video output processing unit 1404, and the encoding / decoding engine 1407. The frame memory 1405 is realized as a semiconductor memory such as a DRAM.
[0512] The memory control unit 1406 receives a synchronization signal from the encoding / decoding engine 1407 and controls access to the frame memory 1405 for writing and reading in accordance with an access schedule to the frame memory 1405 written in an access management table 1406A. The access management table 1406A is updated by the memory control unit 1406 in accordance with the processing executed by the encoding / decoding engine 1407, the first image scaling unit 1402, the second image scaling unit 1403, etc.
[0513] The encoding / decoding engine 1407 performs encoding processing on image data and decoding processing on a video stream, which is data obtained by encoding image data. For example, the encoding / decoding engine 1407 encodes image data read from the frame memory 1405 and sequentially writes the encoded image data to the video ES buffer 1408A as a video stream. For example, the encoding / decoding engine 1407 sequentially reads and decodes the video stream from the video ES buffer 1408B and sequentially writes the decoded image data to the frame memory 1405 as image data. The encoding / decoding engine 1407 uses the frame memory 1405 as a working area for these encoding and decoding operations. The encoding / decoding engine 1407 also outputs a synchronization signal to the memory control unit 1406, for example, at the timing when processing for each macroblock starts.
[0514] The video ES buffer 1408A buffers the video stream generated by the encoding / decoding engine 1407 and supplies it to the multiplexing unit (MUX) 1412. The video ES buffer 1408B buffers the video stream supplied from the demultiplexing unit (DMUX) 1413 and supplies it to the encoding / decoding engine 1407.
[0515] The audio ES buffer 1409A buffers the audio stream generated by the audio encoder 1410 and supplies it to a multiplexer (MUX) 1412. The audio ES buffer 1409B buffers the audio stream supplied from a demultiplexer (DMUX) 1413 and supplies it to an audio decoder 1411.
[0516] The audio encoder 1410 digitally converts an audio signal input from, for example, the connectivity 1321, and encodes it using a predetermined format such as the MPEG audio format or AC3 (Audio Code number 3) format. The audio encoder 1410 sequentially writes an audio stream, which is data obtained by encoding the audio signal, into the audio ES buffer 1409A. The audio decoder 1411 decodes the audio stream supplied from the audio ES buffer 1409B, converts it into an analog signal, and supplies it to, for example, the connectivity 1321 as a reproduced audio signal.
[0517] The multiplexing unit (MUX) 1412 multiplexes a video stream and an audio stream. This multiplexing method (i.e., the format of the bitstream generated by multiplexing) is arbitrary. Furthermore, during this multiplexing, the multiplexing unit (MUX) 1412 can also add predetermined header information, etc. to the bitstream. In other words, the multiplexing unit (MUX) 1412 can convert the format of the stream by multiplexing. For example, the multiplexing unit (MUX) 1412 multiplexes a video stream and an audio stream to convert them into a transport stream, which is a bitstream in a format for transfer. Furthermore, for example, the multiplexing unit (MUX) 1412 multiplexes a video stream and an audio stream to convert them into data (file data) in a file format for recording.
[0518] The demultiplexer (DMUX) 1413 demultiplexes a bitstream in which a video stream and an audio stream are multiplexed, using a method corresponding to the multiplexing performed by the multiplexer (MUX) 1412. That is, the demultiplexer (DMUX) 1413 extracts the video stream and the audio stream from the bitstream read from the stream buffer 1414 (separates the video stream from the audio stream). That is, the demultiplexer (DMUX) 1413 can convert the format of the stream by demultiplexing (reverse conversion of the conversion performed by the multiplexer (MUX) 1412). For example, the demultiplexer (DMUX) 1413 can obtain a transport stream supplied from, for example, the connectivity 1321 or the broadband modem 1333 via the stream buffer 1414 and demultiplex it to convert it into a video stream and an audio stream. Furthermore, for example, the demultiplexing unit (DMUX) 1413 can obtain file data read from various recording media by the connectivity 1321 via the stream buffer 1414 and demultiplex the data to convert it into a video stream and an audio stream.
[0519] The stream buffer 1414 buffers the bit stream. For example, the stream buffer 1414 buffers the transport stream supplied from the multiplexer (MUX) 1412 and supplies it to, for example, the connectivity 1321 or the broadband modem 1333 at a predetermined timing or based on an external request.
[0520] Also, for example, the stream buffer 1414 buffers the file data supplied from the multiplexing unit (MUX) 1412, and supplies it to, for example, connectivity 1321 at a predetermined timing or based on an external request, etc., and records it on various recording media.
[0521] Furthermore, the stream buffer 1414 buffers transport streams acquired, for example, via the connectivity 1321 or broadband modem 1333, and supplies them to the demultiplexer (DMUX) 1413 at a predetermined timing or based on an external request.
[0522] The stream buffer 1414 also buffers file data read from various recording media, for example, in the connectivity 1321, and supplies the data to a demultiplexer (DMUX) 1413 at a predetermined timing or based on an external request.
[0523] Next, an example of the operation of the video processor 1332 configured as described above will be described. For example, a video signal input to the video processor 1332 from the connectivity 1321 or the like is converted into digital image data in a predetermined format such as the 4:2:2 Y / Cb / Cr format by the video input processing unit 1401, and sequentially written to the frame memory 1405. This digital image data is read out to the first image scaling unit 1402 or the second image scaling unit 1403, where it is subjected to format conversion to a predetermined format such as the 4:2:0 Y / Cb / Cr format and scaling processing, and then written back to the frame memory 1405. This image data is coded by the encoding / decoding engine 1407 and written as a video stream to the video ES buffer 1408A.
[0524] Furthermore, an audio signal input to the video processor 1332 from the connectivity 1321 or the like is encoded by the audio encoder 1410 and written as an audio stream to the audio ES buffer 1409A.
[0525] The video stream in video ES buffer 1408A and the audio stream in audio ES buffer 1409A are read out and multiplexed by multiplexing unit (MUX) 1412 and converted into a transport stream, file data, or the like. The transport stream generated by multiplexing unit (MUX) 1412 is buffered in stream buffer 1414 and then output to an external network via, for example, connectivity 1321 or broadband modem 1333. In addition, the file data generated by multiplexing unit (MUX) 1412 is buffered in stream buffer 1414 and then output to, for example, connectivity 1321, and recorded on various recording media.
[0526] Furthermore, a transport stream input to the video processor 1332 from an external network via, for example, the connectivity 1321 or the broadband modem 1333 is buffered in the stream buffer 1414 and then demultiplexed by the demultiplexer (DMUX) 1413. Furthermore, file data read from various recording media, for example, in the connectivity 1321, and input to the video processor 1332 is buffered in the stream buffer 1414 and then demultiplexed by the demultiplexer (DMUX) 1413. In other words, the transport stream or file data input to the video processor 1332 is separated by the demultiplexer (DMUX) 1413 into a video stream and an audio stream.
[0527] The audio stream is supplied to an audio decoder 1411 via an audio ES buffer 1409B, where it is decoded and the audio signal is reproduced. Meanwhile, the video stream is written to a video ES buffer 1408B, and then sequentially read and decoded by an encoding / decoding engine 1407, and written to a frame memory 1405. The decoded image data is enlarged or reduced by a second image scaling unit 1403, and written to the frame memory 1405. The decoded image data is then read to a video output processing unit 1404, where it is format-converted to a predetermined format such as the 4:2:2 Y / Cb / Cr format, and further converted to an analog signal, and the video signal is reproduced and output.
[0528] When the present technology is applied to the video processor 1332 configured in this manner, it is sufficient to apply the present technology according to each of the above-described embodiments to the encoding / decoding engine 1407. That is, for example, the encoding / decoding engine 1407 may have the functions of the above-described image encoding device 100, the functions of the image decoding device 200, or both. In this way, the video processor 1332 can obtain the same effects as those of each of the above-described embodiments with reference to FIGS. 1 to 23.
[0529] In addition, in the encoding / decoding engine 1407, the present technology (i.e., the functions of the image encoding device 100 or the functions of the image decoding device 200, or both) may be realized by hardware such as a logic circuit, or by software such as an embedded program, or by both.
[0530] <Other examples of video processor configurations> Fig. 48 shows another example of a schematic configuration of the video processor 1332 to which the present technology is applied. In the example of Fig. 48, the video processor 1332 has a function of encoding and decoding video data in a predetermined format.
[0531] 48, the video processor 1332 has a control unit 1511, a display interface 1512, a display engine 1513, an image processing engine 1514, and an internal memory 1515. The video processor 1332 also has a codec engine 1516, a memory interface 1517, a multiplexing / demultiplexing unit (MUX / DMUX) 1518, a network interface 1519, and a video interface 1520.
[0532] The control unit 1511 controls the operation of each processing unit in the video processor 1332, such as the display interface 1512, the display engine 1513, the image processing engine 1514, and the codec engine 1516.
[0533] As shown in FIG. 48, the control unit 1511 includes, for example, a main CPU 1531, a sub-CPU 1532, and a system controller 1533. The main CPU 1531 executes programs and the like for controlling the operation of each processing unit in the video processor 1332. The main CPU 1531 generates control signals in accordance with the programs and supplies them to each processing unit (i.e., controls the operation of each processing unit). The sub-CPU 1532 plays an auxiliary role to the main CPU 1531. For example, the sub-CPU 1532 executes child processes and subroutines of programs and the like executed by the main CPU 1531. The system controller 1533 controls the operation of the main CPU 1531 and the sub-CPU 1532, for example, by specifying the programs executed by the main CPU 1531 and the sub-CPU 1532.
[0534] The display interface 1512 outputs image data to, for example, the connectivity 1321 under the control of the control unit 1511. For example, the display interface 1512 converts digital image data into an analog signal and outputs the analog signal to a monitor device of the connectivity 1321 as a reproduced video signal or as the digital image data itself.
[0535] Under the control of the control unit 1511, the display engine 1513 performs various conversion processes such as format conversion, size conversion, and color gamut conversion on the image data so that the image matches the hardware specifications of the monitor device or the like that displays the image.
[0536] Under the control of the control unit 1511, the image processing engine 1514 performs predetermined image processing on the image data, such as filtering to improve image quality.
[0537] The internal memory 1515 is a memory provided inside the video processor 1332 and shared by the display engine 1513, the image processing engine 1514, and the codec engine 1516. The internal memory 1515 is used, for example, for data exchange between the display engine 1513, the image processing engine 1514, and the codec engine 1516. For example, the internal memory 1515 stores data supplied from the display engine 1513, the image processing engine 1514, or the codec engine 1516, and supplies the data to the display engine 1513, the image processing engine 1514, or the codec engine 1516 as needed (for example, in response to a request). The internal memory 1515 may be realized by any storage device, but is generally often used to store small amounts of data such as block-based image data and parameters, and therefore is desirably realized by a semiconductor memory such as an SRAM (Static Random Access Memory) that has a relatively small capacity (compared to, for example, the external memory 1312) but a high response speed.
[0538] The codec engine 1516 performs processing related to encoding and decoding of image data. The codec engine 1516 can support any encoding and decoding method, and the number of methods may be one or more. For example, the codec engine 1516 may be provided with codec functions for multiple encoding and decoding methods, and may encode image data or decode encoded data using a codec selected from among the methods.
[0539] In the example shown in FIG. 48, the codec engine 1516 has, as functional blocks for codec-related processing, for example, MPEG-2 Video 1541, AVC / H.264 1542, HEVC / H.265 1543, HEVC / H.265 (Scalable) 1544, HEVC / H.265 (Multi-view) 1545, and MPEG-DASH 1551.
[0540] MPEG-2 Video 1541 is a functional block that encodes and decodes image data in the MPEG-2 format. AVC / H.264 1542 is a functional block that encodes and decodes image data in the AVC format. HEVC / H.265 1543 is a functional block that encodes and decodes image data in the HEVC format. HEVC / H.265 (Scalable) 1544 is a functional block that scalably encodes and scalably decodes image data in the HEVC format. HEVC / H.265 (Multi-view) 1545 is a functional block that multi-view encodes and multi-view decodes image data in the HEVC format.
[0541] MPEG-DASH1551 is a functional block that transmits and receives image data using the MPEG-DASH (MPEG-Dynamic Adaptive Streaming over HTTP) method. MPEG-DASH is a technology for streaming video using HTTP (HyperText Transfer Protocol), and one of its features is that it selects and transmits the appropriate data segment by segment from multiple pre-prepared encoded data sets with different resolutions, etc. MPEG-DASH1551 generates streams that comply with the standard and controls the transmission of those streams. For encoding and decoding of image data, it uses the above-mentioned MPEG-2 Video1541 to HEVC / H.265 (Multi-view)1545.
[0542] The memory interface 1517 is an interface for the external memory 1312. Data supplied from the image processing engine 1514 or the codec engine 1516 is supplied to the external memory 1312 via the memory interface 1517. Data read from the external memory 1312 is supplied to the video processor 1332 (the image processing engine 1514 or the codec engine 1516) via the memory interface 1517.
[0543] The multiplexing / demultiplexing unit (MUX DMUX) 1518 multiplexes and demultiplexes various types of image-related data, such as coded data bit streams, image data, and video signals. Any method of multiplexing / demultiplexing may be used. For example, during multiplexing, the multiplexing / demultiplexing unit (MUX DMUX) 1518 can not only combine multiple pieces of data into one, but also add predetermined header information, etc. to the data. During demultiplexing, the multiplexing / demultiplexing unit (MUX DMUX) 1518 can not only divide one piece of data into multiple pieces, but also add predetermined header information, etc. to each piece of divided data. In other words, the multiplexing / demultiplexing unit (MUX DMUX) 1518 can convert data formats through multiplexing / demultiplexing. For example, the multiplexing / demultiplexing unit (MUX DMUX) 1518 can multiplex bit streams to convert them into a transport stream, which is a bit stream in a format for transfer, or into data (file data) in a file format for recording. Of course, the reverse conversion is also possible by demultiplexing.
[0544] The network interface 1519 is an interface for, for example, the broadband modem 1333, the connectivity 1321, etc. The video interface 1520 is an interface for, for example, the connectivity 1321, the camera 1322, etc.
[0545] Next, an example of the operation of such a video processor 1332 will be described. For example, when a transport stream is received from an external network via the connectivity 1321, broadband modem 1333, or the like, the transport stream is supplied to a multiplexing / demultiplexing unit (MUX DMUX) 1518 via a network interface 1519, where it is demultiplexed, and decoded by a codec engine 1516. Image data obtained by decoding by the codec engine 1516 is subjected to predetermined image processing by, for example, an image processing engine 1514, and then subjected to predetermined conversion by a display engine 1513. The image data is then supplied to, for example, the connectivity 1321 via a display interface 1512, and the image is displayed on a monitor. Furthermore, for example, image data obtained by decoding by the codec engine 1516 is re-encoded by the codec engine 1516, multiplexed by a multiplexing / demultiplexing unit (MUX DMUX) 1518 and converted into file data, which is output via a video interface 1520 to, for example, connectivity 1321, and recorded on various recording media.
[0546] Furthermore, for example, file data of coded data obtained by coding image data read from a recording medium (not shown) by the connectivity 1321 or the like is supplied to the multiplexing and demultiplexing unit (MUX DMUX) 1518 via the video interface 1520, where it is demultiplexed, and decoded by the codec engine 1516. The image data obtained by decoding by the codec engine 1516 is subjected to predetermined image processing by the image processing engine 1514, subjected to predetermined conversion by the display engine 1513, and supplied to, for example, the connectivity 1321 or the like via the display interface 1512, where the image is displayed on a monitor. Also, for example, image data obtained by decoding by the codec engine 1516 is re-encoded by the codec engine 1516, multiplexed by the multiplexing and demultiplexing unit (MUX DMUX) 1518 and converted into a transport stream, and supplied to, for example, the connectivity 1321 or the broadband modem 1333 via the network interface 1519, and transmitted to another device (not shown).
[0547] Note that image data and other data are exchanged between the processing units in the video processor 1332, for example, using the internal memory 1515 or the external memory 1312. The power management module 1313 controls the power supply to the control unit 1511, for example.
[0548] When the present technology is applied to the video processor 1332 configured in this manner, it is sufficient to apply the present technology according to each of the above-described embodiments to the codec engine 1516. That is, for example, it is sufficient to make the codec engine 1516 have the functions of the above-described image encoding device 100, the functions of the image decoding device 200, or both. In this way, the video processor 1332 can obtain the same effects as those of each of the above-described embodiments with reference to FIGS. 1 to 23.
[0549] In the codec engine 1516, the present technology (i.e., the functions of the image encoding device 100) may be realized by hardware such as a logic circuit, by software such as an embedded program, or by both.
[0550] Although two exemplary configurations of the video processor 1332 have been given above, the configuration of the video processor 1332 is arbitrary and may be other than the two examples described above. Furthermore, the video processor 1332 may be configured as a single semiconductor chip, or may be configured as multiple semiconductor chips. For example, it may be a three-dimensional stacked LSI in which multiple semiconductors are stacked. It may also be realized by multiple LSIs.
[0551] <Example of application to equipment> The video set 1300 can be incorporated into various devices that process image data. For example, the video set 1300 can be incorporated into a television device 900 (FIG. 42), a mobile phone 920 (FIG. 43), a recording / playback device 940 (FIG. 44), an imaging device 960 (FIG. 45), etc. By incorporating the video set 1300, the device can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23.
[0552] Note that even a portion of each component of the video set 1300 described above can be implemented as a configuration to which the present technology is applied, as long as it includes the video processor 1332. For example, only the video processor 1332 can be implemented as a video processor to which the present technology is applied. Also, for example, the processor and video module 1311 indicated by the dotted line 1341 as described above can be implemented as a processor, module, etc. to which the present technology is applied. Furthermore, for example, the video module 1311, external memory 1312, power management module 1313, and front-end module 1314 can be combined and implemented as a video unit 1361 to which the present technology is applied. In any configuration, the same effects as those of the embodiments described above with reference to FIGS. 1 to 23 can be obtained.
[0553] That is, any configuration including the video processor 1332 can be incorporated into various devices that process image data, similar to the case of the video set 1300. For example, the video processor 1332, the processor indicated by the dotted line 1341, the video module 1311, or the video unit 1361 can be incorporated into a television device 900 (FIG. 42), a mobile phone 920 (FIG. 43), a recording / playback device 940 (FIG. 44), an imaging device 960 (FIG. 45), or the like. By incorporating any of the configurations to which the present technology is applied, the device can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23, similar to the case of the video set 1300.
[0554] <Sixth Application Example: Network System> This technology can also be applied to a network system made up of multiple devices. Figure 49 shows an example of a schematic configuration of a network system to which this technology is applied.
[0555] Network system 1600 shown in FIG. 49 is a system in which devices exchange information related to images (video images) with each other via a network. Cloud service 1601 of this network system 1600 is a system that provides services related to images (video images) to terminals such as a computer 1611, an AV (Audio Visual) device 1612, a portable information processing terminal 1613, and an IoT (Internet of Things) device 1614 that are communicatively connected to the system. For example, cloud service 1601 provides an image (video image) content supply service, such as so-called video distribution (on-demand or live distribution), to the terminals. Also, for example, cloud service 1601 provides a backup service that receives and stores image (video image) content from the terminals. Also, for example, cloud service 1601 provides a service that mediates the exchange of image (video image) content between terminals.
[0556] The physical configuration of the cloud service 1601 is arbitrary. For example, the cloud service 1601 may include various servers such as a server that stores and manages moving images, a server that distributes moving images to terminals, a server that acquires moving images from terminals, and a server that manages users (terminals) and billing, and an arbitrary network such as the Internet or a LAN.
[0557] The computer 1611 is, for example, an information processing device such as a personal computer, a server, or a workstation. The AV equipment 1612 is, for example, an image processing device such as a television receiver, a hard disk recorder, a game console, or a camera. The portable information processing terminal 1613 is, for example, a portable information processing device such as a notebook personal computer, a tablet terminal, a mobile phone, or a smartphone. The IoT device 1614 is, for example, any object that performs image processing, such as a machine, a home appliance, furniture, other objects, an IC tag, or a card-type device. All of these terminals have a communication function and can connect to the cloud service 1601 (establish a session) and exchange information with the cloud service 1601 (i.e., communicate). Each terminal can also communicate with other terminals. Communication between terminals may be performed via the cloud service 1601 or may be performed without the cloud service 1601.
[0558] The present technology may be applied to the above-described network system 1600, and when image (video) data is exchanged between terminals or between a terminal and a cloud service 1601, the image data may be encoded and decoded as described above in each embodiment. That is, the terminals (computers 1611 to IoT devices 1614) and the cloud service 1601 may have the functions of the image encoding device 100 and the image decoding device 200, respectively. In this way, the terminals (computers 1611 to IoT devices 1614) and the cloud service 1601 that exchange image data can obtain the same effects as those of the embodiments described above with reference to FIGS. 1 to 23.
[0559] <Other> Various pieces of information related to the coded data (bit stream) may be multiplexed onto the coded data and transmitted or recorded, or may be transmitted or recorded as separate data associated with the coded data without being multiplexed onto the coded data. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. That is, pieces of data associated with each other may be combined into one piece of data, or may be separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of the same recording medium). This "association" may refer not to the entire data, but to only a portion of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.
[0560] Also, as mentioned above, in this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as mentioned above.
[0561] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0562] For example, in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.
[0563] Also, for example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0564] Furthermore, for example, this technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0565] Furthermore, for example, the above-described program can be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.
[0566] For example, each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. Furthermore, if one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0567] In addition, the processing of the steps of the program executed by the computer may be executed in chronological order according to the order described in this specification, or may be executed in parallel, or individually at the required timing such as when called, etc. Furthermore, the processing of the steps of the program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0568] It should be noted that the present technologies described in this specification can be implemented independently and singly, unless a contradiction arises. Of course, any two or more of the present technologies can also be implemented in combination. For example, the present technology described in any of the embodiments can be implemented in combination with the present technology described in another embodiment. Furthermore, any of the present technologies described above can also be implemented in combination with other technologies not described above.
[0569] The present technology can also be configured as follows. (1) a control unit that, when skipping a primary transform that is a transform process on a prediction residual that is a difference between an image and a predicted image of the image and a secondary transform that is a transform process on a primary transform coefficient obtained by subjecting the prediction residual to the primary transform, also skips band limitation on a secondary transform coefficient obtained by subjecting the primary transform coefficient to the secondary transform. An image processing device comprising: (2) When the control unit does not skip the primary transform, and when a transform block size of the primary transform and the secondary transform is equal to or larger than a predetermined size, the control unit performs band limitation on the secondary transform coefficients. It was configured as The image processing device according to (1). (3) When the control unit does not skip the inverse primary transform, the control unit performs band limitation on the secondary transform coefficients based on horizontal and vertical sizes of transform blocks of the inverse primary transform and the inverse secondary transform. It was configured as (2) An image processing device according to the present invention. (4) When the control unit does not skip the inverse primary transform, and when a larger one of the horizontal and vertical sizes of the transform block is equal to or larger than a predetermined value, the control unit performs band limitation on the secondary transform coefficients. It was configured as (3) An image processing device according to the present invention. (5) When the control unit does not skip the inverse primary transform, if the sum or product of the horizontal and vertical sizes of the transform block is equal to or greater than a predetermined value, the control unit performs band limitation on the secondary transform coefficients. It was configured as (3) An image processing device according to the present invention. (6) The band limitation is performed by setting the secondary transform coefficients after band limitation to zero in a region other than a square region of a predetermined size within a rectangular transform block of the inverse primary transform and the inverse secondary transform. It was configured as An image processing device according to any one of (2) to (5). (7) The band limitation is performed by setting the secondary transform coefficients after band limitation to 0 for pixels that are processed in a predetermined order or more among the pixels that constitute the transform blocks of the inverse primary transform and the inverse secondary transform. It was configured as An image processing device according to any one of (2) to (5). (8) The image processing device a control step of skipping band limitation on secondary transform coefficients obtained by subjecting the primary transform coefficients to the secondary transform when skipping a primary transform that is a transform process on a prediction residual that is a difference between an image and a predicted image of the image and a secondary transform that is a transform process on primary transform coefficients obtained by subjecting the prediction residual to the primary transform; An image processing method comprising: (9) a control unit that, when skipping an inverse primary transform that is an inverse transform of a primary transform that is a transform process on a prediction residual that is a difference between an image and a predicted image of the image, and an inverse secondary transform that is an inverse transform of a secondary transform that is a transform process on a primary transform coefficient obtained by subjecting the prediction residual to the primary transform, also skips band limitation on a secondary transform coefficient obtained by subjecting the primary transform coefficient to the secondary transform and band-limiting the primary transform coefficient. An image processing device comprising: (10) When the control unit does not skip the inverse primary transform, and when a transform block size of the inverse primary transform and the inverse secondary transform is equal to or larger than a predetermined size, the control unit performs band limitation on the secondary transform coefficients. It was configured as (9) An image processing device according to (9). (11) When the control unit does not skip the inverse primary transform, the control unit performs band limitation on the secondary transform coefficients based on horizontal and vertical sizes of transform blocks of the inverse primary transform and the inverse secondary transform. It was configured as (10) An image processing device according to (10). (12) When the control unit does not skip the inverse primary transform, and when a larger one of the horizontal and vertical sizes of the transform block is equal to or larger than a predetermined value, the control unit performs band limitation on the secondary transform coefficients. It was configured as The image processing device according to (11). (13) When the control unit does not skip the inverse primary transform, if the sum or product of the horizontal and vertical sizes of the transform block is equal to or greater than a predetermined value, the control unit performs band limitation on the secondary transform coefficients. It was configured as The image processing device according to (11). (14) The band limitation is performed by setting the secondary transform coefficients after band limitation to zero in a region other than a square region of a predetermined size within a rectangular transform block of the inverse primary transform and the inverse secondary transform. It was configured as (10) An image processing device according to (10). (15) The band limitation is performed by setting the secondary transform coefficients after band limitation to 0 for pixels that are processed in a predetermined order or more among the pixels that constitute the transform blocks of the inverse primary transform and the inverse secondary transform. It was configured as (10) An image processing device according to (10). (16) The image processing device a control step of skipping band limitation on secondary transform coefficients obtained by subjecting the primary transform coefficients to the secondary transform and band-limiting the primary transform coefficients when skipping an inverse primary transform that is an inverse transform of a primary transform that is a transform process on a prediction residual that is a difference between an image and a predicted image of the image, and an inverse secondary transform that is an inverse transform of a secondary transform that is a transform process on a primary transform coefficient obtained by subjecting the prediction residual to the primary transform. An image processing method comprising: (17) a control unit that controls, for each sub-block, skipping of a transform process for a transform coefficient obtained from a prediction residual that is a difference between an image and a predicted image of the image, based on the number of non-zero coefficients of the transform coefficients for each sub-block. An image processing device comprising: (18) The control unit controls, for each sub-block, skipping of secondary transform for primary transform coefficients obtained by primary transforming the prediction residual, based on the number of non-zero coefficients of the transform coefficients for each sub-block. (17) An image processing device according to (17). (19) a determination unit that determines, for each of the sub-blocks, whether or not to skip the secondary transform, based on quantized primary transform coefficients obtained by quantizing the primary transform coefficients and quantized secondary transform coefficients obtained by quantizing secondary transform coefficients obtained by performing the secondary transform on the primary transform coefficients; The control unit is configured to control skipping of the secondary transform for the primary transform coefficients for each of the sub-blocks according to a result of the determination by the determination unit. The image processing device according to (17) or (18). (20) The determination unit obtains the number of non-zero coefficients of the quantized primary transform coefficients and the number of non-zero coefficients of the quantized secondary transform coefficients, and determines, for each sub-block, whether to skip the secondary transform based on the obtained number of non-zero coefficients of the quantized primary transform coefficients and the number of non-zero coefficients of the quantized secondary transform coefficients and a predetermined threshold. An image processing device according to any one of (17) to (19). (twenty one) The determination unit determines to skip the secondary transform when the number of non-zero coefficients of the quantized primary transform coefficients is equal to or less than the number of non-zero coefficients of the quantized secondary transform coefficients and when the number of non-zero coefficients of the quantized primary transform coefficients or the number of non-zero coefficients of the quantized secondary transform coefficients is equal to or less than the threshold. An image processing device according to any one of (17) to (20). (twenty two) a quantization unit that quantizes the primary transform coefficients to obtain the quantized primary transform coefficients and quantizes the secondary transform coefficients to obtain the quantized secondary transform coefficients, The determination unit is configured to determine, for each of the sub-blocks, whether or not to skip the secondary transform, based on the quantized primary transform coefficients and the quantized secondary transform coefficients obtained by quantization by the quantization unit. An image processing device according to any one of (17) to (21). (twenty three) the primary transform is an orthogonal transform; The secondary transformation is transforming the primary transform coefficients into a one-dimensional vector; performing a matrix operation on the one-dimensional vector; Scaling the one-dimensional vector on which the matrix operation has been performed; Matrix the scaled one-dimensional vector It is a conversion process An image processing device according to any one of (17) to (22). (twenty four) a primary conversion unit that performs the primary conversion; a secondary conversion unit that performs the secondary conversion under control of the control unit; The image processing device according to any one of (17) to (23), further comprising: (twenty five) a quantization unit that quantizes secondary transform coefficients obtained by the secondary transform of the primary transform coefficients by the secondary transform unit or the primary transform coefficients; a coding unit that codes quantized transform coefficient levels obtained by quantizing the secondary transform coefficients or the primary transform coefficients by the quantization unit; The image processing device according to any one of (17) to (24), further comprising: (26) skipping of transform processing for transform coefficients obtained from a prediction residual, which is a difference between an image and a predicted image of the image, is controlled for each sub-block based on the number of non-zero coefficients of the transform coefficients for each sub-block. Image processing methods. (27) a control unit that controls, for each sub-block, skipping of inverse transform processing for transform coefficients that, when subjected to inverse transform processing, result in a prediction residual, which is a difference between an image and a predicted image of the image, based on the number of non-zero coefficients of the transform coefficients for each sub-block. An image processing device comprising: (28) The control unit controls, for each sub-block, skipping of the inverse secondary transform for the secondary transform coefficients obtained by decoding the encoded data, based on the number of non-zero coefficients of the transform coefficients for each sub-block. (27) An image processing device according to (27). (29) a determination unit that determines, for each of the sub-blocks, whether or not to skip the inverse secondary transform based on the secondary transform coefficients; The control unit is configured to control skipping of the inverse secondary transform for each of the sub-blocks in accordance with a result of the determination by the determination unit. The image processing device according to (27) or (28). (30) The determination unit determines to skip the inverse secondary transform when the number of non-zero coefficients of the secondary transform coefficients is equal to or less than a predetermined threshold. An image processing device according to any one of (27) to (29). (31) The inverse secondary transformation is transforming the secondary transform coefficients into a one-dimensional vector; performing a matrix operation on the one-dimensional vector; Scaling the one-dimensional vector on which the matrix operation has been performed; Matrix the scaled one-dimensional vector It is a conversion process An image processing device according to any one of (27) to (30). (32) an inverse secondary transformation unit that performs the inverse secondary transformation under the control of the control unit; An image processing device according to any one of (27) to (31). (33) The inverse primary transform unit further includes: an inverse primary transform unit that converts primary transform coefficients obtained by the inverse secondary transform of the secondary transform coefficients into the prediction residuals; An image processing device according to any one of (27) to (32). (34) further comprising an inverse quantization unit that inversely quantizes quantized transform coefficient levels obtained by decoding the encoded data, The inverse secondary transform unit is configured to perform the inverse secondary transform on the secondary transform coefficients obtained by inverse quantizing the quantized transform coefficient levels by the inverse quantization unit under control of the control unit. An image processing device according to any one of (27) to (33). (35) further comprising a decoding unit that decodes the encoded data; The inverse quantization unit is configured to inverse quantize the quantized transform coefficient levels obtained by decoding the coded data by the decoding unit. An image processing device according to any one of (27) to (34). (36) A skip of the inverse transform process for the transform coefficients for which a prediction residual, which is a difference between an image and a predicted image of the image, is obtained by performing the inverse transform process is controlled for each sub-block based on the number of non-zero coefficients of the transform coefficients for each sub-block. Image processing methods. (37) a setting unit that sets a matrix of a transformation process for a transformation coefficient based on the content of the transformation process and a scanning method; a rasterization unit that converts a prediction residual, which is a difference between an image and a predicted image of the image, into a transformation coefficient obtained by transformation processing the prediction residual into a one-dimensional vector; a matrix calculation unit that performs a matrix calculation on the one-dimensional vector using the matrix set by the setting unit; a scaling unit that performs scaling on the one-dimensional vector on which the matrix operation has been performed; a matrix generator that converts the scaled one-dimensional vector into a matrix; An image processing device comprising: (38) a storage unit that stores the matrix candidates; The setting unit is configured to set the matrix by selecting the matrix corresponding to the content of the conversion process and the scanning method from among candidates of the matrix stored in the storage unit. (37) An image processing device according to (37). (39) The setting unit sets the matrix corresponding to the content of the conversion process and the scanning method adopted in the rasterizing unit and the matrix generating unit. The image processing device according to (37) or (38). (40) The setting unit sets the matrix based on a conversion identifier indicating the content of the conversion process and a scan identifier that is information about the scan method. An image processing device according to any one of (37) to (39). (41) the transform identifier is a secondary transform identifier indicating the content of a secondary transform for a primary transform coefficient obtained by performing a primary transform on the prediction residual, The setting unit is configured to set a matrix of a secondary transform for primary transform coefficients obtained by primary transforming the prediction residual, based on the secondary transform identifier and the scan identifier. An image processing device according to any one of (37) to (40). (42) The primary transform is an orthogonal transform. An image processing device according to any one of (37) to (41). (43) a primary conversion unit that performs the primary conversion An image processing device according to any one of (37) to (42). (44) The one-dimensional vector scaled by the matrixing unit is matrixed to obtain secondary transform coefficients. An image processing device according to any one of (37) to (43). (45) The image processing system further includes a coding unit that codes quantized transform coefficient levels obtained by quantizing the secondary transform coefficients by the quantization unit. An image processing device according to any one of (37) to (44). (46) A matrix of a transformation process for the transformation coefficients is set based on the content of the transformation process and a scanning method; a prediction residual, which is a difference between an image and a predicted image of the image, is subjected to a transformation process to obtain a transformation coefficient, which is then transformed into a one-dimensional vector; performing a matrix operation on the one-dimensional vector using the set matrix; Scaling is performed on the one-dimensional vector on which the matrix operation has been performed; Matrix the scaled one-dimensional vector Image processing methods. (47) a setting unit that sets a matrix for an inverse transform process for a transform coefficient based on the content of the inverse transform process and a scanning method; a rasterization unit that converts a transform coefficient obtained by performing an inverse transform process to obtain a prediction residual, which is a difference between an image and a predicted image of the image, into a one-dimensional vector; a matrix calculation unit that performs a matrix calculation on the one-dimensional vector using the matrix set by the setting unit; a scaling unit that performs scaling on the one-dimensional vector on which the matrix operation has been performed; a matrix generator that converts the scaled one-dimensional vector into a matrix; An image processing device comprising: (48) a storage unit that stores the matrix candidates; The setting unit is configured to set the matrix by selecting the matrix corresponding to the content of the inverse transformation process and the scanning method from among matrix candidates stored in the storage unit. (47) An image processing device according to (47). (49) The setting unit sets the matrix corresponding to the content of the inverse conversion process and the scanning method adopted in the rasterizing unit and the matrix generating unit. An image processing device according to (47) or (48). (50) The setting unit sets the matrix based on a transformation identifier indicating the content of the inverse transformation process and a scan identifier that is information about the scanning method. An image processing device according to any one of (47) to (49). (51) the transform identifier is a secondary transform identifier indicating the content of an inverse secondary transform for secondary transform coefficients obtained by decoding the coded data, the setting unit is configured to set a matrix of the inverse secondary transform based on the secondary transform identifier and the scan identifier; the rasterization unit is configured to transform the secondary transform coefficients into the one-dimensional vector; The matrixing unit is configured to matrix the scaled one-dimensional vector to obtain primary transform coefficients. An image processing device according to any one of (47) to (50). (52) an inverse primary transform unit that performs the inverse primary transform on the primary transform coefficients obtained by the matrix generation unit to obtain the prediction residuals; An image processing device according to any one of (47) to (51). (53) The inverse primary transform is an inverse orthogonal transform. An image processing device according to any one of (47) to (52). (54) further comprising an inverse quantization unit that inversely quantizes quantized transform coefficient levels obtained by decoding the coded data; The rasterization unit is configured to convert the secondary transform coefficients obtained by the inverse quantization unit inversely quantizing the quantized transform coefficient levels into the one-dimensional vectors. An image processing device according to any one of (47) to (53). (55) further comprising a decoding unit that decodes the encoded data; The inverse quantization unit is configured to inverse quantize the quantized transform coefficient levels obtained by decoding the coded data by the decoding unit. An image processing device according to any one of (47) to (54). (56) A matrix for an inverse transform process for the transform coefficients is set based on the content of the inverse transform process and a scanning method; converting the transform coefficients obtained by inverse transform processing into a one-dimensional vector, the transform coefficients being the difference between the image and a predicted image of the image; performing a matrix operation on the one-dimensional vector using the set matrix; Scaling is performed on the one-dimensional vector on which the matrix operation has been performed; Matrix the scaled one-dimensional vector Image processing methods. [Explanation of symbols]
[0570] 100 Image encoding device, 101 Control unit, 111 Arithmetic unit, 112 Transform unit, 113 Quantization unit, 114 Encoding unit, 115 Inverse quantization unit, 116 Inverse transform unit, 117 Arithmetic unit, 118 Frame memory, 119 Prediction unit, 131 Primary transform unit, 132 Secondary transform unit, 141 Rasterization unit, 142 Matrix operation unit, 143 Scaling unit, 144 Matrix generation unit, 145 Secondary transform selection unit, 146 Quantization unit, 147 Non-zero coefficient number determination unit, 148 Switch, 200 Image decoding device, 211 Decoding unit, 212 Inverse quantization unit, 213 Inverse transform unit, 214 Arithmetic unit, 215 Frame memory, 216 Prediction unit 231 inverse secondary transform unit, 232 inverse primary transform unit, 241 non-zero coefficient number determination unit, 242 switch, 243 rasterization unit, 244 matrix operation unit, 245 scaling unit, 246 matrix generation unit, 247 inverse secondary transform selection unit, 301 secondary transform selection unit, 311 secondary transform derivation unit, 312 secondary transform storage unit, 321 inverse secondary transform selection unit, 331 inverse secondary transform derivation unit, 332 inverse secondary transform storage unit, 421, 441 switch
Claims
1. a matrix operation unit that performs a matrix operation on a one-dimensional vector obtained by transforming transform coefficients, the one-dimensional vector being obtained by performing an inverse transform process using a matrix for an inverse transform process on the transform coefficients, the matrix being set based on the secondary transform identifier and the scanning method, and the prediction residual being a difference between an image and a predicted image of the image; a matrixing unit that matrix-generates the scaled one-dimensional vector with respect to the one-dimensional vector on which the matrix operation has been performed; An image processing device comprising:
2. performing a matrix operation on a one-dimensional vector obtained by transforming the transform coefficients, which is obtained by performing an inverse transform process using a matrix of an inverse transform process on the transform coefficients set based on the secondary transform identifier and the scanning method, to obtain a prediction residual, which is a difference between an image and a predicted image of the image, by performing an inverse transform process; matrixing the scaled one-dimensional vector with respect to the one-dimensional vector on which the matrix operation has been performed; An image processing method comprising:
3. performing a matrix operation on a one-dimensional vector obtained by transforming the transform coefficients, which is obtained by performing an inverse transform process using a matrix of an inverse transform process on the transform coefficients set based on the secondary transform identifier and the scanning method, to obtain a prediction residual, which is a difference between an image and a predicted image of the image, by performing an inverse transform process; matrixing the scaled one-dimensional vector with respect to the one-dimensional vector on which the matrix operation has been performed; A program for causing a computer to execute a process including the above.
4. performing a matrix operation on a one-dimensional vector obtained by transforming the transform coefficients, which is obtained by performing an inverse transform process using a matrix of an inverse transform process on the transform coefficients set based on the secondary transform identifier and the scanning method, to obtain a prediction residual, which is a difference between an image and a predicted image of the image, by performing an inverse transform process; matrixing the scaled one-dimensional vector with respect to the one-dimensional vector on which the matrix operation has been performed; A computer-readable recording medium on which a program for executing a process including the steps of:
Citation Information
Patent Citations
Arithmetic decoding device, arithmetic coding device, image decoding device and image coding device
WO2013190990A1
Systems and methods for low complexity forward transforms using zeroed-out coefficients
WO2015183375A2
Systems and methods for coding transform data
WO2017191782A1