Decoder, encoder, decoding method and encoding method

The decoding apparatus optimizes processing by selectively applying inverse transforms based on intra prediction modes, addressing the challenge of maintaining efficiency in video coding.

JP2025106454APending Publication Date: 2025-07-15PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025064042
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-06-01
Filing Date
2025-04-09
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in reducing processing load while maintaining compression efficiency.

Method used

A decoding apparatus that determines the use of intra prediction and its mode for blocks, performing a second inverse transform only when necessary, thereby optimizing processing by skipping it when the intra prediction mode is predetermined.

Benefits of technology

This approach reduces processing load while preserving compression efficiency by selectively applying inverse transforms based on intra prediction modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106454000001_ABST
    Figure 2025106454000001_ABST
Patent Text Reader

Abstract

To provide an encoder that can achieve processing load reduction while suppressing degradation of compression efficiency.SOLUTION: A decoder 200 includes a processor and a memory. Using the memory, the processor determines whether intra prediction is to be used for a target block to be decoded, and whether the target block to be decoded has a predetermined size, if it is determined that intra prediction is to be used for the target block to be decoded and the target block to be decoded has the predetermined size, further determines whether the intra prediction mode of the target block to be decoded is a predetermined mode, if the intra prediction mode is not the predetermined mode, performs second inverse transform on inverse quantization coefficients of the target block to be decoded, performs first inverse transform on the coefficients obtained by the second inverse transform, if the intra prediction mode is the predetermined mode, skips the second inverse transform and performs the first inverse transform on the inverse quantization coefficients of the target block to be decoded.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to encoding and decoding of images / videos in block units.

Background Art

[0002] A video coding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In such encoding and decoding technologies, it is required to reduce the processing load while suppressing a decrease in compression efficiency.

[0005] Therefore, the present disclosure provides an encoding device, a decoding device, an encoding method, or a decoding method that can reduce the processing load while suppressing a decrease in compression efficiency.

Means for Solving the Problems

[0006] A decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine whether to use intra prediction for a block to be decoded and whether the block to be decoded has a predetermined size. When it is determined to use intra prediction for the block to be decoded and it is determined that the block to be decoded has the predetermined size, the processor further determines whether an intra prediction mode of the block to be decoded is a predetermined mode. When the intra prediction mode is not the predetermined mode, a second inverse transformation is performed on the inverse quantization coefficients of the block to be decoded, and then a first inverse transformation is performed on the coefficients obtained by the second inverse transformation. When the intra prediction mode is the predetermined mode, the second inverse transformation is skipped, and a first inverse transformation is performed on the inverse quantization coefficients of the block to be decoded.

[0007] Note that these general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

Advantages of the Invention

[0008] The present disclosure can provide an encoding apparatus, a decoding apparatus, an encoding method, or a decoding method that can reduce a processing load while suppressing a decrease in compression efficiency.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be specifically described with reference to the drawings.

[0011] Note that all the embodiments described below show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. In addition, among the components in the following embodiments, the components not described in the independent claims indicating the most general concept are described as optional components.

[0012] (Embodiment 1) First, as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure to be described later are applicable, an overview of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.

[0013] When applying the processes and / or configurations described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.

[0014] (1) For the encoding device or decoding device of Embodiment 1, among the plurality of components constituting the encoding device or decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. (2) For the encoding device or decoding device of Embodiment 1, after performing any changes such as addition, replacement, deletion, etc. of functions or processes to be performed on some of the components constituting the encoding device or decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. For the method implemented by the encoding device or decoding device of Embodiment 1, after adding processing and / or making any changes such as replacement or deletion to some of the multiple processes included in the method, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure. (4) Implement a combination of some of the multiple components that make up the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components that include some of the functions of the components described in each aspect of the present disclosure, or components that implement some of the processes implemented by the components described in each aspect of the present disclosure. (5) Implement a combination of components that include some of the functions of some of the multiple components that make up the encoding device or decoding device of Embodiment 1, or components that implement some of the processes implemented by some of the multiple components that make up the encoding device or decoding device of Embodiment 1, with the components described in each aspect of the present disclosure, components that include some of the functions of the components described in each aspect of the present disclosure, or components that implement some of the processes implemented by the components described in each aspect of the present disclosure. (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure among the multiple processes included in the method. (7) Implement a combination of some of the multiple processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processes described in each aspect of the present disclosure.

[0015] Note that the ways of implementing the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Also, the processes and / or configurations described in different aspects may be implemented in combination.

[0016] [Overview of the Encoding Device] First, the overview of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of an encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.

[0017] As shown in FIG. 1, the encoding device 100 is a device that encodes images in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0018] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0019] Next, each component included in the encoding device 100 will be described.

[0020] [Division Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These blocks of fixed size are sometimes called Coding Tree Units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into variable-size blocks (e.g., 64x64 or smaller) based on recursive quadtree and / or binary tree block splitting. These variable-size blocks are sometimes called Coding Units (CUs), Prediction Units (PUs), or Transformation Units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be the processing units of CUs, PUs, and TUs.

[0021] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, solid lines represent block boundaries by quadtree block splitting, and dashed lines represent block boundaries by binary tree block splitting.

[0022] Here, block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).

[0023] The upper-left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper-left 64x64 block is split into two 16x64 blocks 11, 12 and a 32x64 block 13.

[0024] The upper-right 64x64 block is horizontally split into two rectangular 64x32 blocks 14, 15 (binary tree block splitting).

[0025] The 64x64 block in the lower left is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the 64x64 block in the lower left is divided into 16 16x32 blocks, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.

[0026] The 64x64 block 23 in the lower right is not divided.

[0027] As described above, in FIG. 2, the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0028] Note that in FIG. 2, although one block is divided into four or two blocks (quad-tree or binary-tree block division), the division is not limited to this. For example, one block may be divided into three blocks (ternary-tree block division). Such division including ternary-tree block division is sometimes called MBT (multi type tree) division.

[0029] [Subtraction unit] The subtraction unit 104 subtracts the prediction signal (predicted sample) from the original signal (original sample) in units of blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0030] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, a luminance signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0031] [Transformation unit] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0032] Note that the transformation unit 106 may adaptively select a transformation type from a plurality of transformation types and transform the prediction error into transformation coefficients using a transformation basis function corresponding to the selected transformation type. Such a transformation may be referred to as an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).

[0033] The plurality of transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the transformation basis functions corresponding to each transformation type. In FIG. 3, N indicates the number of input pixels. The selection of the transformation type from these plurality of transformation types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0034] Information indicating whether or not to apply such an EMT or AMT (for example, called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that the signaling of this information need not be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0035] Also, the conversion unit 106 may re-convert the conversion coefficient (conversion result). Such re-conversion is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (e.g., 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0036] Here, the separable transform is a method of performing multiple conversions by separating in each direction by the number of dimensions of the input, and the non-separable transform is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing the conversion collectively.

[0037] For example, as an example of the non-separable transform, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.

[0038] Also, similarly, after regarding a 4×4 input block as an array having 16 elements, a method of performing multiple Givens rotations on the array (Hypercube Givens Transform) is also an example of the non-separable transform.

[0039] [Quantization unit] The quantization unit 108 quantizes the transform coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the scanned transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0040] The predetermined order is the order for quantization / inverse quantization of the transform coefficients. For example, the predetermined scanning order is defined as ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0041] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0042] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable-length encoding on the quantization coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.

[0043] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantization coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse conversion unit 114.

[0044] [Inverse Conversion Unit] The inverse transform unit 114 restores the prediction error by inversely transforming the transform coefficients which are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients, thereby restoring the prediction error of the current block. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0045] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.

[0046] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error which is the input from the inverse transform unit 114 and the prediction sample which is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.

[0047] [Block memory] The block memory 118 is a storage unit for storing blocks within the coded target picture (hereinafter referred to as the current picture) which are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.

[0048] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116, and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0049] In ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 sub-block within a current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.

[0050] Specifically, first, sub-blocks (for example, 2x2 sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (for example, 15 or 25 classes).

[0051] The gradient direction value D is derived, for example, by comparing gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding gradients in a plurality of directions and quantizing the addition result.

[0052] Based on the results of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0053] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Note that the signaling of information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (for example, sequence level, slice level, tile level, CTU level, or CU level).

[0054] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance difference, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or CU level. Note that the signaling of the information indicating the on / off of ALF does not have to be limited to the picture level or CU level, and it may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0055] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets does not have to be limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0056] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0057] [Intra prediction unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to the blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0058] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0059] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).

[0060] The plurality of directional prediction modes includes, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0061] In addition, in the intra prediction of the chrominance blocks, the luminance blocks may be referred to. That is, based on the luminance component of the current block, the chrominance components of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of the chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0062] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction. Such intra prediction with such correction may be called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called a PDPC flag) is signaled at, for example, the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0063] [Inter prediction unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0064] The motion information used for motion compensation is signaled. A motion vector predictor may be used for the signaling of the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.

[0065] In addition to the motion information of the current block obtained by motion search, the motion information of adjacent blocks may also be used to generate an inter prediction signal. Specifically, an inter prediction signal may be generated for each sub-block within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).

[0066] In such an OBMC mode, information indicating the size of the sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be at other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).

[0067] The OBMC mode will be described more specifically. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process by OBMC processing.

[0068] First, a prediction image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.

[0069] Next, the motion vector (MV_L) of the coded left adjacent block is applied to the block to be coded to obtain a prediction image (Pred_L), and the first correction of the prediction image is performed by weighting and superimposing the prediction image and Pred_L.

[0070] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_U). The predicted image after the first correction and Pred_U are weighted and superimposed to perform the second correction on the predicted image, which is used as the final predicted image.

[0071] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to configure to perform corrections more times than two stages using the right adjacent block or the lower adjacent block.

[0072] Note that the region for superimposition may be only a partial region near the block boundary instead of the pixel region of the entire block.

[0073] Here, the predicted image correction process from a single reference picture has been described. However, the same applies to the case of correcting the predicted image from a plurality of reference pictures. After obtaining the predicted images corrected from each reference picture, the obtained predicted images are further superimposed to obtain the final predicted image.

[0074] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.

[0075] As a method for determining whether to apply OBMC processing, for example, there is a method using an obmc_flag which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, a value of 1 is set as the obmc_flag and encoding is performed by applying OBMC processing. If it does not belong to a region with complex motion, a value of 0 is set as the obmc_flag and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC processing according to the value.

[0076] Note that the motion information may be derived on the decoder side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also for example, the motion information may be derived by performing motion search on the decoder side. In this case, the motion search is performed without using the pixel values of the current block.

[0077] Here, a mode in which motion search is performed on the decoder side will be described. This mode in which motion search is performed on the decoder side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.

[0078] An example of FRUC processing is shown in FIG. 5D. First, with reference to the motion vectors of encoded blocks that are spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value for each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0079] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, a motion vector for the current block may be derived. That is, search is performed in the same manner for the region around the best candidate MV, and if there is an MV for which the evaluation value is a good value, the best candidate MV may be updated to the MV, and that may be used as the final MV of the current block. Note that a configuration in which this process is not performed is also possible.

[0080] When performing processing in sub-block units, the same processing may be used as well.

[0081] Note that the evaluation value is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that in addition to the difference value, other information may be used to calculate the evaluation value.

[0082] As the pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.

[0083] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the candidate described above, the region in another reference picture along the motion trajectory of the current block is used.

[0084] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is advisable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.

[0085] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0086] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidates.

[0087] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that most closely matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or either one of the left and upper adjacent blocks and the reconstructed image at the equivalent position within the encoded reference picture (Ref0) specified by the candidate MV is derived, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.

[0088] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0089] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called the BIO (bi-directional optical flow) mode.

[0090] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (vx, vy) indicates a velocity vector, and τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.

[0091] At this time, under the assumption of uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (1) holds.

[0092]

Equation

[0093] Here, I(k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from the merge list or the like are corrected in pixel units.

[0094] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0095] Here, a mode of deriving a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.

[0096] FIG. 9A is a diagram for explaining the derivation of a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, a current block includes 16 4x4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (vx, vy) of each sub-block within the current block is derived by the following equation (2).

[0097]

Equation

[0098] Here, x and y indicate the horizontal position and vertical position of the sub-block, respectively, and w indicates a predetermined weight coefficient.

[0099] Such an affine motion compensation prediction mode may include several modes with different methods for deriving the motion vectors of the upper left and upper right control points. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode does not necessarily have to be limited to the CU level, and it may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0100] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0101] Here, an example of deriving the motion vector of the picture to be coded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.

[0102] First, a prediction MV list in which candidates for the prediction MV are registered is generated. Examples of candidates for the prediction MV include a spatial adjacent prediction MV that is the MV of a plurality of coded blocks located spatially adjacent to the block to be coded, a temporal adjacent prediction MV that is the MV of a nearby block obtained by projecting the position of the block to be coded in the coded reference picture, a combined prediction MV that is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV that is an MV with a value of zero.

[0103] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be coded.

[0104] Furthermore, in the variable length coding unit, a merge_idx, which is a signal indicating which prediction MV has been selected, is described in the stream and coded.

[0105] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are just an example, and the number may be different from that in the figure, or the configuration may not include some of the types of prediction MVs in the figure, or prediction MVs other than the types of prediction MVs in the figure may be added.

[0106] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be coded derived in the merge mode.

[0107] Here, an example of determining the MV using the DMVR process will be described.

[0108] FIG. 9C is a conceptual diagram for explaining the outline of the DMVR process.

[0109] First, using the optimal MVP set for the processing target block as a candidate MV, according to the candidate MV, reference pixels are respectively obtained from the first reference picture which is the processed picture in the L0 direction and the second reference picture which is the processed picture in the L1 direction, and a template is generated by taking the average of each reference pixel.

[0110] Next, using the template, the peripheral areas of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, the MV value, etc.

[0111] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.

[0112] Note that even if it is not the process itself described here, other processes may be used as long as it is a process capable of searching the periphery of the candidate MV to derive the final MV.

[0113] Here, the mode of generating a predicted image using the LIC process will be described.

[0114] FIG. 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction process by the LIC process.

[0115] First, an MV for obtaining a reference image corresponding to the block to be encoded is derived from the reference picture which is the encoded picture.

[0116] Next, for the block to be encoded, information indicating how the luminance values change between the reference picture and the picture to be encoded is extracted using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, and the luminance correction parameter is calculated.

[0117] The luminance correction process is performed on the reference image in the reference picture specified by the MV using the luminance correction parameter to generate a predicted image for the block to be encoded.

[0118] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0119] Also, although the process of generating a predicted image from a single reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures. The luminance correction process is performed on the reference images obtained from each reference picture in the same way and then the predicted image is generated.

[0120] As a method for determining whether to apply the LIC process, for example, there is a method that uses a lic_flag, which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and the LIC process is applied for encoding. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply the LIC process according to the value.

[0121] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the peripheral encoded blocks selected during the derivation of the MV in the merge mode process are encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. In the case of this example, the processing in decoding is exactly the same.

[0122] [Overview of Decoder] Next, an overview of a decoder capable of decoding the encoded signal (encoded bitstream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing the functional configuration of a decoder 200 according to Embodiment 1. The decoder 200 is a moving image / image decoder that decodes moving images / images in units of blocks.

[0123] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0124] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0125] Each component included in the decoder 200 will be described below.

[0126] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202 arithmetic-decodes the encoded bit stream into a binary signal, for example. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients in block units to the inverse quantization unit 204.

[0127] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0128] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients, which is the input from the inverse quantization unit 204.

[0129] For example, when the information decoded from the encoded bit stream indicates that EMT or AMT is to be applied (for example, the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.

[0130] Also, for example, when the information decoded from the encoded bit stream indicates that NSST is to be applied, the inverse transform unit 206 applies an inverse reverse transform to the transform coefficients.

[0131] [Addition Unit] The adder 208 reconstructs the current block by adding the prediction error, which is the input from the inverse conversion unit 206, and the prediction sample, which is the input from the predictive control unit 220. Then, the adder 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0132] [Block Memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture), which are blocks referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0133] [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device, etc.

[0134] When the information indicating the on / off of the ALF read from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0135] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0136] [Intra Prediction Unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bitstream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0137] In addition, when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0138] Also, when the information decoded from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0139] [Inter Prediction Unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bitstream, and outputs the inter prediction signal to the prediction control unit 220.

[0140] In addition, when the information decoded from the encoded bitstream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.

[0141] In addition, when the information decoded from the encoded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0142] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0143] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the addition unit 208 as the prediction signal.

[0144] (Embodiment 2) Next, Embodiment 2 will be described. In this embodiment, the conversion and inverse conversion will be described in detail. Note that since the configurations of the encoding apparatus and the decoding apparatus according to this embodiment are substantially the same as those of Embodiment 1, illustration and description thereof are omitted.

[0145] [Processing of the conversion unit and quantization unit of the encoding apparatus] First, the processing of the conversion unit 106 and the quantization unit 108 of the encoding apparatus 100 according to this embodiment will be specifically described with reference to FIG. 11. FIG. 11 is a flowchart showing the conversion and quantization processing in the encoding apparatus 100 according to Embodiment 2.

[0146] First, the conversion unit 106 determines whether to use intra prediction or inter prediction for the block to be encoded (S101). For example, the conversion unit 106 determines whether to use intra prediction or inter prediction based on the difference between the original image and the reconstructed image obtained by locally decoding the compressed image and / or the cost based on the amount of code. Also, for example, the conversion unit 106 may determine whether to use intra prediction or inter prediction based on information different from the cost based on the difference and / or the amount of code (e.g., picture type).

[0147] Here, when it is determined to use inter prediction for the block to be encoded (inter in S101), the conversion unit 106 selects a first conversion basis for the block to be encoded from among one or more candidates for the first conversion basis (S102). For example, the conversion unit 106 fixedly selects the conversion basis of DCT-II as the first conversion basis for the block to be encoded. Also, for example, the conversion unit 106 may select the first conversion basis from among a plurality of candidates for the first conversion basis.

[0148] Then, the conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S102 (S103). The quantization unit 108 quantizes the generated first conversion coefficients (S110) and ends the conversion and quantization processing.

[0149] On the other hand, when it is determined to use intra prediction for the block to be coded (intra in S101), the conversion unit 106 selects a first conversion basis for the block to be coded from among one or more candidates for the first conversion basis (S104). For example, the conversion unit 106 can select the first conversion basis using an adaptive basis selection mode. The adaptive basis selection mode is a mode in which the conversion basis is adaptively selected from among a plurality of predetermined candidates for the conversion basis based on the cost based on the difference between the original image and the reconstructed image and / or the amount of coding. This adaptive basis selection mode may also be referred to as the EMT mode or the AMT mode. As the plurality of candidates for the conversion basis, for example, the plurality of conversion bases shown in FIG. 6 can be used. Note that the plurality of candidates for the conversion basis is not limited to the plurality of conversion bases in FIG. 6. The plurality of candidates for the conversion basis may include, for example, a conversion basis equivalent to not performing conversion.

[0150] Also, for example, the conversion unit 106 may select the first conversion basis using a non-adaptive basis selection mode (that is, without using the adaptive basis selection mode). In the non-adaptive basis selection mode, for example, the conversion unit 106 can select the first conversion basis based on coding parameters (such as block size, quantization parameter, intra prediction mode, etc.). Further, the conversion unit 106 can also fixedly select a conversion basis (such as the conversion basis of DCT-II) defined in advance in a standard specification or the like. In this case, the selection of the conversion basis means fixedly adopting one predefined conversion basis. Also, the conversion unit 106 may adaptively switch between the adaptive basis conversion mode and the non-adaptive basis selection mode.

[0151] Subsequently, the conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S104 (S105). The conversion unit 106 determines whether the intra prediction mode of the block to be encoded is a predetermined mode (S106). For example, the conversion unit 106 determines whether the intra prediction mode is a predetermined mode based on the cost based on the difference between the original image and the reconstructed image and / or the amount of code. Note that the determination of whether the intra prediction mode is a predetermined mode may be performed based on information different from the cost.

[0152] The predetermined mode may be defined in advance by, for example, a standard specification or the like. Also, for example, the predetermined mode may be determined based on encoding parameters or the like.

[0153] When the intra prediction mode is a predetermined mode (YES in S106), the conversion unit 106 determines whether the first conversion basis selected in step S104 matches the predetermined conversion basis (S107). The predetermined conversion basis may be defined in advance by, for example, a standard specification or the like. Also, for example, the predetermined conversion basis may be determined based on encoding parameters or the like.

[0154] When the intra prediction mode is not a predetermined mode (NO in S106), or when the first conversion basis matches the predetermined conversion basis (YES in S107), the conversion unit 106 selects a second conversion basis for the block to be encoded from among one or more second conversion basis candidates (S108). The conversion unit 106 generates second conversion coefficients by performing a second conversion on the first conversion coefficients using the selected second conversion basis (S109). The quantization unit 108 quantizes the generated second conversion coefficients (S110), and ends the conversion and quantization processing.

[0155] In the second conversion, a secondary conversion called NSST may be performed, or a conversion that selectively uses any of a plurality of candidates for the second conversion basis may be performed. At this time, in the selection of the second conversion basis, the selected conversion basis may be fixed. That is, a predetermined fixed conversion basis may be selected as the second conversion basis. Also, as the second conversion basis, a conversion basis equivalent to not performing the second conversion may be used.

[0156] When the intra prediction mode is a predetermined mode (YES in S106) and the first conversion basis is different from the predetermined conversion basis (NO in S107), the conversion unit 106 skips the second conversion basis selection step (S108) and the second conversion step (S109). That is, the conversion unit 106 does not perform the second conversion. In this case, the first conversion coefficients generated in step S105 are quantized (S110), and the conversion and quantization processes are completed.

[0157] When the second conversion step is skipped in this way, information indicating that the second conversion is not performed may be notified to the decoding apparatus. Also, when the second conversion step is skipped, the second conversion may be performed using a second conversion basis equivalent to not performing the conversion, and information indicating the second conversion basis may be notified to the decoding apparatus.

[0158] Note that the inverse quantization unit 112 and the inverse conversion unit 114 of the encoding apparatus 100 can reconstruct the encoding target block by performing processes reverse to those of the conversion unit 106 and the quantization unit 108.

[0159] [Processes of the inverse quantization unit and the inverse conversion unit of the decoding apparatus] Next, the processes of the inverse quantization unit 204 and the inverse conversion unit 206 of the decoding apparatus 200 according to the present embodiment will be specifically described with reference to FIG. 12. FIG. 12 is a flowchart showing the inverse quantization and inverse conversion processes in the decoding apparatus 200 according to Embodiment 2.

[0160] First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S501). The inverse transform unit 206 determines whether to use intra prediction or inter prediction for the block to be decoded (S502). For example, the inverse transform unit 206 determines whether to use intra prediction or inter prediction based on the information obtained from the bit stream.

[0161] When it is determined to use inter prediction for the block to be decoded (Inter in S502), the inverse transform unit 206 selects a first inverse transform basis for the block to be decoded (S503). Selecting an inverse transform basis (the first inverse transform basis or the second inverse transform basis) in the decoding apparatus 200 means determining the inverse transform basis based on predetermined information. As the predetermined information, for example, a basis selection signal can be used. Also, as the predetermined information, an intra prediction mode or a block size or the like can be used.

[0162] The inverse transform unit 206 performs a first inverse transform on the inverse quantized coefficients of the block to be decoded using the first inverse transform basis selected in step S503 (S504), and ends the inverse quantization and inverse transform processing.

[0163] When it is determined to use intra prediction for the block to be decoded (Intra in S502), the inverse transform unit 206 determines whether the intra prediction mode of the block to be decoded is a predetermined mode (S505). The predetermined mode used in the decoding apparatus 200 is the same as the predetermined mode used in the encoding apparatus 100.

[0164] When the intra prediction mode is the predetermined mode (YES in S505), the inverse transform unit 206 determines whether the first inverse transform basis matches the predetermined inverse transform basis (S506). As the predetermined inverse transform basis, the inverse transform basis corresponding to the predetermined transform basis used in the encoding apparatus 100 is used.

[0165] When the intra prediction mode is not a predetermined mode (NO in S505), or when the first inverse transform basis matches the predetermined inverse transform basis (YES in S506), the inverse transform unit 206 selects a second inverse transform basis for the block to be decoded (S507). The inverse transform unit 206 performs a second inverse transform on the inverse quantized coefficients of the block to be decoded using the selected second inverse transform basis (S508). The inverse transform unit 206 selects the first inverse transform basis (S509). The inverse transform unit 206 performs a first inverse transform on the coefficients obtained by the second inverse transform in step S508 using the selected first inverse transform basis (S510), and ends the inverse quantization and inverse transform processing.

[0166] On the other hand, when the intra prediction mode is a predetermined mode (YES in S505) and the first inverse transform basis is different from the predetermined inverse transform basis (NO in S506), the inverse transform unit 206 skips the second inverse transform basis selection step (S507) and the second inverse transform step (S508). That is, the inverse transform unit 206 selects the first inverse transform basis without performing the second inverse transform (S509). The inverse transform unit 206 performs a first inverse transform on the coefficients inverse quantized in step S501 using the selected first inverse transform basis (S510), and ends the inverse quantization and inverse transform processing.

[0167] [Effects, etc.] The inventors have found that in conventional coding, there is a problem that the amount of processing for searching for an optimal combination of a transform basis and transform parameters (for example, filter coefficients) in both the first transform and the second transform is enormous. In contrast, according to the coding apparatus 100 and the decoding apparatus 200 according to the present embodiment, the second transform can be skipped according to the intra prediction mode and the first transform basis. As a result, it is possible to reduce the processing for searching for an optimal combination of the transform basis and the transform parameters in both the first transform and the second transform, and it is possible to reduce the processing load while suppressing a decrease in compression efficiency.

[0168] In addition, in the present embodiment, although the second transformation is not performed when inter prediction is used for an encoding target block, the present invention is not limited to this. That is, when inter prediction is used for an encoding target block, the second transformation may be performed on the first transformation coefficients generated by the first transformation. In this case, the second transformation coefficients generated by the second transformation are quantized.

[0169] Note that the order of steps in the flowcharts of FIGS. 11 and 12 is not limited to the order described in FIGS. 11 and 12. For example, in FIG. 11, the determination step (S106) as to whether the intra prediction mode is a predetermined mode and the determination step (S107) as to whether the first transformation basis matches the predetermined transformation basis may be in the reverse order or may be performed simultaneously.

[0170] Note that this aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0171] (Embodiment 3) Next, Embodiment 3 will be described. In the present embodiment, the difference from the above-described Embodiment 2 is that the predetermined mode used for determining the intra prediction mode is limited to the non-directional prediction mode. Hereinafter, the present embodiment will be described with reference to the drawings, centering on the differences from the above-described Embodiment 2. In the following respective drawings, steps substantially the same as those in Embodiment 2 are denoted by the same reference numerals, and overlapping descriptions are omitted or simplified.

[0172] [Processing of the transformation unit and quantization unit of the encoding apparatus] First, the processing of the transformation unit 106 and quantization unit 108 of the encoding apparatus 100 according to the present embodiment will be specifically described with reference to FIG. 13. FIG. 13 is a flowchart showing the transformation and quantization processing in the encoding apparatus 100 according to Embodiment 3.

[0173] First, the conversion unit 106 determines whether to use intra prediction or inter prediction for the block to be encoded (S101). Here, when it is determined to use inter prediction for the block to be encoded (Inter in S101), the conversion unit 106 selects a first conversion basis (S102), and generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the selected first conversion basis (S103). The quantization unit 108 quantizes the generated first conversion coefficients (S110), and ends the conversion and quantization processes.

[0174] On the other hand, when it is determined to use intra prediction for the block to be encoded (Intra in S101), the conversion unit 106 determines whether the intra prediction mode of the block to be encoded is a non-directional prediction mode (S201). The non-directional prediction mode is a mode that does not use a specific direction for predicting the block to be decoded. Specifically, the non-directional prediction mode is, for example, the DC prediction mode and / or the Planar prediction mode. In the non-directional prediction mode, for example, the pixel value is predicted using the average value of the reference pixels or the interpolated value of the reference pixels. Conversely, a mode that uses a specific direction for predicting the block to be decoded is called a directional prediction mode. In the directional prediction mode, the pixel value is predicted by extending the values of the reference pixels in a specific direction. Note that the pixel value is the value of each pixel constituting the picture, and is, for example, a luminance value or a color difference value.

[0175] Here, when the intra prediction mode is different from the non-directional prediction mode (NO in S201), the conversion unit 106 selects a first conversion basis for the block to be encoded (S202). The conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S202 (S203). Further, the conversion unit 106 selects a second conversion basis for the block to be encoded (S204). The conversion unit 106 generates second conversion coefficients by performing a second conversion on the first conversion coefficients generated in step S203 using the second conversion basis selected in step S204 (S205). The processing of steps S202 to S205 is substantially the same as the processing of steps S104 to S109 when step S106 is NO in FIG. 11. Thereafter, the quantization unit 108 quantizes the second conversion coefficients generated in step S205 (S110), and ends the conversion and quantization processing.

[0176] On the other hand, when the intra prediction mode matches the non-directional prediction mode (YES in S201), the conversion unit 106 selects a first conversion basis for the block to be encoded (S206). The conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S206 (S207). The conversion unit 106 determines whether or not the first conversion basis selected in step S206 matches a predetermined conversion basis (S208). As the predetermined conversion basis, for example, a conversion basis of DCT-II and / or a conversion basis similar thereto can be used.

[0177] Here, when the first conversion basis matches the predetermined conversion basis (YES in S208), the conversion unit 106 selects a second conversion basis for the block to be encoded (S209). Then, the conversion unit 106 performs a second conversion on the first conversion coefficients generated in step S207 using the second conversion basis selected in step S209, thereby generating second conversion coefficients (S210). Thereafter, the quantization unit 108 quantizes the second conversion coefficients generated in step S210 (S110), and ends the conversion and quantization processes.

[0178] On the other hand, when the first conversion basis is different from the predetermined conversion basis (NO in S208), the conversion unit 106 skips the selection step (S209) and the second conversion step (S210) of the second conversion basis. That is, the conversion unit 106 does not perform the second conversion. In this case, the first conversion coefficients generated in step S207 are quantized (S110), and the conversion and quantization processes are ended.

[0179] The processes of steps S206 to S209 are substantially the same as the processes of steps S104 to S109 when step S106 is YES in FIG. 11.

[0180] [Processing of the Inverse Quantization Unit and Inverse Conversion Unit of the Decoder] Next, the processing of the inverse quantization unit 204 and the inverse conversion unit 206 of the decoder 200 according to the present embodiment will be specifically described with reference to FIG. 14. FIG. 14 is a flowchart showing the inverse quantization and inverse conversion processes in the decoder 200 according to Embodiment 3.

[0181] First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S501). The inverse conversion unit 206 determines whether to use intra prediction or inter prediction for the block to be decoded (S502).

[0182] When it is determined to use inter prediction for the block to be decoded (inter in S502), the inverse transform unit 206 selects a first inverse transform basis for the block to be decoded (S503). The inverse transform unit 206 performs a first inverse transform on the inverse quantized coefficients of the block to be decoded using the first inverse transform basis selected in step S503 (S504), and ends the inverse quantization and inverse transform processing.

[0183] On the other hand, when it is determined to use intra prediction for the block to be decoded (intra in S502), the inverse transform unit 206 determines whether the intra prediction mode of the block to be decoded is a non-directional prediction mode (S601).

[0184] Here, when the intra prediction mode is not the non-directional prediction mode (NO in S601), the inverse transform unit 206 selects a second inverse transform basis for the block to be decoded (S602). The inverse transform unit 206 performs a second inverse transform on the inverse quantized coefficients of the block to be decoded using the selected second inverse transform basis (S603). The inverse transform unit 206 selects a first inverse transform basis (S604). The inverse transform unit 206 performs a first inverse transform on the coefficients obtained by the second inverse transform in step S603 using the selected first inverse transform basis (S605), and ends the inverse quantization and inverse transform processing.

[0185] On the other hand, when the intra prediction mode is the non-directional prediction mode (YES in S601), the inverse transform unit 206 determines whether the first inverse transform basis matches a predetermined inverse transform basis (S606). The predetermined inverse transform basis used in the decoding apparatus 200 is an inverse transform basis corresponding to the predetermined transform basis used in the encoding apparatus 100.

[0186] Here, when the first inverse transformation basis matches the predetermined inverse transformation basis (YES in S606), the inverse transformation unit 206 selects a second inverse transformation basis for the block to be decoded (S607). The inverse transformation unit 206 performs a second inverse transformation on the inverse quantized coefficients of the block to be decoded using the selected second inverse transformation basis (S608). The inverse transformation unit 206 selects the first inverse transformation basis (S609). The inverse transformation unit 206 performs a first inverse transformation on the coefficients obtained by the second inverse transformation in step S608 using the selected first inverse transformation basis (S610), and ends the inverse quantization and inverse transformation processing.

[0187] On the other hand, when the first inverse transformation basis is different from the predetermined inverse transformation basis (NO in S606), the inverse transformation unit 206 skips the selection step (S607) and the second inverse transformation step (S608) of the second inverse transformation basis. That is, the inverse transformation unit 206 selects the first inverse transformation basis without performing the second inverse transformation (S609). The inverse transformation unit 206 performs a first inverse transformation on the coefficients inverse quantized in step S501 using the selected first inverse transformation basis (S610), and ends the inverse quantization and inverse transformation processing.

[0188] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to the present embodiment, the second transformation can be skipped when the intra prediction mode is a non-directional prediction mode. In the non-directional prediction mode, the residual is often flat within the block. Therefore, if a transformation basis other than the DCT-II transformation basis and a transformation basis similar thereto is used, high-frequency components are likely to remain, and the distribution of the transformation coefficients is likely to become random. In this case, since the effect of improving the compression efficiency by the second transformation is reduced, it is possible to suppress a decrease in the compression efficiency and reduce the processing load by skipping the second transformation.

[0189] Note that, in the present embodiment, although the second transformation is not performed when inter prediction is used for an encoding target block, the present invention is not limited to this. That is, when inter prediction is used for an encoding target block, the second transformation may be performed on the first transformation coefficients generated by the first transformation. In this case, the second transformation coefficients generated by the second transformation are quantized.

[0190] Note that the order of the steps in the flowcharts of FIGS. 13 and 14 is not limited to the order described in FIGS. 13 and 14.

[0191] Note that this aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0192] (Embodiment 4) Next, Embodiment 4 will be described. In the present embodiment, the point that the first transformation basis is fixed according to the block size in the adaptive basis selection mode is different from Embodiment 2 described above. Hereinafter, the present embodiment will be described with reference to the drawings, centering on the points different from Embodiments 2 and 3 described above. In each of the following figures, steps substantially the same as those in Embodiments 2 and 3 are denoted by the same reference numerals, and overlapping descriptions are omitted or simplified.

[0193] [Processing of the Transformation Unit and Quantization Unit of the Encoding Apparatus] First, the processing of the transformation unit 106 and quantization unit 108 of the encoding apparatus 100 according to the present embodiment will be specifically described with reference to FIG. 15. FIG. 15 is a flowchart showing the transformation and quantization processing in the encoding apparatus 100 according to Embodiment 4.

[0194] First, the conversion unit 106 determines whether to use intra prediction or inter prediction for the block to be encoded (S101). Here, when it is determined to use inter prediction for the block to be encoded (Inter in S101), the conversion unit 106 selects the first conversion basis (S102), and generates the first conversion coefficients by performing the first conversion on the residual of the block to be encoded using the selected first conversion basis (S103). The quantization unit 108 quantizes the generated first conversion coefficients (S110) and ends the conversion and quantization processes.

[0195] On the other hand, when it is determined to use intra prediction for the block to be encoded (Intra in S101), the conversion unit 106 determines whether the size of the block to be encoded matches a predetermined size and whether to use the adaptive basis selection mode for the block to be encoded (S301). Whether to use the adaptive basis selection mode can be determined based on, for example, the difference between the original image and the reconstructed image and / or the cost based on the amount of code.

[0196] As the predetermined size, for example, a specific block size defined in advance by a standard specification or the like can be used. Specifically, as the predetermined size, for example, 4x4 pixels can be used. Also, as the predetermined size, a plurality of block sizes may be used. Specifically, as the predetermined size, for example, 4x4 pixels, 8x4 pixels, and 4x8 pixels may be used. Also, whether the size of the block to be encoded matches the predetermined size may be determined by determining whether the size of the block to be encoded satisfies a predetermined condition. In this case, as the predetermined condition, for example, a condition that both the horizontal size and the vertical size are less than or equal to a predetermined number of pixels, or at least one of the horizontal size and the vertical size is less than or equal to a predetermined number of pixels can be used.

[0197] When the size of the block to be encoded is different from the predetermined size or when the adaptive basis selection mode is not used (NO in S301), the conversion unit 106 determines whether the intra prediction mode of the block to be encoded is a non-directional prediction mode (S201).

[0198] Here, when the intra prediction mode is different from the non-directional prediction mode (NO in S201), the conversion unit 106 selects a first conversion basis for the block to be encoded (S202). For example, when it is determined that the adaptive basis selection mode is used, the conversion unit 106 adaptively selects the first conversion basis from among a plurality of candidates for the first conversion basis. Also, for example, when it is determined that the adaptive basis selection mode is not used, the conversion unit 106 fixedly selects a predefined conversion basis (for example, the conversion basis of DCT-II).

[0199] The conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S202 (S203). Further, the conversion unit 106 selects a second conversion basis for the block to be encoded (S204). The conversion unit 106 generates second conversion coefficients by performing a second conversion on the first conversion coefficients generated in step S203 using the second conversion basis selected in step S204 (S205). Thereafter, the quantization unit 108 quantizes the second conversion coefficients generated in step S205 (S110), and ends the conversion and quantization processing.

[0200] On the other hand, when the intra prediction mode matches the non-directional prediction mode (YES in S201), the conversion unit 106 selects a first conversion basis for the block to be encoded (S206). For example, when it is determined that the adaptive basis selection mode is used, the conversion unit 106 adaptively selects the first conversion basis from among a plurality of candidates for the first conversion basis. Also, for example, when it is determined that the adaptive basis selection mode is not used, the conversion unit 106 fixedly selects a predefined conversion basis (for example, the conversion basis of DCT-II).

[0201] The conversion unit 106 generates first conversion coefficients (S207) by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S206. The conversion unit 106 determines whether the first conversion basis selected in step S206 matches the second predetermined conversion basis (S208). As the second predetermined conversion basis, for example, a conversion basis of DCT-II and / or a conversion basis similar thereto can be used.

[0202] Here, when the first conversion basis matches the second predetermined conversion basis (YES in S208), the conversion unit 106 selects a second conversion basis for the block to be encoded (S209). Then, the conversion unit 106 generates second conversion coefficients (S210) by performing a second conversion on the first conversion coefficients generated in step S207 using the second conversion basis selected in step S209. Thereafter, the quantization unit 108 quantizes the second conversion coefficients generated in step S210 (S110), and ends the conversion and quantization processing.

[0203] On the other hand, when the first conversion basis is different from the second predetermined conversion basis (NO in S208), the conversion unit 106 skips the second conversion basis selection step (S209) and the second conversion step (S210). That is, the conversion unit 106 does not perform the second conversion. In this case, the first conversion coefficients generated in step S207 are quantized (S110), and the conversion and quantization processing ends.

[0204] When the size of the block to be encoded matches the predetermined size and the adaptive basis selection mode is used (YES in S301), the conversion unit 106 fixes the first conversion basis to the first predetermined conversion basis (S302). As the first predetermined conversion basis, for example, a conversion basis of DST-VII can be used. Note that the first predetermined conversion basis is not limited to the conversion basis of DST-VII. For example, a conversion basis of DCT-V may be used as the first predetermined conversion basis.

[0205] The conversion unit 106 generates first conversion coefficients (S303) by performing a first conversion on the residual of the block to be encoded using the first conversion basis fixed in step S302. The conversion unit 106 determines whether the intra prediction mode of the block to be encoded is a non-directional prediction mode (S304).

[0206] Here, when the intra prediction mode is different from the non-directional prediction mode (NO in S304), the conversion unit 106 selects a second conversion basis (S305). Then, the conversion unit 106 generates second conversion coefficients (S306) by performing a second conversion on the first conversion coefficients generated in step 303 using the second conversion basis selected in step S305. Thereafter, the quantization unit 108 quantizes the second conversion coefficients generated in step S306 (S110), and ends the conversion and quantization processes.

[0207] On the other hand, when the intra prediction mode matches the non-directional prediction mode (YES in S304), the conversion unit 106 skips the second conversion basis selection step (S305) and the second conversion step (S306). That is, the conversion unit 106 does not perform the second conversion. In this case, the first conversion coefficients generated in step S303 are quantized (S110), and the conversion and quantization processes are ended.

[0208] [Processing of the Entropy Encoding Unit of the Encoding Device] Next, the encoding process related to the conversion of the entropy encoding unit 110 of the encoding device 100 according to the present embodiment will be specifically described with reference to FIG. 16. FIG. 16 is a flowchart showing the encoding process in the encoding device 100 according to Embodiment 4.

[0209] When inter prediction is used for the block to be encoded (inter in S401), the entropy encoding unit 110 encodes the first basis selection signal into the bit stream (S402). Here, the first basis selection signal is information or data indicating the first conversion basis selected in step S102.

[0210] Encoding a signal in a bitstream means placing a code indicating information within the bitstream. The code is generated, for example, by context-adaptive binary arithmetic coding (CABAC). Note that CABAC is not necessarily used for code generation, nor is entropy coding necessarily used. For example, the code may be the information itself (e.g., a flag of 0 or 1).

[0211] Next, the entropy encoding unit 110 encodes the coefficients quantized in step S110 (S403) and ends the encoding process.

[0212] When intra prediction is used for the block to be encoded (intra in S401), the entropy encoding unit 110 encodes an intra prediction mode signal indicating the intra prediction mode of the block to be encoded in the bitstream (S404). Further, the entropy encoding unit 110 encodes an adaptive selection mode signal indicating whether the adaptive basis selection mode is used for the block to be encoded in the bitstream (S405).

[0213] Here, when the adaptive basis selection mode is used and the size of the block to be encoded is different from the predetermined size (YES in S406), the entropy encoding unit 110 encodes a first basis selection signal in the bitstream (S407). Here, the first basis selection signal is information or data indicating the first transform basis selected in step S202 or S206. On the other hand, when the adaptive basis selection mode is not used, or when the adaptive basis selection mode is used and the size of the block to be encoded matches the predetermined size (NO in S406), the entropy encoding unit 110 skips the encoding step of the first basis selection signal (S407). That is, the entropy encoding unit 110 does not encode the first basis selection signal.

[0214] Here, when the second transformation is performed (YES in S408), the entropy encoding unit 110 encodes the second basis selection signal into the bit stream (S409). Here, the second basis selection signal is information or data indicating the second transformation basis selected in step S204, S209, or S305. On the other hand, when the second transformation is not performed (NO in S408), the entropy encoding unit 110 skips the encoding step (S409) of the second basis selection signal. That is, the entropy encoding unit 110 does not encode the second basis selection signal.

[0215] Finally, the entropy encoding unit 110 encodes the coefficients quantized in step S110 (S410) and ends the encoding process.

[0216] [Processing of the Entropy Decoding Unit of the Decoder] Next, the processing of the entropy decoding unit 202 of the decoder 200 according to the present embodiment will be specifically described with reference to FIG. 17. FIG. 17 is a flowchart showing the decoding process in the decoder 200 according to Embodiment 4.

[0217] When inter prediction is used for the decoding target block (inter in S701), the entropy decoding unit 202 decodes the first basis selection signal from the bit stream (S702).

[0218] Decoding a signal from a bit stream means reading a code indicating information from the bit stream and restoring the information from the read code. For example, context-adaptive binary arithmetic decoding (CABAD) is used for restoring information from the code. Note that CABAD is not necessarily used for restoring information from the code, and entropy decoding is not necessarily used either. For example, when the read code itself indicates information (for example, a flag of 0 or 1), it is only necessary to read the code.

[0219] Next, the entropy decoding unit 202 decodes quantization coefficients from the bit stream (S703) and ends the decoding process.

[0220] When intra prediction is used for the block to be decoded (intra in S701), the entropy decoding unit 202 decodes an intra prediction mode signal from the bit stream (S704). Further, the entropy decoding unit 202 decodes an adaptive selection mode signal (S705).

[0221] Here, when the adaptive basis selection mode is used and the size of the block to be decoded is different from the predetermined size (YES in S706), the entropy decoding unit 202 decodes a first basis selection signal from the bit stream (S707). On the other hand, when the adaptive basis selection mode is not used, or when the adaptive basis selection mode is used and the size of the block to be decoded matches the predetermined size (NO in S706), the entropy decoding unit 202 skips the decoding step of the first basis selection signal (S707). That is, the entropy decoding unit 202 does not decode the first basis selection signal.

[0222] Here, when performing the second inverse transformation (YES in S708), the entropy decoding unit 202 decodes a second basis selection signal from the bit stream (S709). On the other hand, when not performing the second inverse transformation (NO in S708), the entropy decoding unit 202 skips the decoding step of the second basis selection signal (S709). That is, the entropy decoding unit 202 does not decode the second basis selection signal.

[0223] Finally, the entropy decoding unit 202 decodes quantization coefficients from the bit stream (S710) and ends the decoding process.

[0224] [Processing of the Inverse Quantization Unit and Inverse Transformation Unit of the Decoding Device] Next, the processing of the inverse quantization unit 204 and the inverse transform unit 206 of the decoding apparatus 200 according to the present embodiment will be specifically described with reference to FIG. 18. FIG. 18 is a flowchart showing the inverse quantization and inverse transform processing in the decoding apparatus 200 according to Embodiment 4.

[0225] First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S501). The inverse transform unit 206 determines whether to use intra prediction or inter prediction for the block to be decoded (S502). When it is determined to use inter prediction for the block to be decoded (inter in S502), the inverse transform unit 206 selects the first inverse transform basis for the block to be decoded (S503). The inverse transform unit 206 performs the first inverse transform on the inverse quantized coefficients of the block to be decoded using the first inverse transform basis selected in step S503 (S504), and ends the inverse quantization and inverse transform processing.

[0226] When it is determined to use intra prediction for the block to be decoded (intra in S502), the inverse transform unit 206 determines whether the size of the block to be decoded matches a predetermined size and whether an adaptive basis selection mode is used for the block to be decoded (S801). For example, the inverse transform unit 206 determines whether an adaptive basis selection mode is used based on the adaptive selection mode signal decoded in step S705 of FIG. 17.

[0227] When the size of the block to be decoded is different from the predetermined size or the adaptive basis selection mode is not used (NO in S801), the inverse transform unit 206 determines whether the intra prediction mode of the block to be decoded is a non-directional prediction mode (S601).

[0228] Here, when the intra prediction mode is not the non-directional prediction mode (NO in S601), the inverse transform unit 206 selects a second inverse transform basis for the block to be decoded (S602). For example, the inverse transform unit 206 selects the second inverse transform basis based on the second basis selection signal decoded in step S709 of FIG. 17. The inverse transform unit 206 performs a second inverse transform on the inverse quantized coefficients of the block to be decoded using the selected second inverse transform basis (S603). The inverse transform unit 206 selects a first inverse transform basis (S604). For example, when the adaptive basis selection mode is used, the inverse transform unit 206 selects the first inverse transform basis based on the first basis selection signal decoded in step S707 of FIG. 17. The inverse transform unit 206 performs a first inverse transform on the coefficients obtained by the second inverse transform in step S603 using the selected first inverse transform basis (S605), and ends the inverse quantization and inverse transform processing.

[0229] On the other hand, when the intra prediction mode is the non-directional prediction mode (YES in S601), the inverse transform unit 206 determines whether the first inverse transform basis matches the second predetermined inverse transform basis (S606). For example, when the adaptive basis selection mode is used, the inverse transform unit 206 determines whether the first inverse transform basis matches the second predetermined inverse transform basis based on the first basis selection signal decoded in step S707 of FIG. 17. As the second predetermined inverse transform basis, an inverse transform basis corresponding to the second predetermined transform basis used in the encoding device 100 is used.

[0230] Here, when the first inverse transformation basis coincides with the second predetermined inverse transformation basis (YES in S606), the inverse transformation unit 206 selects a second inverse transformation basis for the block to be decoded (S607). For example, the inverse transformation unit 206 selects the second inverse transformation basis based on the second basis selection signal decoded in step S709 of FIG. 17. The inverse transformation unit 206 performs a second inverse transformation on the inverse quantized coefficients of the block to be decoded using the selected second inverse transformation basis (S608). The inverse transformation unit 206 selects the first inverse transformation basis (S609). For example, when the adaptive basis selection mode is used, the inverse transformation unit 206 selects the first inverse transformation basis based on the first basis selection signal decoded in step S707 of FIG. 17. The inverse transformation unit 206 performs a first inverse transformation on the coefficients obtained by the second inverse transformation in step S608 using the selected first inverse transformation basis (S610), and ends the inverse quantization and inverse transformation processing.

[0231] When the size of the block to be decoded matches the predetermined size and the adaptive basis selection mode is used (YES in S801), the inverse transformation unit 206 determines whether the intra prediction mode of the block to be decoded is a non-directional prediction mode (S802).

[0232] Here, when the intra prediction mode is not the non-directional prediction mode (NO in S802), the inverse transformation unit 206 selects a second inverse transformation basis for the block to be decoded (S803). For example, the inverse transformation unit 206 selects the second inverse transformation basis based on the second basis selection signal decoded in step S709 of FIG. 17. The inverse transformation unit 206 performs a second inverse transformation on the inverse quantized coefficients of the block to be decoded using the selected second inverse transformation basis (S804). The inverse transformation unit 206 fixes the first inverse transformation basis to the first predetermined inverse transformation basis. As the first predetermined inverse transformation basis, an inverse transformation basis corresponding to the first predetermined transformation basis used in the encoding device 100 is used. The inverse transformation unit 206 performs a first inverse transformation on the coefficients obtained by the second inverse transformation in step S804 using the fixed first inverse transformation basis (S806), and ends the inverse quantization and inverse transformation processing.

[0233] On the other hand, when the intra prediction mode is the non-directional prediction mode (YES in S802), the inverse transform unit 206 skips the second inverse transform basis selection step (S803) and the second inverse transform step (S804). That is, the inverse transform unit 206 fixes the first inverse transform basis to the first predetermined inverse transform basis without performing the second inverse transform (S805). The inverse transform unit 206 performs the first inverse transform on the coefficients inverse quantized in step S501 using the fixed first inverse transform basis (S806), and ends the inverse quantization and inverse transform processing.

[0234] [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 according to the present embodiment, when the adaptive basis selection mode is used, the first transform basis can be fixed according to the block size. Therefore, the load of the first transform in the adaptive basis selection mode can be reduced.

[0235] Note that in the present embodiment, the second transform is not performed when inter prediction is used for the encoding target block, but the present invention is not limited to this. That is, when inter prediction is used for the encoding target block, the second transform may be performed on the first transform coefficients generated by the first transform. In this case, the second transform coefficients generated by the second transform are quantized.

[0236] Note that the order of the steps in the flowcharts of FIGS. 15 to 18 is not limited to the order described in FIGS. 15 to 18. For example, in FIG. 16, the encoding order of the signal may be another order defined in advance by a standard or the like.

[0237] Note that in the present embodiment, a plurality of signals (intra prediction mode signal, adaptive selection mode signal, first basis selection signal, and second basis selection signal) are encoded in the bitstream, but these plurality of signals may not be encoded in the bitstream. For example, these plurality of signals may be notified from the encoding device 100 to the decoding device 200 separately from the bitstream.

[0238] In addition, in the present embodiment, the positions of each of the plurality of signals (intra prediction mode signal, adaptive selection mode signal, first basis selection signal, and second basis selection signal) within the bit stream are not particularly limited. The plurality of signals are encoded, for example, in at least one of the plurality of headers. As the plurality of headers, for example, a video parameter set, a sequence parameter set, a picture parameter set, and a slice header can be used. When the signal is in a plurality of layers (for example, a picture parameter set and a slice header), the signal in the lower layer (for example, a slice header) overwrites the signal in the higher layer (for example, a picture parameter set).

[0239] Note that this aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0240] (Embodiment 5) In the above embodiments and each modification, each of the functional blocks can usually be realized by an MPU, a memory, etc. Also, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (a dedicated circuit).

[0241] Also, the processing described in the embodiments and each modification may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing may be performed, or distributed processing may be performed.

[0242] The present invention is not limited to the above embodiments, and various modifications are possible, and they are also included within the scope of the present invention.

[0243] Furthermore, here, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in the above embodiments and each modification example, and a system using the same will be described. The system is characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, and an image encoding / decoding device having both. Regarding other configurations in the system, they can be appropriately changed as the case may be.

[0244] [Usage Example] FIG. 19 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed radio stations, are installed in each cell.

[0245] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above elements. Without passing through the base stations ex106 to ex110, which are fixed radio stations, the devices may be directly or indirectly connected to each other via a telephone network or short-range wireless, etc. Further, the streaming server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101 or the like. Also, the streaming server ex103 is connected to terminals in a hotspot in an airplane ex117 via a satellite ex116.

[0246] Note that a wireless access point, a hot spot, or the like may be used instead of the base stations ex106 to ex110. Further, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0247] The camera ex113 is a device capable of taking still images and video, such as a digital camera. Further, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handyphone System) or the like that generally supports the mobile communication system standards known as 2G, 3G, 3.9G, 4G, and in the future 5G.

[0248] The home appliance ex118 is a device included in a refrigerator or a household fuel cell cogeneration system.

[0249] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 through the base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal in the airplane ex117, etc.) performs the encoding process described in the above embodiments and each modification on the still image or video content photographed by the user using the terminal, multiplexes the video data obtained by encoding with the audio data obtained by encoding the audio corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present invention.

[0250] On the one hand, the streaming server ex103 streams the transmitted content data to the requested client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc., which is capable of decrypting the encoded data. Each device that receives the distributed data decrypts and plays back the received data. That is, each device functions as an image decoding device according to an aspect of the present invention.

[0251] [Distributed Processing] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world. In a CDN, an edge server physically close to the client is dynamically assigned according to the client. Then, by caching and distributing the content to the edge server, the delay can be reduced. Also, when some error occurs or the communication state changes due to an increase in traffic, etc., processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the distribution, so high-speed and stable distribution can be realized.

[0252] Furthermore, not limited to the distributed processing of the distribution itself, the encoding process of the captured data may be performed on each terminal, on the server side, or shared between them. As an example, in general encoding processes, the processing loop is performed twice. In the first loop, the complexity of the image or the amount of code is detected in units of frames or scenes. In the second loop, processing is performed to improve the encoding efficiency while maintaining the image quality. For example, if the terminal performs the first encoding process and the server that receives the content performs the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode in almost real time, the already encoded data from the first encoding performed by the terminal can be received and played back by other terminals, enabling more flexible real-time distribution.

[0253] As another example, cameras such as ex113 perform feature extraction from the image, compress the data related to the features as metadata, and transmit it to the server. The server performs compression according to the meaning of the image, such as determining the importance of the object from the features and switching the quantization accuracy. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.

[0254] As yet another example, in a stadium, shopping mall, or factory, etc., there may be a case where there are multiple video data in which substantially the same scene is captured by multiple terminals. In this case, using the multiple terminals that performed the shooting, and other terminals and the server that did not perform the shooting as necessary, encoding processes are respectively assigned and distributed processing is performed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture. This can reduce the delay and achieve more real-time performance.

[0255] In addition, since the plurality of video data is of almost the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can be referred to each other. Alternatively, the server may receive the encoded data from each terminal and change the reference relationship between the plurality of data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.

[0256] In addition, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert the MPEG-based encoding method to the VP-based method, or convert H.264 to H.265.

[0257] In this way, the encoding process can be performed by the terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the subject performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.

[0258] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and use different scenes captured by terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are almost synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal are integrated based on the relative positional relationship between the terminals obtained separately, or the area where the feature points included in the videos match.

[0259] The server may not only encode two-dimensional moving images, but also automatically encode still images based on scene analysis of the moving images or at a time specified by the user, and transmit them to the receiving terminal. Further, when the server can acquire the relative positional relationship between the shooting terminals, it can generate the three-dimensional shape of the scene based not only on two-dimensional moving images, but also on videos shot from different angles of the same scene. Note that the server may separately encode three-dimensional data generated by a point cloud or the like, or select or reconstruct the video to be transmitted to the receiving terminal based on the result of recognizing or tracking a person or an object using the three-dimensional data from the videos shot by multiple terminals.

[0260] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or enjoy the content obtained by cutting out the video from an arbitrary viewpoint from the three-dimensional data reconstructed using multiple images or videos. Further, similar to the video, sound is also collected from multiple different angles, and the server may multiplex and transmit the sound from a specific angle or space together with the video according to the video.

[0261] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become widespread. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by Multi-View Coding (MVC) or the like, or encode them as separate streams without referring to each other. At the time of decoding the separate streams, it is preferable to synchronize and reproduce them so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0262] In the case of an AR image, the server superimposes virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold the virtual object information and the three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting the virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like, and encode it in a state where the portion is transparent. Alternatively, the server may set an RGB value of a predetermined value as a background like a chroma key and generate data with the portion other than the object being the background color.

[0263] Similarly, the decoding process of the distributed data may be performed on each terminal that is a client, on the server side, or may be shared between them. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request on another terminal, perform the decoding process, and send the decoded signal to the device having the display. By dispersing the process regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with good image quality. As another example, while receiving large-size image data on a TV or the like, a part of the area such as a tile in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to check at hand the area of one's own field of responsibility or the area that one wants to check in more detail.

[0264] In the future, it is expected that, regardless of whether indoors or outdoors, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used, content will be received seamlessly while switching appropriate data for the ongoing communication by utilizing a distribution system standard such as MPEG-DASH. As a result, the user can freely select not only their own terminal but also a decoding device or display device such as a display installed indoors or outdoors and switch in real time. Also, based on their own location information and the like, decoding can be performed while switching the terminal for decoding and the terminal for display. This makes it possible to move while displaying map information on a part of the wall surface or ground of the adjacent building where a displayable device is embedded during movement to the destination. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time or being copied to an edge server in a content delivery service, it is also possible to switch the bitrate of the received data.

[0265] [Scalable Encoding] Regarding content switching, it will be described using a scalable stream compressed and encoded by applying the moving image encoding method shown in the above-described embodiments and each modification example shown in FIG. 20. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of the temporally / spatially scalable stream realized by performing encoding in layers as shown in the figure, a configuration for switching content may also be used. That is, by determining up to which layer to decode according to internal factors such as the performance on the decoding side and external factors such as the state of the communication bandwidth, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on your smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, so the burden on the server side can be reduced.

[0266] Furthermore, as described above, pictures are encoded for each layer, and in addition to the configuration that realizes scalability where an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of the image, etc., and the decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta information includes information for specifying linear or non-linear filter coefficients used for super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least-squares operation used for super-resolution processing, etc.

[0267] Alternatively, the picture may be divided into tiles or the like according to the meaning of an object or the like in the image, and the decoding side may decode only a part of the region by selecting the tile to be decoded. Also, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 21, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0268] Also, the meta information may be stored in a unit composed of a plurality of pictures such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining it with the information for each picture, can specify the picture in which the object exists and the position of the object in the picture.

[0269] [Optimization of Web Page] FIG. 22 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 23 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 22 and 23, a web page may include a plurality of link images that are links to image contents, and the appearance thereof varies depending on the device for viewing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) displays a still image or an I picture that each content has as a link image, displays a video like a gif animation with a plurality of still images or I pictures, etc., or receives only the base layer and decodes and displays the video.

[0270] When a link image is selected by the user, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Also, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay between the decoding time and the display time of the leading picture (the delay from the start of content decoding to the start of display) by decoding and displaying only the forward-reference pictures (I pictures, P pictures, B pictures with only forward reference). Also, the display device may deliberately ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures with forward reference, and perform normal decoding as the received pictures increase over time.

[0271] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode these in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.

[0272] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal move, the receiving terminal can achieve seamless reception and decoding by transmitting the position information of the receiving terminal at the time of the reception request while switching between the base stations ex106 to ex110. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication band.

[0273] As described above, in the content supply system ex100, the client can receive, decode, and play back the encoded information transmitted by the user in real time.

[0274] [Delivery of Personal Content] Also, in the content supply system ex100, not only high-quality and long-duration content by video delivery providers but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. Also, it is considered that such personal content will increase in the future. In order to make personal content into better content, the server may perform an encoding process after performing an editing process. This can be realized, for example, with the following configuration.

[0275] During shooting in real time or accumulating and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection on the original image or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes unimportant scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone for editing. The server encodes the edited data based on the editing results. Also, it is known that if the shooting time is too long, the viewing rate will decrease. The server may automatically clip not only unimportant scenes but also scenes with little movement within a specific time range according to the shooting time so that the content is within that time range, based on the image processing results. Or, the server may generate a digest based on the result of semantic analysis of the scene, encode it, and output it.

[0276] In addition, in the case of personal content, there may be cases where it contains content that would directly infringe copyright, moral rights of the author, or portrait rights, etc., and it may be inconvenient for the individual, such as the sharing scope exceeding the intended scope. Therefore, for example, the server may deliberately change the image of a person's face in the peripheral part of the screen or inside a house to an out-of-focus image and then encode it. Also, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, the user designates a person or background area that the user wants to process the image from the perspective of copyright, etc., and the server can perform processing such as replacing the designated area with another video or blurring the focus. In the case of a person, the video of the face part can be replaced while tracking the person in the moving image.

[0277] In addition, since the viewing of personal content with a small data volume has a strong requirement for real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and performs decoding and playback. During this time, the decoding device receives the enhancement layer, and when playback is looped or played back two or more times, such as when the enhancement layer is also included, high-quality video may be played back. For a stream encoded in a scalable manner like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played back for the first time and a second stream encoded with reference to the first video are configured as one stream.

[0278] [Other usage examples] In addition, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processing may be performed using the software. Further, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.

[0279] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the content encoding method or has the ability to execute a specific service. If the terminal does not support the content encoding method or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays back the content.

[0280] Moreover, not limited to the content supply system ex100 via the Internet ex101, at least either a moving image encoding device (image encoding device) or a moving image decoding device (image decoding device) of the above-described embodiments and each modification example can be incorporated into a digital broadcast system. In order to transmit and receive multiplexed data in which video and audio are multiplexed on a broadcast radio wave using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the configuration of the content supply system ex100 that is easy for unicast, but the same application is possible for the encoding process and the decoding process.

[0281] [Hardware Configuration] FIG. 24 is a diagram showing a smartphone ex115. FIG. 25 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from a base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by the camera unit ex465 and video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 that is an interface unit with a SIM ex468 for identifying a user and authenticating access to various data including a network. Note that an external memory may be used instead of the memory unit ex467.

[0282] In addition, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a bus ex470.

[0283] When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack to activate the smartphone ex115 in an operable state.

[0284] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, inverse spectral spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. In the data communication mode, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body unit, and the same transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in the above embodiment and each modification example, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while a video or still image is being captured by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.

[0285] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in the above embodiment and each modification example, and the video or still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming has become widespread, depending on the user's situation, there may be a situation where it is not socially appropriate to play audio. Therefore, as an initial value, it is desirable to have a configuration that plays only video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.

[0286] Also, here, the smartphone ex115 has been described as an example, but as a terminal, in addition to a transmission / reception type terminal having both an encoder and a decoder, there are three possible implementation forms: a transmission terminal having only an encoder and a reception terminal having only a decoder. Furthermore, in the digital broadcast system, although it has been described that multiplexed data in which video data and audio data are multiplexed is received or transmitted, the multiplexed data may include character data related to the video in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0287] Although the main control unit ex460 including a CPU has been described as controlling the encoding or decoding process, many terminals also include a GPU. Therefore, a configuration in which a large area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so as to be commonly used may be adopted. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be realized. In particular, it is efficient to perform processes such as motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures using the GPU instead of the CPU.

Industrial Applicability

[0288] The present disclosure can be applied to, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, or a digital video camera.

Explanation of Signs

[0289] 100 Encoding device 102 Splitting unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 200 Decoding device 202 Entropy decoding unit

Claims

1. A decoding apparatus, comprising: a processor and a memory, wherein the processor uses the memory to determine whether to use intra prediction for a block to be decoded and whether the block to be decoded has a predetermined size, when it is determined to use intra prediction for the block to be decoded and it is determined that the block to be decoded has the predetermined size, further determine whether the intra prediction mode of the block to be decoded is a predetermined mode, when the intra prediction mode is not the predetermined mode, perform a second inverse transformation on the inverse quantization coefficients of the block to be decoded, and further perform a first inverse transformation on the coefficients obtained by the second inverse transformation, when the intra prediction mode is the predetermined mode, skip the second inverse transformation and perform a first inverse transformation on the inverse quantization coefficients of the block to be decoded, a decoding apparatus.

2. An encoding apparatus, comprising: a processor and a memory, wherein the processor uses the memory to determine whether to use intra prediction for a block to be encoded and whether the block to be encoded has a predetermined size, when it is determined to use intra prediction for the block to be encoded and it is determined that the block to be encoded has the predetermined size, further determine whether the intra prediction mode of the block to be encoded is a predetermined mode, generate first transformation coefficients by performing a first transformation on the residual signal of the block to be encoded, when the intra prediction mode is not the predetermined mode, generate second transformation coefficients by performing a second transformation on the first transformation coefficients, and further quantize the second transformation coefficients, when the intra prediction mode is the predetermined mode, quantize the first transformation coefficients, an encoding apparatus.

3. A decoding method, comprising: determine whether to use intra prediction for a block to be decoded and whether the block to be decoded has a predetermined size, when it is determined to use intra prediction for the block to be decoded and it is determined that the block to be decoded has the predetermined size, further determine whether the intra prediction mode of the block to be decoded is a predetermined mode, when the intra prediction mode is not the predetermined mode, perform a second inverse transformation on the inverse quantization coefficients of the block to be decoded, and further perform a first inverse transformation on the coefficients obtained by the second inverse transformation, When the intra prediction mode is the predetermined mode, skip the second inverse transformation and perform a first inverse transformation on the inverse quantization coefficients of the block to be decoded. Decoding method. **Claim 4** An encoding method, comprising: determining whether to use intra prediction for a block to be encoded and whether the block to be encoded has a predetermined size; when it is determined to use intra prediction for the block to be encoded and the block to be encoded is determined to have the predetermined size, further determining whether the intra prediction mode of the block to be encoded is a predetermined mode; generating first transformation coefficients by performing a first transformation on a residual signal of the block to be encoded; when the intra prediction mode is not the predetermined mode, generating second transformation coefficients by performing a second transformation on the first transformation coefficients, and further quantizing the second transformation coefficients; when the intra prediction mode is the predetermined mode, quantizing the first transformation coefficients. Encoding method.