Encoding device, encoding method, decoding device, decoding method, and method for transmitting bit streams

TWI937575BActive Publication Date: 2026-09-01PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
TW113137890
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-07-13
Filing Date
2018-07-10
Publication Date
2026-09-01
Estimated Expiration
2038-07-09

Smart Images

  • Figure TWG2TB001908553_001
    Figure TWG2TB001908553_001
  • Figure TWG2TB001908553_002
    Figure TWG2TB001908553_002
  • Figure TWG2TB001908553_003
    Figure TWG2TB001908553_003
Patent Text Reader

Abstract

The encoding device of the present invention encodes the encoding object block of an image and includes a circuit and a memory. The circuit uses the memory and a first conversion substrate to perform a first conversion on the residual signal of the encoding object block, thereby generating a first conversion coefficient. When the first conversion substrate is consistent with a predetermined conversion substrate, a second conversion is performed on the first conversion coefficient using a second conversion substrate, thereby generating a second conversion coefficient, and the second conversion coefficient is quantized. When the first conversion substrate is different from the predetermined conversion substrate, the second conversion is not performed and the first conversion coefficient is quantized.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding Device, Encoding Method, Decoding Device, and Decoding Method Field of the Invention The present disclosure relates to encoding and decoding of images / videos in block units. Background of the Invention An image encoding standard specification called HEVC (High-Efficiency Video Coding) was standardized by JCT-VC (Joint Collaborative Team on Video Coding). Prior Art Documents Non-Patent Document [Non-Patent Document 1] H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding)) Summary of the Invention Problems to be Solved by the Invention Such encoding and decoding technologies require suppressing a decrease in compression efficiency and reducing the processing load. Therefore, the present disclosure provides an encoding device, a decoding device, an encoding method, or a decoding method that can achieve suppressing a decrease in compression efficiency and reducing the processing load. Means for Solving the Problems An encoding device according to an aspect of the present disclosure encodes an encoding target block of a picture, and includes a circuit and a memory. The circuit uses the memory to perform a first transformation on the residual signal of the encoding target block using a first transformation basis, thereby generating first transformation coefficients. When the first transformation basis matches a predetermined transformation basis, the circuit performs a second transformation on the first transformation coefficients using a second transformation basis, thereby generating second transformation coefficients, and quantizes the second transformation coefficients. When the first transformation basis is different from the predetermined transformation basis, the circuit quantizes the first transformation coefficients without performing the second transformation. Furthermore, these general or specific aspects can also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium. Advantages of the Invention The present invention can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve suppressing a decrease in compression efficiency and reducing the processing load. Modes for Carrying Out the Invention (Insight Underlying the Present Invention) In the JEM (Joint Exploration Test Model) software of the JVET (Joint Video Exploration Team), a proposal is made for a two-stage frequency conversion of blocks applicable to intra-frame prediction. In the two-stage frequency conversion, the first conversion uses EMT (Explicit Multiple Core Transform), and the second conversion uses NSST (Non-separable Secondary Transform). In EMT, a plurality of transform bases are adaptively selected to perform a conversion from the spatial domain to the frequency domain. In such a two-stage frequency conversion, there is still room for improvement from the perspective of processing volume. Hereinafter, embodiments based on this insight will be specifically described with reference to the drawings. Furthermore, all the embodiments described below represent comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are examples, and their gist is not intended to limit the scope of the patent application. Also, among the components of the following embodiments, the components not described in the independent claims representing the highest-level concept are described as optional components. (Embodiment 1) First, the outline of Embodiment 1 will be described as an example of an encoding device and a decoding device that can apply the processing and / or configuration described in each aspect of the present disclosure to be described later. However, Embodiment 1 is only an example of an encoding device and a decoding device that can apply the processing and / or configuration described in each aspect of the present disclosure, and the processing and / or configuration described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1. When applying the processing and / or configuration described in each aspect of the present disclosure to Embodiment 1, any of the following may be performed, for example. (1) For the encoding device or decoding device of Embodiment 1, replace the component corresponding to the component described in each aspect of the present disclosure among the plurality of components constituting the encoding device or decoding device with the component described in each aspect of the present disclosure; (2) For the encoding device or decoding device of Embodiment 1, after arbitrarily changing the functions or implementation processes of some of the plurality of components constituting the encoding device or decoding device, such as adding, replacing, or deleting, replace the component corresponding to the component described in each aspect of the present disclosure with the component described in each aspect of the present disclosure; (3) For the method implemented by the encoding device or decoding device of Embodiment 1, after arbitrarily changing the addition of processing and / or a part of the plurality of types of processing included in the method, such as replacement or deletion, replace the processing corresponding to the processing described in each aspect of the present disclosure with the processing described in each aspect of the present disclosure; (4) Combine and implement a part of the constituent elements among the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1 with the constituent elements described in each aspect of the present disclosure, the constituent elements having a part of the functions possessed by the constituent elements described in each aspect of the present disclosure, or the constituent elements implementing a part of the processing implemented by the constituent elements described in each aspect of the present disclosure; (5) Combine and implement the constituent elements having a part of the functions possessed by a part of the constituent elements among the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1, or the constituent elements implementing a part of the processing implemented by a part of the constituent elements among the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1, with the constituent elements described in each aspect of the present disclosure, the constituent elements having a part of the functions possessed by the constituent elements described in each aspect of the present disclosure, or the constituent elements implementing a part of the processing implemented by the constituent elements described in each aspect of the present disclosure; (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processing corresponding to the processing described in each aspect of the present disclosure among the plurality of types of processing included in the method with the processing described in each aspect of the present disclosure; (7) Combine and implement a part of the plurality of types of processing included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processing described in each aspect of the present disclosure. Furthermore, the implementation manners of the processing and / or configuration described in each aspect of the present disclosure are not limited to the above examples. For example, it can be implemented in a device used for different purposes from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processing and / or configuration described in each aspect can be implemented alone. Also, the processing and / or configuration described in different aspects can be combined and implemented. [Outline of Encoding Device] First, an outline of the encoding device of Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 of Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units. As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a segmentation unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128. The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. At this time, when the processor executes a software program stored in the memory, the processor functions as the segmentation unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filtering unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 can also be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filtering unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The following describes each component included in the encoding device 100. [Segmentation unit] The segmentation unit 102 divides each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first divides the picture into blocks of a fixed size (for example, 128×128). Such a block of a fixed size is sometimes referred to as a coding tree unit (CTU). Then, the segmentation unit 102 divides each of the fixed-size blocks into variable-size blocks (for example, 64×64) according to recursive quadtree and / or binary tree block partitioning. Such a variable-size block is sometimes referred to as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). Furthermore, in this embodiment, it is not necessary to distinguish between CU, PU, and TU, and a part or all of the blocks in the picture may be the processing units of CU, PU, and TU. FIG. 2 is a diagram showing an example of block partitioning according to Embodiment 1. In FIG. 2, solid lines represent block boundaries of quadtree block partitioning, and dashed lines represent block boundaries of binary tree block partitioning. Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into 4 square 64×64 blocks (quadtree block partitioning). The upper left 64×64 block is further vertically divided into two rectangular 32×64 blocks, and the left 32×64 block is further vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the upper left 64×64 block is divided into two 16×64 blocks 11, 12 and a 32×64 block 13. The upper right 64×64 block is horizontally divided into two rectangular 64×32 blocks 14, 15 (binary tree block division). The lower left 64×64 block is divided into four square 32×32 blocks (quadtree block division). Among the four 32×32 blocks, the upper left block and the lower right block are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the lower left 64×64 block is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20 and two 32×16 blocks 21, 22. The lower right 64×64 block 23 is not divided. As described above, in FIG. 2, the block 10 is divided into 13 variable-size blocks 11 to 23 according to recursive quadtree and binary tree block division. Such division is sometimes referred to as QTBT (quad-tree plus binary tree) division. Furthermore, in FIG. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary tree block division). Such division including ternary tree block division is sometimes referred to as MBT (multi type tree) division. [Subtraction unit] The subtraction unit 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division unit 102. In short, the subtraction unit 104 calculates the prediction error (also referred to as the residual) of the coding target block (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106. The original signal is the input signal of the coding device 100, which is a signal representing the images of the respective pictures constituting the dynamic image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing the image is sometimes also referred to as a sample. [Conversion unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on, for example, the prediction error in the spatial domain. Furthermore, the conversion unit 106 can also adaptively select a conversion type from a plurality of conversion types, and use a transform basis function corresponding to the selected conversion type to convert the prediction error into conversion coefficients. This type of conversion is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform). The plurality of conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the conversion basis functions corresponding to the respective conversion types. In FIG. 3, N represents the number of input pixels. When selecting a conversion type from the plurality of conversion types, it may depend on, for example, the type of prediction (intra-frame prediction and inter-frame prediction), or it may depend on the intra-frame prediction mode. Information indicating whether EMT or AMT is applicable (for example, referred to as an AMT flag), and information indicating the selected conversion type are signaled at the CU level. Furthermore, the signaling of this information need not be limited to the CU level, and may also be at other levels (for example, sequence level, picture level, slice level, block level, or CTU level). Also, the conversion unit 106 can also re-convert the conversion coefficients (conversion results). This type of re-conversion is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion one by one on sub-blocks (for example, 4×4 sub-blocks) included in a block of conversion coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applicable, and information related to the conversion matrix for NSST are signaled at the CU level. Furthermore, the signaling of this information need not be limited to the CU level, and may also be at other levels (for example, sequence level, picture level, slice level, block level, or CTU level). Here, Separable conversion refers to a method of performing a plurality of conversions by only separating the dimensions of the input in each direction, and Non-Separable conversion refers to a method of unifying and treating two or more dimensions as one dimension and performing a unified conversion when the input is multi-dimensional. For example, as an example of Non-Separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and conversion processing is performed on this array using a 16×16 conversion matrix. Also, similarly regarding a 4×4 input block as an array having 16 elements, performing a transformation of a plurality of Givens rotations (Hypercube Givens Transform) on this array is also an example of a Non-Separable transformation. [Quantization Unit] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients according to the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112. The predetermined order is the order for quantization / inverse quantization of the transform coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency). The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the quantization parameter value is increased, the quantization step also increases. In short, if the quantization parameter value increases, the quantization error increases. [Entropy Encoding Unit] The entropy encoding unit 110 performs variable length encoding on the quantization coefficients input from the quantization unit 108, thereby generating an encoded signal (encoded bit stream). Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal. [Inverse Quantization Unit] The inverse quantization unit 112 inverse quantizes the quantization coefficients, which are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114. [Inverse Transform Unit] The inverse transform unit 114 inverse transforms the transform coefficients, which are the input from the inverse quantization unit 112, thereby restoring the prediction error. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116. Furthermore, since the restored prediction error is information lost due to quantization, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error contains quantization error. [Addition Unit] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. Then, the adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes also referred to as a partial decoded block. [Block Memory] The block memory 118 is a memory unit for storing blocks within a frame that are referenced for prediction and are within the picture to be encoded (hereinafter referred to as the current picture). Specifically, the block memory 118 stores the reconstructed block output from the adder 116. [Loop Filter Unit] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter used within the encoding loop (in-loop filter), and includes, for example, a deblocking filter (DF), sample adaptivity offset (SAO), and adaptive loop filter (ALF). The ALF applies a least squares error filter for removing encoding distortion. For example, for a 2×2 sub-block within the current block, one filter selected from a plurality of filters is applied one by one according to the direction and activity of the local gradient. Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into a plurality of groups (e.g., 15 or 25 groups). The classification of the sub-blocks is performed according to the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C (e.g., C = 5D + A) is calculated. Then, according to the classification value C, the sub-blocks are classified into a plurality of groups (e.g., 15 or 25 groups). The direction value D of the gradient is derived by, for example, comparing the gradients in a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Also, the activity value A of the gradient is derived by, for example, adding the gradients in a plurality of directions and quantifying the addition result. Based on the results of such classification, the filter for the sub-block is determined from a plurality of filters. The shape of the filter used by the ALF can use, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used by the ALF. FIG. 4A shows a 5×5 diamond-shaped filter, FIG. 4B shows a 7×7 diamond-shaped filter, and FIG. 4C shows a 9×9 diamond-shaped filter. The information indicating the filter shape is signaled at the picture level. Furthermore, the signaling of the information indicating the filter shape does not have to be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, block level, CTU level, or CU level). The enabling / disabling of ALF is determined at, for example, the picture level or the CU level. For example, in terms of luminance, it is determined at the CU level whether to apply ALF, and in terms of chrominance difference, it is determined at the picture level whether to apply ALF. The information indicating the enabling / disabling of ALF is signaled at the picture level or the CU level. Furthermore, the signaling of the information indicating the enabling / disabling of ALF need not be limited to the CU level and may also be at other levels (such as the sequence level, slice level, block level, or CTU level). The set of coefficients of a plurality of selectable filters (for example, filters from 15 to 25) is signaled at the picture level. Furthermore, the signaling of the set of coefficients need not be limited to the picture level and may also be at other levels (such as the sequence level, slice level, block level, CTU level, CU level, or sub-block level). [Frame memory] The frame memory 122 is a memory unit for storing reference pictures used for inter-frame prediction and is sometimes also referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120. [Intra-frame prediction unit] The intra-frame prediction unit 124 performs intra-frame prediction (also referred to as in-picture prediction) of the current block by referring to the blocks within the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 performs intra-frame prediction by referring to the samples (such as luminance values, chrominance difference values) of the blocks adjacent to the current block, thereby generating an intra-frame prediction signal and outputting the intra-frame prediction signal to the prediction control unit 128. For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of predefined intra-frame prediction modes. The plurality of intra-frame prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes. The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1). The plurality of directional prediction modes include, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Furthermore, in addition to the 33 directions, the plurality of directional prediction modes may further include the 32-direction prediction mode (a total of 65 directional prediction modes). FIG. 5A is a diagram showing the 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) of intra-frame prediction. The solid arrows indicate the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows indicate the additional 32 directions. Furthermore, in the intra-frame prediction of the color difference block, the luminance block can also be referred to. In summary, the color difference component of the current block can also be predicted based on the luminance component of the current block. This type of intra-frame prediction is sometimes referred to as CCLM (cross-component linear model) prediction. This type of intra-frame prediction mode of the color difference block that refers to the luminance block (for example, called the CCLM mode) can also be added as one of the intra-frame prediction modes of the color difference block. The intra-frame prediction unit 124 can also correct the pixel value after intra-frame prediction according to the gradient of the reference pixels in the horizontal / vertical direction. The intra-frame prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (referred to as, for example, the PDPC flag) is signaled at the CU level. Furthermore, the signaling of this information does not have to be limited to the CU level and can also be at other levels (such as the sequence level, picture level, slice level, block level, or CTU level). [Inter-frame prediction unit] The inter-frame prediction unit 126 refers to the reference picture stored in the frame memory 122 and a reference picture different from the current picture to perform inter-frame prediction (also referred to as inter-picture prediction) of the current picture, thereby generating a prediction signal (inter-frame prediction signal). The inter-frame prediction is performed in units of the current block or a sub-block within the current block (for example, a 4×4 block). For example, the inter-frame prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter-frame prediction unit 126 uses the motion information (such as a motion vector) obtained by the motion estimation to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block. Then, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128. The motion information for motion compensation is signaled. The signaling of the motion vector can also use a motion vector predictor. In summary, the difference between the motion vector and the motion vector predictor can also be signaled. Furthermore, not only the motion information of the current block obtained by motion estimation can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame prediction signal. Specifically, the prediction signal based on the motion information obtained by motion estimation and the prediction signal based on the motion information of adjacent blocks can be weighted and added to generate an inter-frame prediction signal in units of sub-blocks within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation). In such an OBMC mode, information indicating the size of the sub-block for OBMC (e.g., referred to as the OBMC block size) is signaled at the sequence level. Also, information indicating whether the OBMC mode is applied or not (e.g., referred to as the OBMC flag) is signaled at the CU level. Furthermore, the signaling levels of such information do not have to be limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, block level, CTU level, or sub-block level). More specifically, the OBMC mode will be described. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process for OBMC processing. First, using the motion vector (MV) assigned to the coding target block, a prediction image (Pred) for general motion compensation is obtained. Next, the motion vector (MV_L) of the left adjacent block that has been encoded is applied to the coding target block to obtain a prediction image (Pred_L). The aforementioned prediction image and Pred_L are weighted and overlapped to perform the first correction of the prediction image. Similarly, the motion vector (MV_U) of the upper adjacent block that has been encoded is applied to the coding target block to obtain a prediction image (Pred_U). The prediction image that has undergone the aforementioned first correction and Pred_U are weighted and overlapped to perform the second correction of the prediction image, which is used as the final prediction image. Furthermore, a method of two-stage correction using the right adjacent block and the upper adjacent block has been described here, but a configuration in which correction more than two stages is performed using the right adjacent block or the lower adjacent block can also be adopted. Furthermore, the overlapping area is not the entire pixel area of the block, and can also be only a partial area near the block boundary. Furthermore, the prediction image correction process from one reference picture has been described here, but the same applies to the case of correcting the prediction image from a plurality of reference pictures. After obtaining the prediction images corrected from each reference picture, the obtained prediction images are further overlapped to obtain the final prediction image. Furthermore, the aforementioned processing target block can be in units of prediction blocks or sub-blocks obtained by further dividing the prediction blocks. As a method for determining whether to apply OBMC processing, it includes, for example, a method using a signal obmc_flag indicating whether to apply OBMC processing. A specific example is in an encoding device. It determines whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the obmc_flag is set to value 1, and OBMC processing is applied for encoding. If it does not belong to a region with complex motion, the obmc_flag is set to value 0, and encoding is performed without applying OBMC processing. Additionally, in a decoding device, by decoding the obmc_flag described in the stream, decoding is performed while switching whether to apply OBMC processing according to its value. Furthermore, the motion information may not be signaled and may be derived from the decoding device side. For example, the merge mode specified in the H.265 / HEVC standard may also be used. Also, for example, motion information may be derived by performing motion estimation on the decoding device side. At this time, motion estimation is performed without using the pixel values of the current block. Here, a mode of performing motion estimation on the decoding device side will be described. This mode of performing motion estimation on the decoding device side is sometimes referred to as the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode. An example of FRUC processing is shown in FIG. 5D. First, referring to the motion vectors of the encoded blocks adjacent to the current block spatially or temporally, a plurality of candidate lists each having a motion vector predictor are generated (which may also be common to the merge list). Then, from among the plurality of candidate MVs registered in the candidate list, the best candidate MV is selected. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected according to the evaluation value. Then, based on the selected candidate motion vector, the motion vector for the current block is derived. Specifically, for example, the selected candidate motion vector (the best candidate MV) is directly derived as the motion vector for the current block. Also, for example, pattern matching may be performed in the peripheral region of the position in the reference picture corresponding to the selected candidate motion vector, thereby deriving the motion vector for the current block. That is, for the peripheral region of the best candidate MV, estimation is performed in the same manner. If there is an MV with a better evaluation value, the best candidate MV is updated to the aforementioned MV, and it may be used as the final MV for the current block. Furthermore, a configuration that does not perform this processing may also be adopted. When processing is performed in sub-block units, exactly the same processing may also be adopted. Furthermore, the evaluation value is calculated from the difference value of the reconstructed image obtained by pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Furthermore, in addition to the difference value, information other than this may also be used to calculate the evaluation value. Pattern matching utilizes first pattern matching and second pattern matching. The first pattern matching and the second pattern matching are sometimes respectively referred to as bilateral matching and template matching. In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures and between two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, the area in other reference pictures along the motion trajectory of the current block is used as the predetermined area for calculating the above-mentioned candidate evaluation value. FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating two blocks along the motion trajectory of the current block (Cur block) and the most matching pair among the pairs of two blocks in two different reference pictures (Ref0, Ref1). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV is derived, and the obtained difference value is used to calculate the evaluation value. The symmetric MV is obtained by scaling the candidate MV by the display time interval. The candidate MV with the best evaluation value among multiple candidate MVs is selected as the final MV. Assuming a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the time distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is located in time between the two reference pictures and the time distances from the current picture to the two reference pictures are equal, in the first pattern matching, a reflection-symmetric bilateral motion vector is derived. In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (such as the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as the predetermined area for calculating the above-mentioned candidate evaluation value. FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of the current block is derived by estimating, in the reference picture (Ref0), a block that best matches a block adjacent to the current block in the current picture (Cur Pic). Specifically, for the current block, a difference is obtained between a reconstructed image of an encoded region on either or both of the left and upper adjacent sides and a reconstructed image at the same position in the encoded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and a candidate MV having the best evaluation value among a plurality of candidate MVs is selected as the best candidate MV. Information indicating whether or not to apply this type of FRUC mode (for example, referred to as an FRUC flag) is signaled at the CU level. Further, when applying the FRUC mode (for example, when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (for example, referred to as an FRUC flag) is signaled at the CU level. Furthermore, the signaling of such information is not limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, block level, CTU level, or sub-block level). Here, a mode for deriving a motion vector based on a model assuming uniform linear motion is described. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode. FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x , v y ) represents a velocity vector, and τ 0 , τ 1 respectively represent time distances between the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) represents a motion vector corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) represents a motion vector corresponding to the reference picture Ref 1 The motion vector. At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1 , MVy 1 ) are respectively expressed as (vxτ 0 , vyτ 0 ) and (-vxτ 1 , -vyτ 1 ), and the following optical flow equation (1) holds. [Equation 1] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i), (ii), and (iii) is equal to zero, where (i) is the time differential of the luminance value, (ii) is the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) is the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the merge list, etc. is corrected in pixel units. Furthermore, the motion vector can also be derived on the decoder side by a method different from the derivation of the motion vector according to the model assuming uniform linear motion. For example, the motion vector can also be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks. Here, a mode of deriving the motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks is described. This mode is sometimes referred to as the affine motion compensation prediction mode. FIG. 9A is a diagram for explaining the derivation of the motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes sixteen 4×4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v of the control point at the upper left corner of the current block is derived 0, the motion vector of the upper-right control point of the current block is derived based on the motion vectors of adjacent sub-blocks as v 1 . Then, using the two motion vectors v 0 and v 1 , the motion vectors (v x , v y ) of each sub-block within the current block are derived by the following formula (2). [Equation 2] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a predetermined weighting coefficient. In this type of affine motion compensation prediction mode, several modes with different methods for deriving the motion vectors of the upper-left and upper-right control points may also be included. Information indicating this type of affine motion compensation prediction mode (such as an affine flag) is signaled at the CU level. Furthermore, the signaling of the information indicating this affine motion compensation prediction mode need not be limited to the CU level and may also be at other levels (such as the sequence level, picture level, slice level, block level, CTU level, or sub-block level). [Prediction control unit] The prediction control unit 128 selects either the intra-frame prediction signal or the inter-frame prediction signal within the frame and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116. Here, an example of deriving the motion vector of the coded picture by the merge mode is described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process using the merge mode. First, a prediction MV list in which candidates of the prediction MV are registered is generated. The candidates of the prediction MV include: a spatial neighboring prediction MV, which is the MV possessed by a plurality of coded blocks spatially adjacent to the coded object block; a temporal neighboring MV, which is the MV possessed by a block near the position of the coded block projecting the coded reference picture; a combined prediction MV, which is an MV generated by combining the MV values of the spatial neighboring prediction MV and the temporal neighboring MV; and an MV with a value of zero, that is, a zero prediction MV, etc. Then, by selecting one prediction MV from among the plurality of prediction MVs registered in the prediction MV list, the MV of the coded object block is determined. Furthermore, in the variable length coding unit, a signal indicating which prediction MV is selected, that is, merge_idx, is described in the stream and coded. Furthermore, the predicted MVs registered in the predicted MV list illustrated in FIG. 9B are just examples. The number of predicted MVs may be different from the number in the figure, or it may be a composition that does not include some types of the predicted MVs in the figure, or it may be a composition with predicted MVs other than the types of predicted MVs in the figure added. Furthermore, the final MV can be determined by performing the following-described DMVR processing on the MVs of the coding target blocks derived using the merge mode. Here, an example of determining the MV using the DMVR processing will be described. FIG. 9C is a conceptual diagram for explaining the outline of the DMVR processing. First, use the best MVP set for the processing target block as the candidate MV. According to the aforementioned candidate MV, obtain reference pixels from the first reference picture of the processed picture in the L0 direction and the second reference picture of the processed picture in the L1 direction respectively, and take the average of each reference pixel to generate a template. Next, use the aforementioned template to estimate the peripheral areas of the candidate MVs of the first reference picture and the second reference picture respectively, and determine the MV with the minimum cost as the final MV. Furthermore, the cost value is calculated using the difference values and MV values of each pixel value of the template and each pixel value of the estimated area. Furthermore, in the coding device and the decoding device, the outline of the processing described here is basically common. Furthermore, instead of using the processing described here itself, other processing can also be used as long as it can estimate the periphery of the candidate MV and derive the final MV. Here, a mode of generating a predicted picture using the LIC processing will be described. FIG. 9D is a diagram for explaining the outline of a method for generating a predicted picture of a luminance correction process using the LIC processing. First, derive the MV for obtaining the reference picture corresponding to the coding target block from the reference picture of the coded picture. Next, for the coding target block, use the luminance pixel values of the left adjacent and upper adjacent coded peripheral reference areas and the luminance pixel values at the same position in the reference picture specified by the MV to extract the information indicating how the luminance value changes in the reference picture and the coding target picture, and calculate the luminance correction parameter. Perform a luminance correction process on the reference picture in the reference picture specified by the MV using the aforementioned luminance correction parameter, thereby generating a predicted picture for the coding target block. Furthermore, the shape of the aforementioned peripheral reference area in FIG. 9D is just an example, and shapes other than this shape can also be used. Also, although the processing of generating a predicted picture from one reference picture has been described here, the same applies to the case of generating a predicted picture from a plurality of reference pictures. After performing the luminance correction process on the reference pictures obtained from each reference picture in the same way, a predicted picture is generated. As a method for determining whether to apply LIC processing, it includes, for example, a method using a signal lic_flag indicating whether LIC processing is applicable. A specific example is in an encoding device. It is determined whether an encoding target block belongs to an area where a luminance change occurs. When it belongs to an area where a luminance change occurs, the lic_flag is set to 1, and LIC processing is applied for encoding. When it does not belong to an area where a luminance change occurs, the lic_flag is set to 0, and encoding is performed without applying LIC processing. Further, in a decoding device, by means of the lic_flag described in the decoded stream, decoding is performed by switching whether to apply LIC processing according to its value. As a method for determining whether to apply LIC processing, it also includes, for example, a method of determining according to whether peripheral blocks are applicable for LIC processing. As a specific example, when the encoding target block is in a merge mode, it is determined whether the peripheral encoded completed blocks selected at the time of MV derivation in the merge mode processing have been encoded by applying LIC processing, and encoding is performed by switching whether to apply LIC processing according to the result. Furthermore, in the case of this example, the decoding process is exactly the same. [Outline of the decoding device] Next, an outline of a decoding device that can decode the encoded signal (encoded bitstream) output from the above encoding device 100 will be described. FIG. 10 is a block diagram showing the functional configuration of the decoding device 200 according to Embodiment 1. The decoding device 200 is a dynamic image / image decoding device that decodes dynamic images / images in block units. As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. At this time, when the processor executes a software program stored in the memory, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filtering unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. Also, the decoding device 200 can also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filtering unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. The following describes each constituent element included in the decoding device 200. [Entropy decoding unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202 arithmetic-decodes the encoded bit stream into a binary signal, for example. Then, the entropy decoding unit 202 de-binarizes the binary signal. Thereby, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the input from the entropy decoding unit 202, that is, the quantization coefficients of the block to be decoded (hereinafter referred to as the current block). Specifically, the inverse quantization unit 204 inverse-quantizes each quantization coefficient of the current block according to the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206. [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the input from the inverse quantization unit 204, that is, the transform coefficients. For example, when the information read from the encoded bit stream is applicable to EMT or AMT (e.g., when the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block according to the information indicating the read transform type. Also, for example, when the information read from the encoded bit stream is applicable to NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients. [Addition Unit] The addition unit 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction sample input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212. [Block Memory] The block memory 210 is a memory unit for storing the blocks referred to in intra-frame prediction and is a block within the picture to be decoded (hereinafter referred to as the current picture). Specifically, the block memory 210 stores the reconstructed block output from the addition unit 208. [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214, the display device, etc. The information indicating the ON / OFF of ALF read from the encoded bit stream indicates that when ALF is ON, one filter is selected from a plurality of filters according to the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block. [Frame Memory] The frame memory 214 is a memory unit for storing the reference pictures used in inter-frame prediction and is sometimes also referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212. [Intra-Frame Prediction Unit] The intra-frame prediction unit 216 performs intra-frame prediction by referring to the blocks in the current picture stored in the block memory 210 according to the intra-frame prediction mode decoded from the encoded bitstream, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 216 refers to the samples (e.g., luminance values, chrominance differences) of the blocks adjacent to the current block to perform intra-frame prediction, thereby generating an intra-frame prediction signal and outputting the intra-frame prediction signal to the prediction control unit 220. Furthermore, in the intra-frame prediction of the chrominance blocks, when selecting the intra-frame prediction mode of the reference luminance block, the intra-frame prediction unit 216 can also predict the chrominance component of the current block according to the luminance component of the current block. Also, when the information decoded from the encoded bitstream indicates that PDPC is applicable, the intra-frame prediction unit 216 corrects the pixel value after intra-frame prediction according to the gradients of the reference pixels in the horizontal / vertical directions. [Inter-frame prediction unit] The inter-frame prediction unit 218 predicts the current block by referring to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 218 performs motion compensation by using the motion information (e.g., motion vector) decoded from the encoded bitstream, thereby generating an inter-frame prediction signal for the current block or sub-block and outputting the inter-frame prediction signal to the prediction control unit 220. Furthermore, when the information decoded from the encoded bitstream indicates that the OBMC mode is applicable, the inter-frame prediction unit 218 can generate an inter-frame prediction signal by using not only the motion information of the current block obtained by motion estimation but also the motion information of the adjacent blocks. Also, when the information decoded from the encoded bitstream indicates that the FRUC mode is applicable, the inter-frame prediction unit 218 performs motion estimation according to the mode matching method (bidirectional matching or template matching) decoded from the encoded bitstream, thereby deriving motion information. Then, the inter-frame prediction unit 218 uses the derived motion information to perform motion compensation. Also, when the BIO mode is applicable, the inter-frame prediction unit 218 derives a motion vector according to the model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is applicable, the inter-frame prediction unit 218 derives a motion vector in units of sub-blocks according to the motion vectors of a plurality of adjacent blocks. [Prediction control unit] The prediction control unit 220 selects either the intra-frame prediction signal or the inter-frame prediction signal and outputs the selected signal as the prediction signal to the addition unit 208. (Embodiment 2) Next, Embodiment 2 will be described. In the aspect of this embodiment, conversion and inverse conversion will be described in detail. Furthermore, since the configurations of the encoding device and the decoding device in this embodiment are substantially the same as those in Embodiment 1, illustration and description thereof are omitted. [Processing of the conversion unit and quantization unit of the encoding device] First, with reference to FIG. 11, the processing of the conversion unit 106 and the quantization unit 108 of the encoding device 100 in this embodiment will be specifically described. FIG. 11 is a flowchart showing the conversion and quantization processing of the encoding device 100 in Embodiment 2. First, the conversion unit 106 selects a first conversion basis for the block to be encoded from candidates of the first conversion basis of 1 or more (S101). For example, the conversion unit 106 fixedly selects the conversion basis of DCT-II as the first conversion basis for the block to be encoded. Also, for example, the conversion unit 106 may select the first conversion basis using an adaptive basis selection mode. The adaptive basis selection mode is a mode in which a conversion basis is adaptively selected from a plurality of pre-determined conversion basis candidates according to cost, and the cost is based on the difference between the original image and the reconstructed image and / or the amount of code. This adaptive basis selection mode is sometimes referred to as the EMT mode or the AMT mode. The plurality of conversion basis candidates may use, for example, the plurality of conversion bases shown in FIG. 6. Furthermore, the plurality of conversion basis candidates are not limited to the plurality of conversion bases in FIG. 6. The plurality of conversion basis candidates may, for example, also include a conversion basis equivalent to not performing conversion. By encoding the encoding identification information in the bitstream, the adaptive basis selection mode and the basis fixed mode can be selectively adopted, and the identification information indicates which of the adaptive basis selection mode and the basis fixed mode using a fixed conversion basis (for example, the basis of DCT of type II) is an effective mode. This identification information corresponds to the identification information indicating whether the adaptive basis selection mode is effective. At this time, it may be possible to determine whether the first conversion basis is consistent with the predetermined conversion basis based on the identification information. For example, in EMT, there is identification information (emt_cu_flag) indicating which of the adaptive basis selection mode and the basis fixed mode is effective in units such as CU, so using this identification information, it is possible to determine whether the first conversion basis is consistent with the predetermined conversion basis. Then, the conversion unit 106 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected in step S102 (S102). The first conversion corresponds to a primary conversion. The conversion unit 106 determines whether the first conversion basis selected in step S101 is the same as the predetermined conversion basis (S103). For example, the conversion unit 106 determines whether the first conversion basis is the same as any one of a plurality of predetermined conversion bases. Also, for example, the conversion unit 106 determines whether the first conversion basis is the same as one predetermined conversion basis. The predetermined conversion basis may employ, for example, a conversion basis of DCT of type II (i.e., DCT-II), and / or a conversion basis similar thereto. Such a predetermined conversion basis may also be predefined in a standard mode or the like. Also, for example, the predetermined conversion basis may be determined according to encoding parameters or the like. Here, when the first conversion basis is the same as the predetermined conversion basis (S103, YES), the conversion unit 106 selects, from candidates of the second conversion basis of 1 or more, the second conversion basis for the block to be encoded (S104). The conversion unit 106 generates second conversion coefficients (S105) by performing a second conversion on the first conversion basis using the selected second conversion basis. The second conversion is equivalent to a secondary conversion. The quantization unit 108 quantizes the generated second conversion coefficients (S106), and ends the conversion and quantization processing. For the second conversion, a secondary conversion called NSST, or a conversion that selectively uses any one of candidates of a plurality of second conversion bases may be performed. At this time, for the selection of the second conversion basis, the selected conversion basis may also be fixed. In short, a predetermined fixed conversion basis may also be selected as the second conversion basis. Also, a conversion basis equivalent to not performing the second conversion may be used as the second conversion basis. Also, NSST may be a frequency space conversion after DCT or DST. For example, NSST represents the KLT (Karhunen Loveve Transform (K-L transform)) of the conversion coefficients for DCT or DST obtained offline, or a basis equivalent to KLT, and a HyGT (Hypercube-Givens Transform) represented by a combination of rotation conversions may also be used. On the other hand, when the first conversion basis is different from the predetermined conversion basis (S103, NO), the conversion unit 106 skips the selection step (S104) of the second conversion basis and the second conversion step (S105). In short, the conversion unit 106 does not perform the second conversion. At this time, the first conversion coefficients generated in step S207 are quantized (S106), and the conversion and quantization processing ends. Thus, when skipping the second conversion step, information indicating that the second conversion is not performed may be notified to the decoding device. Also, when skipping the second conversion step, a second conversion may be performed using a second conversion basis equivalent to not performing the conversion, and information indicating the second conversion basis may be notified to the decoding device. Furthermore, the inverse quantization unit 112 and the inverse transform unit 114 of the encoding device 100 can reconstruct the block to be encoded by performing processes opposite to those of the transform unit 106 and the quantization unit 108. [Processes of the inverse quantization unit and the inverse transform unit of the decoding device] Next, with reference to FIG. 12, the processes of the inverse quantization unit 204 and the inverse transform unit 206 of the decoding device 200 of the present embodiment will be specifically described. FIG. 12 is a flowchart showing the inverse quantization and inverse transform processes of the decoding device 200 of Embodiment 2. First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S601). The inverse transform unit 206 determines whether the first inverse transform basis for the block to be decoded is the same as a predetermined inverse transform basis (S602). The predetermined inverse transform basis is an inverse transform basis corresponding to the predetermined transform basis used by the encoding device 100. When the first inverse transform basis is the same as the predetermined inverse transform basis (S602, YES), the inverse transform unit 206 selects the second inverse transform basis for the block to be decoded (S603). Selecting the inverse transform basis (the first inverse transform basis or the second inverse transform basis) in the decoding device 200 means determining the inverse transform basis according to predetermined information. The predetermined information can be, for example, a basis selection signal. Also, the predetermined information can be the intra prediction mode or the block size within the frame, etc. The inverse transform unit 206 generates second inverse transform coefficients by performing a second inverse transform on the inverse quantized coefficients of the block to be decoded using the selected second inverse transform basis (S604). Further, the inverse transform unit 206 selects the first inverse transform basis (S605). The inverse transform unit 206 performs a first inverse transform on the second inverse transform coefficients generated in step S605 using the selected first inverse transform basis (S606), and ends the inverse quantization and inverse transform processes. On the other hand, when the first inverse transform basis is different from the predetermined inverse transform basis (S602, NO), the inverse transform unit 206 skips the selection step (S603) of the second inverse transform basis and the second inverse transform step (S604). In short, the inverse transform unit 206 does not perform the second inverse transform but selects the first inverse transform basis (S605). The inverse transform unit 206 performs a first inverse transform on the coefficients inverse quantized in step S501 using the selected first inverse transform basis (S606), and ends the inverse quantization and inverse transform processes. [Effects, etc.] The inventor has found that in the conventional encoding, there is a problem of a large amount of processing for estimating the optimal combination of the conversion basis and conversion parameters (such as the coefficients of the filter) for both the first conversion and the second conversion. In response to this, if the encoding device 100 and the decoding device 200 according to the present embodiment are used, the second conversion basis can be skipped in response to the first conversion basis. As a result, the processing for estimating the optimal combination of the conversion basis and conversion parameters for both the first conversion and the second conversion can be reduced, the reduction in compression efficiency can be suppressed, and the processing load can be reduced. As described above, if the encoding device 100 and the decoding device 200 according to the present embodiment are used, when the first conversion basis is different from the predetermined conversion basis, the second conversion can be skipped. The first conversion coefficients generated by the first conversion are affected by the first conversion basis. Therefore, the effect of improving the compression rate obtained by performing the second conversion on the first conversion coefficients often depends on the first conversion basis. Here, when the first conversion basis is different from the predetermined conversion basis with a high compression rate improvement effect, by skipping the second conversion, the reduction in compression efficiency can be suppressed, and the processing load can be reduced. Especially in the case of type II DCT, since the concentration in the low-frequency region often increases, the possibility that the effect of the second conversion becomes high is high. Therefore, when the compression efficiency improvement effect of the second conversion is high by using the basis of type II DCT as the predetermined conversion basis, the second conversion is performed, and in other cases, the second conversion is skipped, whereby it is possible to expect to suppress the reduction in compression efficiency and reduce the processing load. Furthermore, the above processing can be applied to either the luminance signal or the color difference signal. If the input signal is in RGB format, it can also be applied to each of the R, G, and B signals. Furthermore, in the luminance signal and the color difference signal, the bases that can be selected in the first conversion or the second conversion can also be different. For example, the frequency band region of the luminance signal is wider than that of the color difference signal. Therefore, in the conversion of the luminance signal, a basis with more types can also be selected than in the color difference signal. Furthermore, the predetermined conversion basis is not limited to one conversion basis. In short, the predetermined conversion basis can also be a plurality of conversion bases. In this case, it is only necessary to determine whether the first conversion basis is consistent with any of the plurality of predetermined conversion bases. Furthermore, this aspect can also be implemented in combination with at least a part of other aspects of the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the grammar, etc. can be combined with other aspects and implemented. (Embodiment 3) Next, Embodiment 3 will be described. The difference between the aspect of this embodiment and the above-described Embodiment 2 is that the conversion process varies depending on whether intra-frame prediction is used for the coding / decoding target block. Hereinafter, for this embodiment, with reference to the drawings, the differences from the above-described Embodiment 2 will be mainly described. Furthermore, in the following figures, the same reference numerals are assigned to steps that are substantially the same as those in Embodiment 2, and repeated descriptions are omitted or simplified. [Processing of the conversion unit and quantization unit of the encoding device] First, with reference to FIG. 13, the processing of the conversion unit 106 and the quantization unit 108 of the encoding device 100 in this embodiment will be specifically described. FIG. 13 is a flowchart showing the conversion and quantization processing of the encoding device 100 according to Embodiment 3. First, the conversion unit 106 determines whether intra-frame prediction or inter-frame prediction is used for the coding target block (S201). For example, the conversion unit 106 determines which of intra-frame prediction and inter-frame prediction to use based on cost, and the aforementioned cost is based on the difference between the original image and the reconstructed image obtained from the partial decoded compressed image and / or the code amount. Also, for example, the conversion unit 106 may also determine which of intra-frame prediction and inter-frame prediction to use based on information different from the cost based on the difference and / or code amount (e.g., picture type). Here, when it is determined that inter-frame prediction is used for the coding target block (S201, inter-frame), the conversion unit 106 selects the first conversion basis for the coding target block from candidates of the first conversion basis of 1 or more (S202). For example, the conversion unit 106 fixedly selects the conversion basis of DCT-II as the first conversion basis for the coding target block. Also, for example, the conversion unit 106 may also select the first conversion basis from a plurality of candidates of the first conversion basis. Then, the conversion unit 106 performs a first conversion on the residual of the coding target block by using the first conversion basis selected in step S202 to generate first conversion coefficients (S203). The quantization unit 108 quantizes the generated first conversion coefficients (S204), and ends the conversion and quantization processing. On the other hand, when it is determined that intra-frame prediction is used for the coding target block (S201, intra-frame), the conversion unit 106 executes steps S101 to S105 in the same manner as in Embodiment 2. Then, the quantization unit 108 quantizes the first conversion coefficients generated in step S102 or the second conversion coefficients generated in step S105 (S204), and ends the conversion and quantization processing. [Processing of the inverse quantization unit and inverse conversion unit of the decoding device] Next, with reference to FIG. 14, the processing of the inverse quantization unit 204 and the inverse conversion unit 206 of the decoding device 200 in this embodiment will be specifically described. FIG. 14 is a flowchart showing the inverse quantization and inverse conversion processing of the decoding device 200 according to Embodiment 3. First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S601). The inverse transform unit 206 determines whether intra-frame prediction or inter-frame prediction is to be used for the block to be decoded (S701). For example, the inverse transform unit 206 determines whether intra-frame prediction or inter-frame prediction is to be used based on the information obtained from the bitstream. When it is determined that inter-frame prediction is to be used for the block to be decoded (S701, inter-frame), the inverse transform unit 206 selects the first inverse transform basis for the block to be decoded (S702). The inverse transform unit 206 performs the first inverse transform on the inverse quantized coefficients of the block to be decoded using the first inverse transform basis selected in step S503 (S703), and ends the inverse quantization and inverse transform processing. On the other hand, when it is determined that intra-frame prediction is to be used for the block to be decoded (S701, intra-frame), the inverse transform unit 206 executes steps S602 to S606 in the same manner as in Embodiment 2, and ends the inverse quantization and inverse transform processing. [Effects, etc.] According to the encoding device 100 and the decoding device 200 of the present embodiment, the second transform can be skipped in response to intra-frame / inter-frame prediction and the first transform basis. As a result, the process of estimating the optimal combination of the transform basis and the transform parameters for both the first transform and the second transform can be reduced, the reduction in compression efficiency can be suppressed, and the processing load can be reduced. Furthermore, this aspect can also be implemented in combination with at least a part of other aspects of the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the grammar, etc. can be implemented in combination with other aspects. (Embodiment 4) Next, Embodiment 4 will be described. The aspect of the present embodiment is different from those of the above-described Embodiments 2 and 3 in that the transform processing differs according to the intra-frame prediction mode for the encoding / decoding target block. Hereinafter, with reference to the drawings, the present embodiment will be described centering on the differences from the above-described Embodiments 2 and 3. Furthermore, in the following drawings, the same reference numerals are assigned to steps that are substantially the same as those in Embodiment 2 or 3, and redundant explanations are omitted or simplified. [Processing of the transform unit and the quantization unit of the encoding device] First, with reference to FIG. 15, the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 of the present embodiment will be specifically described. FIG. 15 is a flowchart showing the transform and quantization processing of the encoding device 100 of Embodiment 4. Similar to Embodiment 2, the conversion unit 106 determines whether to use intra-frame prediction or inter-frame prediction for the coding target block (S201). Here, when it is determined to use inter-frame prediction for the coding target block (S201, inter-frame), the conversion unit 106 executes step S202 and step S203 in the same manner as in Embodiment 2. Further, the quantization unit 108 quantizes the first conversion coefficient generated in step S203 (S302). On the other hand, when it is determined to use intra-frame prediction for the coding target block (S201, intra-frame), the conversion unit 106 executes step S101 and step S102 in the same manner as in Embodiment 1. Then, the conversion unit 106 determines whether the intra-frame prediction mode of the coding target block is a predetermined mode (S106). For example, the conversion unit 106 determines whether the intra-frame prediction mode adopted is a predetermined mode based on the cost, and the aforementioned cost is based on the difference between the original image and the reconstructed image and / or the code amount. Furthermore, the determination of whether the intra-frame prediction mode is a predetermined mode can also be performed based on information different from the cost. The predetermined mode can also be defined by, for example, a standard specification or the like. Also, for example, the predetermined mode can be determined based on coding parameters or the like. The predetermined mode can also adopt, for example, an oblique directional prediction mode. The directional prediction mode is a prediction of the coding target block, using an intra-frame prediction mode in a specific direction. In the directional prediction mode, the pixel value is predicted by extending the value of the reference pixel in a specific direction. Furthermore, the pixel value is the value of the pixel unit constituting the picture, such as a luminance value or a color difference value. For example, the directional prediction mode is an intra-frame prediction mode other than the DC prediction mode and the Planar prediction mode. The oblique directional prediction mode is a directional prediction mode having a direction inclined with respect to the horizontal direction and the vertical direction. For example, the oblique directional prediction mode can also be a 3-directional prediction mode identified by 2 (lower left), 34 (upper left), and 66 (upper right) in the 65-directional prediction mode that is sequentially identified by numbers 2 to 66 from the lower left to the upper right (refer to FIG. 5A). Also, for example, the oblique directional prediction mode can also be a 7-directional prediction mode identified by 2 to 3 (lower left), 33 to 35 (upper left), and 65 to 66 (upper right) in the 65-directional prediction mode. When the intra-frame prediction mode is not a predetermined mode (S301, NO), the conversion unit 106 determines whether the first conversion basis selected in step S101 is consistent with a predetermined conversion basis (S103). When the prediction mode within the frame is the predetermined mode (S301, YES), or when the first conversion basis matches the predetermined conversion basis (S103, YES), the conversion unit 106 selects, from among candidates for the second conversion basis of 1 or more, the second conversion basis for the block to be encoded (S104). The conversion unit 106 generates the second conversion coefficients by performing a second conversion on the first conversion coefficients using the selected second conversion basis (S105). The quantization unit 108 quantizes the generated second conversion coefficients (S302), and ends the conversion and quantization processing. When the prediction mode within the frame is different from the predetermined mode (S301, NO), and the first conversion basis is different from the predetermined conversion basis (S103, NO), the conversion unit 106 skips the step of selecting the second conversion basis (S104) and the second conversion step (S105). In short, the conversion unit 106 does not perform the second conversion. At this time, the first conversion coefficients generated in step S102 are quantized (S302), and the conversion and quantization processing ends. [Processing of the inverse quantization unit and inverse conversion unit of the decoding device] Next, with reference to FIG. 16, the processing of the inverse quantization unit 204 and inverse conversion unit 206 of the decoding device 200 of the present embodiment will be specifically described. FIG. 16 is a flowchart showing the inverse quantization and inverse conversion processing of the decoding device 200 of Embodiment 4. First, the inverse quantization unit 204 inverse quantizes the quantization coefficients of the block to be decoded (S601). The inverse conversion unit 206 determines whether intra-frame prediction or inter-frame prediction is used for the block to be decoded (S701). When it is determined that inter-frame prediction is used for the block to be decoded (S701, inter-frame), the inverse conversion unit 206 executes steps S702 and S703 in the same manner as in Embodiment 3, and ends the inverse quantization and inverse conversion processing. Alternatively, when it is determined that intra-frame prediction is used for the block to be decoded (S701, intra-frame), the inverse conversion unit 206 determines whether the intra-frame prediction mode of the block to be decoded is the predetermined mode (S801). The predetermined mode used by the decoding device 200 is the same as the predetermined mode used by the encoding device 100. When the intra-frame prediction mode is not the predetermined mode (S801, NO), the inverse conversion unit 206 determines whether the first inverse conversion basis for the block to be decoded matches the predetermined inverse conversion basis (S602). When the intra-frame prediction mode is the predetermined mode (S801, YES), or when the first inverse conversion basis matches the predetermined inverse conversion basis (S602, YES), steps S603 to S606 are executed in the same manner as in Embodiment 2, and the inverse quantization and inverse conversion processing ends. Further, when the in-frame prediction mode is different from the predetermined mode (S801, NO), and the first inverse transform basis is different from the predetermined inverse transform basis (S602, NO), the inverse transform unit 206 skips the selection step (S603) and the second inverse transform step (S604) of the second inverse transform basis. In short, the inverse transform unit 206 does not perform the second inverse transform and selects the first inverse transform basis (S605). The inverse transform unit 206 uses the selected first inverse transform basis to perform the first inverse transform on the coefficients that have been inverse quantized in step S501 (S606), and ends the inverse quantization and inverse transform processing. [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 of the present embodiment, the second transform can be skipped in response to the in-frame prediction mode and the first transform basis. As a result, the process of estimating the optimal combination of the transform basis and the transform parameters for both the first transform and the second transform can be reduced, the reduction of the compression efficiency can be suppressed, and the processing load can be reduced. In particular, if an oblique directional prediction mode is adopted as the predetermined mode, the second transform can be performed when the oblique directional prediction mode is adopted for the encoding / decoding target block, and the second transform can be skipped in other cases. Thereby, the reduction of the compression efficiency can be suppressed, and the processing load can be reduced. Generally speaking, for the first transform, a DCT or DST that can be separated into a vertical direction and a horizontal direction is performed. At this time, the oblique correlation is not used in the first transform. Therefore, when an oblique directional prediction mode with a high oblique correlation is adopted, it is difficult to sufficiently concentrate the coefficients only with the first transform. Therefore, when the oblique directional prediction mode is adopted for in-frame prediction, by performing the second transform using the second transform basis that utilizes the oblique correlation, the coefficients can be further concentrated, and the compression efficiency can be improved. Furthermore, the order of the steps in the flowcharts of FIGS. 15 and 16 is not limited to the order described in FIGS. 15 and 16. For example, in FIG. 15, the determination step (S801) of whether the in-frame prediction mode is the predetermined mode and the determination step (S602) of whether the first transform basis is consistent with the predetermined transform basis can be in the reverse order or performed simultaneously. Furthermore, this aspect can also be implemented in combination with at least a part of other aspects of the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the grammar, etc. can be combined with other aspects and implemented. (Embodiment 5) Next, Embodiment 5 will be described. In the aspect of this embodiment, encoding / decoding of information related to conversion / inverse conversion will be described. Hereinafter, for this embodiment, with reference to the drawings, the description will be centered on the points different from those of the above-described Embodiments 2 to 4. Further, in this embodiment, since the conversion and quantization processing, and the inverse quantization and inverse conversion processing are substantially the same as those of the above-described Embodiment 4, the description thereof will be omitted. [Processing of the entropy encoding unit of the encoding device] With reference to FIG. 17, the encoding process of the information of the conversion of the entropy encoding unit 110 of the encoding device 100 according to this embodiment will be specifically described. FIG. 17 is a flowchart showing the encoding process of the encoding device 100 of Embodiment 5. When inter-frame prediction is used for the encoding target block (S401, inter-frame), the entropy encoding unit 110 encodes the first base selection signal into the bit stream (S402). Here, the first base selection signal is information or data indicating the first conversion base selected in step S202 of FIG. 15. Encoding a signal into the bit stream means arranging a code representing the information in the bit stream. The code is generated by, for example, the context adaptive binary arithmetic coding method (CABAC). Further, the generation of the code does not necessarily have to use CABAC, nor does it necessarily have to use entropy encoding. For example, the code can also be the information itself (for example, a flag of 0 or 1). Next, the entropy encoding unit 110 encodes the coefficients quantized in step S302 of FIG. 15 (S403), and ends the encoding process. When intra-frame prediction is used for the encoding target block (S401, intra-frame), the entropy encoding unit 110 encodes the intra-frame prediction mode signal indicating the intra-frame prediction mode of the encoding target block into the bit stream (S404). Further, the entropy encoding unit 110 encodes the first base selection signal into the bit stream (S405). Here, the first base selection signal is information or data indicating the first conversion base selected in step S101 of FIG. 15. Here, when the second conversion is performed (S406, YES), the entropy encoding unit 110 encodes the second base selection signal into the bit stream (S407). Here, the second base selection signal is information or data indicating the second conversion base selected in step S104. Further, when the second conversion is not performed (S406, NO), the entropy encoding unit 110 skips the encoding step of the second base selection signal (S407). In short, the entropy encoding unit 110 does not encode the second base selection signal. Finally, the entropy encoding unit 110 encodes the coefficients quantized in step S302 (S408), and ends the encoding process. [Syntax] FIG. 18 is a specific example showing the syntax of Embodiment 5. In FIG. 18, the prediction mode signal (pred_mode), the in-frame prediction mode signal (pred_mode_dir), and the adaptive selection mode signal (emt_mode), and as required, the first basis selection signal (primary_transform_type) and the second basis selection signal (secondary_transform_type) are encoded in the bitstream. The prediction mode signal (pred_mode) indicates whether in-frame prediction or inter-frame prediction is used for the coding / decoding target block (here, the coding unit). The inverse transform unit 206 of the decoding device 200 can determine whether to use in-frame prediction for the decoding target block based on this prediction mode signal. The in-frame prediction mode signal (pred_mode_dir) indicates the in-frame prediction mode when in-frame prediction is used for the coding / decoding target block. The inverse transform unit 206 of the decoding device 200 can determine whether the in-frame prediction mode of the decoding target block is a predetermined mode based on this in-frame prediction mode signal. The adaptive selection mode signal (emt_mode) indicates whether to use the adaptive basis selection mode of adaptively selecting a transform basis from among candidates of a plurality of transform bases for the coding / decoding target block. Here, when the adaptive selection mode signal is "ON (on)", a transform basis is selected from among DCT of type V, DCT of type VIII, DST of type I, and DST of type VII. Conversely, when the adaptive selection mode signal is "OFF (off)", DCT of type II is selected. The inverse transform unit 206 of the decoding device 200 can determine whether the first inverse transform basis of the decoding target block is the same as a predetermined inverse transform basis based on this adaptive selection mode signal. The first basis selection signal (primary_transform_type) indicates the first transform basis / inverse transform basis used for the transform / inverse transform of the coding / decoding target block. The first basis selection signal is encoded in the bitstream when the adaptive selection mode signal is "ON". Conversely, when the adaptive selection mode signal is "OFF", the first basis selection signal is not encoded. The inverse transform unit 206 of the decoding device 200 can select the first inverse transform basis based on this first basis selection signal. The second base selection signal (secondary_transform_type) indicates the second transform base / inverse transform base used for the transform / inverse transform of the coding / decoding target block. The second base selection signal is encoded in the bitstream when the adaptive selection mode signal is "ON" and the in-frame prediction mode signal is "2", "34", or "66". The "2", "34", and "66" of the in-frame prediction mode signal all indicate the oblique directional prediction mode. In summary, when the first transform base is the same as the base of the type II DCT and the in-frame prediction mode is the oblique directional prediction mode, the second base selection signal is encoded in the bitstream. Conversely, when the in-frame prediction mode is not the oblique directional prediction mode, the second base selection signal is not encoded in the bitstream. The inverse transform unit 206 of the decoding device 200 can select the second inverse transform base according to the second base selection signal. Furthermore, here, the transform bases that can be selected in the adaptive base selection mode adopt the bases of type V DCT, type VIII DCT, type I DST, and type VII DST, but are not limited thereto. For example, type IV DCT can also be used instead of type V DCT. For type IV DCT, since part of the processing of type II DCT can be reused, the processing load can be reduced. Also, type IV DST can be used. For type IV DST, since part of the processing of type IV DCT can be reused, the processing load can be reduced. [Processing of the entropy decoding unit of the decoding device] Next, with reference to FIG. 19, the processing of the entropy decoding unit 202 of the decoding device 200 according to the present embodiment will be specifically described. FIG. 19 is a flowchart showing the decoding process of the decoding device 200 of Embodiment 5. When inter-frame prediction is used for the decoding target block (S901, inter-frame), the entropy decoding unit 202 decodes the first base selection signal from the bitstream (S902). Decoding a signal from the bitstream means reading the code representing the information from the bitstream and restoring the information from the read code. Restoring the information from the code uses, for example, the context-adaptive binary arithmetic decoding method (CABAD). Furthermore, restoring the information from the code does not necessarily have to use CABAD, nor does it necessarily have to use entropy decoding. For example, when the read code itself represents the information (such as a flag of 0 or 1), only the code needs to be decoded. Next, the entropy decoding unit 202 decodes the quantization coefficient from the bitstream (S903), and ends the decoding process. When in-frame prediction is used for the decoding target block (S901, in-frame), the entropy decoding unit 202 decodes the in-frame prediction mode signal from the bitstream (S904). Furthermore, the entropy decoding unit 202 decodes the first base selection signal from the bitstream (S905). Here, when the second conversion is performed (S906, YES), the entropy decoding unit 202 decodes the second base selection signal from the bit stream (S907). On the other hand, when the second conversion is not performed (S906, NO), the entropy decoding unit 202 skips the decoding step of the second base selection signal (S907). In short, the entropy decoding unit 202 does not decode the second base selection signal. Finally, the entropy decoding unit 202 decodes the quantization coefficient from the bit stream (S908), and ends the decoding process. [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 of this embodiment, the first base selection signal and the second base selection signal can be encoded in the bit stream. Then, by encoding the intra-frame prediction mode signal and the first base selection signal earlier than the second base selection signal in the frame, it is possible to determine whether to skip the second inverse conversion before decoding the second base selection signal. Therefore, when the second inverse conversion is skipped, the encoding of the second base selection signal can also be skipped, improving the compression efficiency. (Embodiment 6) Next, Embodiment 6 will be described. In the aspect of this embodiment, the difference from the above-described Embodiment 5 is that the information of the intra-frame prediction mode for performing the second conversion is encoded. Hereinafter, this embodiment will be described centering on the differences from the above-described Embodiment 5 while referring to the drawings. Furthermore, in the following figures, the same reference numerals are attached to steps that are substantially the same as those in Embodiment 5, and repeated descriptions are omitted or simplified. [Processing of the entropy encoding unit of the encoding device] The encoding process of the information of the conversion related to the entropy encoding unit 110 of the encoding device 100 of this embodiment will be specifically described with reference to FIG. 20. FIG. 20 is a flowchart showing the encoding process of the encoding device 100 of Embodiment 6. When inter-frame prediction is used for the encoding target block (S401, inter-frame), the entropy encoding unit 110 executes steps S402 and S403 in the same manner as in Embodiment 5, and ends the encoding process. On the other hand, when intra-frame prediction is used for the encoding target block (S401, intra-frame), the entropy encoding unit 110 encodes the second conversion target prediction mode signal in the bit stream (S501). The second conversion target prediction mode signal represents a predetermined mode for determining whether to perform the second inverse conversion. Specifically, the second conversion target prediction mode signal represents, for example, the numbers of intra-frame prediction modes (for example, 2, 34, and 66). Furthermore, the coding unit of the second conversion target prediction mode signal can also be a CU (Coding Unit) or CTU (Coding Tree Unit), or equivalent to the SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) of the H.265 / HEVC standard, or a slice unit. Thereafter, the entropy encoding unit 110 performs steps S404 to S408 in the same manner as in Embodiment 5, and ends the encoding process. [Processing of the entropy decoding unit of the decoding device] Next, with reference to FIG. 21, the processing of the entropy decoding unit 202 of the decoding device 200 according to this embodiment will be specifically described. FIG. 21 is a flowchart showing the decoding process of the decoding device 200 according to Embodiment 6. When performing inter-frame prediction on the decoding target block (S901, inter-frame), the entropy decoding unit 202 performs steps S902 and S903 in the same manner as in Embodiment 5, and ends the decoding process. Alternatively, when performing intra-frame prediction on the decoding target block (S901, intra-frame), the entropy decoding unit 202 decodes the second conversion target prediction mode signal from the bitstream (S1001). Thereafter, the entropy decoding unit 202 performs steps S904 to S908 in the same manner as in Embodiment 5, and ends the decoding process. [Effects, etc.] As described above, according to the encoding device 100 and the decoding device 200 of this embodiment, the second conversion target prediction mode signal representing a predetermined mode of the intra-frame prediction mode for performing the second conversion / inverse conversion can be encoded in the bitstream. Therefore, the predetermined mode can be arbitrarily determined on the side of the encoding device 100, and the compression efficiency can be improved. Furthermore, the encoding order of each signal can also be predetermined, and various signals can be encoded in an order different from the above encoding order. Furthermore, this aspect can also be implemented in combination with at least a part of other aspects of this disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. can be implemented in combination with other aspects. (Embodiment 7) Various modifications can also be made to the above Embodiments 2 to 6. For example, in each of the above embodiments, the first conversion basis can be fixed according to the size of the encoding / decoding target block. For example, when the block size is smaller than a certain size (e.g., 4×4 pixels, 4×8 pixels, or 8×4 pixels), the first conversion basis is fixed to the conversion basis of DST of type VII, and in this case, the encoding of the first basis selection signal can also be skipped. Also, for example, in each of the above embodiments, a signal indicating whether to skip the selection and the first conversion of the first conversion basis, or the selection and the second conversion of the second conversion basis may be used. For example, if the process of skipping the second conversion is effective, since the second basis selection signal may not be encoded, the decoding operation will be different from the case where the skipping of the second conversion is ineffective. The encoding unit of such a signal may also be in units of CU (Coding Unit) or CTU (Coding Tree Unit), or equivalent to SPS (Sequence Parameter Set) or PPS (Picture Parameter Set) in the H.265 / HEVC standard, or a slice unit. Also, for example, in each of the above embodiments, it is possible to skip the selection and the first conversion of the first conversion basis, or skip the selection and the second conversion of the second conversion basis according to the picture type (I, P, B) or slice type (I, P, B), block size, number of non-zero coefficients, quantization parameter, Temporal_id (hierarchical coding layer). Furthermore, when the encoding device performs the above actions, the decoding device also performs corresponding actions. For example, when information indicating whether to skip the first conversion or the second conversion is encoded, the decoding device decodes this information and determines whether the first or second conversion is effective, and whether the first or second basis selection signal is encoded. Furthermore, in the above Embodiments 5 and 6, a plurality of signals (for example, in-frame prediction mode signal, adaptive selection mode signal, first basis selection signal, and second basis selection signal) are encoded in the bitstream, but in the above Embodiments 2 to 4, these plurality of signals may not be encoded in the bitstream. For example, these plurality of signals may also be notified to the decoding device 200 from the encoding device 100 separately from the bitstream. Furthermore, in this embodiment, the positions of the bits of a plurality of signals (for example, in-frame prediction mode signal, adaptive selection mode signal, first basis selection signal, and second basis selection signal) in the stream are not particularly limited. The plurality of signals are encoded in at least one of a plurality of headers, for example. The plurality of headers may use, for example, video parameter sets, sequence parameter sets, picture parameter sets, and slice headers. Furthermore, when the signals are located in a plurality of hierarchies (for example, picture parameter sets and slice headers), the signals located in the lower hierarchy (for example, slice headers) overwrite the signals located in the higher hierarchy (for example, picture parameter sets). (Embodiment 8) In the above embodiments and various modifications, each of the functional blocks can generally be implemented by an MPU, a memory, etc. Also, the processing of each of the functional blocks is generally implemented by a program execution unit such as a processor, which reads and executes software (program) recorded on a recording medium such as a ROM. This software can be distributed by downloading or the like, or can be distributed by being recorded on a recording medium such as a semiconductor memory. Furthermore, of course, each functional block can also be implemented by hardware (a dedicated circuit). Also, the processing described in the embodiments and various modifications can be implemented by centralized processing using a single device (system) or by distributed processing using a plurality of devices. Also, the number of processors executing the above program can be single or plural. That is, either centralized processing or distributed processing is possible. The present invention is not limited to the above embodiments and can be variously modified, and such modifications are also included in the scope of the present invention. Furthermore, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in the above embodiments and various modifications, and a system using the same will be further described herein. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Regarding other configurations of the system, they can be appropriately changed according to circumstances. [Usage example] FIG. 22 is an overall configuration diagram showing a content supply system ex100 that realizes a content distribution service. The area for providing the communication service is divided into required sizes, and fixed radio stations, i.e., base stations ex106, ex107, ex108, ex109, ex110, are respectively provided in each cell. In this content supply system ex100, machines such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102, a communication network ex104, and base stations ex106 to ex110. This content supply system ex100 can also be connected by combining any of the above elements. Without passing through the fixed radio stations, i.e., base stations ex106 to ex110, the machines can also be directly or indirectly connected to each other via a telephone network or short-range wireless, etc. Also, a streaming server ex103 is connected to machines such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101, etc. Also, the streaming server ex103 is connected to terminal devices, etc. in a hotspot in an airplane ex117 via a satellite ex116. Furthermore, a wireless access point or a hotspot can be used to replace the base stations ex106 to ex110. Also, the streaming server ex103 can be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly connected to the aircraft ex117 without going through the satellite ex116. The camera ex113 is a machine such as a digital camera that can capture still images and animations. Also, the smartphone ex115 is an intelligent device, mobile phone, or PHS (Personal Handyphone System) that generally supports mobile communication system modes such as 2G, 3G, 3.9G, 4G, and what will be called 5G in the future. The home appliance ex118 is a machine included in a refrigerator or a home fuel cell cogeneration system. In the content supply system ex100, a terminal device with a photographing function can be connected to the streaming server ex103 through the base station ex106 etc., to achieve live broadcast etc. In live broadcast, the terminal device (the terminal device in the computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and aircraft ex117, etc.) performs the encoding process described in the above embodiments and each modification on the still image or animation content captured by the user using the terminal device, multiplexes the image data obtained by encoding and the sound data encoded corresponding to the image, and sends the obtained data to the streaming server ex103. That is, each terminal device functions as an image encoding device according to an aspect of the present invention. In addition, the streaming server ex103 performs stream publishing on the content data sent by customers with demands. The customers are the computer ex111, game console ex112, camera ex113, home appliance ex114, smartphone ex115, and the terminal device in the aircraft ex117 etc. that can decode the above encoded data. Each machine that receives the published data decodes and plays the received data. That is, each machine functions as an image decoding device according to an aspect of the present invention. [Distributed processing] Furthermore, the streaming server Ex103 can also be a plurality of servers or a plurality of computers for distributed processing, recording, or publishing of data. For example, the streaming server Ex103 can also be implemented by a CDN (Content Delivery Network). By connecting the networks between many edge servers distributed around the world, content publishing can be achieved. In a CDN, physically proximate edge servers are dynamically assigned according to customers. Then, by caching and publishing content to the edge servers, latency can be reduced. Also, when a certain error occurs, or when the communication state changes due to increased traffic, etc., distributed processing is performed by a plurality of edge servers, or the publishing entity is switched to another edge server, and the network part with a fault can be bypassed to continue publishing. Therefore, high-speed and stable publishing can be achieved. Furthermore, not only the distributed processing of the publishing itself, but the encoding process of the captured data can be performed on each terminal device or on the server side, or can also be shared mutually. As an example, the encoding process generally performs two processing cycles. In the first cycle, the picture complexity or the amount of code of a frame or scene unit is detected. Also, in the second cycle, processing for maintaining the picture quality or improving the encoding efficiency is performed. For example, the terminal device performs the first encoding process, and the server side that receives the content performs the second encoding process. Thereby, the processing load of each terminal device can be reduced, and at the same time, the quality and efficiency of the content can be improved. At this time, when almost instant reception and decoding are required, the data completed by the first encoding performed by the terminal device can be received and played on other terminal devices. Therefore, more flexible real-time communication can also be achieved. As another example, a camera Ex113, etc. extracts feature amounts from an image, compresses the data related to the feature amounts as metadata, and sends it to the server. The server performs compression according to the meaning of the image. For example, from the feature amounts, the importance of an object is judged to switch the quantum precision, etc. Feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction when recompressing on the server. Also, simple encoding such as VLC (Variable Length Coding) can be performed on the terminal device, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed on the server. Furthermore, as another example, in a stadium, a shopping center, a factory, etc., there are sometimes a plurality of image data capturing approximately the same scene by a plurality of terminal devices. At this time, using the plurality of terminal devices performing the photography, other terminal devices not performing the photography as needed, and the server, encoding processes are respectively assigned and distributed in units such as GOP (Group of Picture), picture unit, or divided block unit of a picture. Thereby, latency can be reduced, and more real-time performance can be achieved. Furthermore, since the plurality of image data is roughly of the same scene, it is also possible to manage and / or instruct the server to mutually reference the image data captured by each terminal device. Alternatively, the server may receive the encoded data from each terminal device, change the reference relationship among the plurality of data, or correct or replace the picture itself, and then re-encode it. Thereby, a stream with improved quality and efficiency for each piece of data can be generated. Furthermore, the server may also perform transcoding to change the encoding method of the image data and then publish the image data. For example, the server can convert the encoding method of the MPEG system to the VP system, or convert H.264 to H.265. In this way, the encoding process can be performed by the terminal device or one or more servers. Therefore, although the following describes the "server" or "terminal" as the main body for performing the process, part or all of the processes performed by the server can also be performed by the terminal device, or part or all of the processes performed by the terminal device can also be performed by the server. Also, the same applies to the decoding process. [3D, Multi-angle] In recent years, there has been an increasing integration and utilization of images or videos of different scenes captured by a plurality of cameras ex113 and / or smartphones ex115 that are roughly synchronized with each other, or of the same scene captured from different angles. The images captured by each terminal device are integrated based on the relative position relationship between the terminal devices obtained separately, or the regions where the feature points contained in the images are consistent. The server can not only encode two-dimensional dynamic images, but also, based on scene analysis of the dynamic images, etc., automatically or at a time specified by the user, encode still images and send them to the receiving terminal device. When the server can further obtain the relative position relationship between the imaging terminal devices, it can generate the three-dimensional shape of the scene not only based on two-dimensional dynamic images, but also based on images of the same scene captured from different angles. Furthermore, the server can separately encode three-dimensional data such as point clouds generated, or use the three-dimensional data to identify people or objects, or select or reconstruct the images to be sent to the receiving terminal device from the images captured by a plurality of terminal devices according to the tracking results. In this way, the user can arbitrarily select each image corresponding to each imaging terminal device to view the scene, and can also view the content of the image cut from an arbitrary viewpoint of the three-dimensional data reconstructed from a plurality of images or videos. Furthermore, similar to the images, the sound can also be collected from a plurality of different angles, and the server multiplexes the sound from a specific angle or space with the image and sends them together. In recent years, content that correlates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become increasingly popular. In the case of VR images, the server separately creates viewpoint images for the right eye and the left eye, and through Multi-View Coding (MVC) or the like, coding that allows reference between each viewpoint image can be performed, or they can be coded as different streams without mutual reference. When decoding different streams, in response to the user's viewpoint, they are synchronized and played back to reproduce a virtual three-dimensional space. In the case of AR images, the server overlays virtual object information on the virtual space based on the camera information of the real space according to the three-dimensional position or the movement of the user's viewpoint. The decoding device can also acquire or hold virtual object information and three-dimensional data, and generate two-dimensional images in response to the movement of the user's viewpoint, and smoothly combine them to produce overlapping data. Also, in addition to sending a request for virtual object information, the decoding device also sends the movement of the user's viewpoint to the server. The server creates overlapping data in accordance with the movement of the viewpoint received from the three-dimensional data held in the server, encodes the overlapping data, and publishes it to the decoding device. Furthermore, in addition to RGB, the overlapping data has an α value indicating transparency. The server can also set the α value of the part other than the object created from the three-dimensional data to 0, etc., and encode it in a state where it will penetrate in that part. Or, the server can also generate data by setting a predetermined RGB value for the background like chroma key and setting the part other than the object to the background color. Similarly, the decoding process of the published data can be performed on each terminal device as a client or on the server side, or they can also be shared with each other. As an example, a certain terminal device can temporarily send a reception request to the server, receive the content corresponding to the request by another terminal device, perform the decoding process, and send a signal indicating the completion of decoding to the device with a display. By dispersing the processing and selecting appropriate content regardless of the performance of the communicable terminal device itself, data with good picture quality can be played. Also, as another example, large-size image data can be received on a TV or the like, and a part of the area such as a divided block of the decoded image can be displayed on the personal terminal device of the viewer. Thereby, the entire image can be shared, and at the same time, the viewer's responsible area or the area to be confirmed in more detail can be confirmed on hand. Furthermore, in the future, it is expected that in a situation where wireless communication of multiple types of short-distance, medium-distance, or long-distance can be used without being affected by the inside or outside of the house, by using a distribution system standard such as MPEG-DASH, appropriate data is switched while receiving content seamlessly for the ongoing communication. Thereby, the user is not limited to their own terminal device, and can freely select a decoding device or a display device such as a monitor installed inside or outside the house and switch instantaneously. Also, based on their own location information, etc., the decoding terminal device and the display terminal device can be switched while decoding. Thereby, even while moving towards a destination, a part of the wall surface or the ground of an adjacent building in which a display device is embedded can display map information while moving. Also, based on the ease of access to the encoded data on the network, such as caching the encoded data on a server that can be accessed in a short time from the receiving terminal device, or replicating it on an edge server of the content distribution service, etc., the bit rate of the received data can be switched. [Adaptive Coding] Regarding content switching, as shown in FIG. 23, the characteristics of the adaptive stream compressed and encoded by applying the dynamic image encoding method shown in the above-described embodiments and each modification example will be described. It is also possible for the server to have multiple streams with the same content but different qualities as individual streams, but as shown in the figure, it is also possible to utilize the characteristics of the temporal / spatial adaptive stream achieved by encoding in layers to switch the content and configure it. In short, the decoding side determines the layer to which decoding reaches based on internal factors such as performance and external factors such as the communication band state. Thereby, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, after watching an image on a smartphone ex115 while moving and then wanting to watch it on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side. Furthermore, in addition to the above-described configuration where for each layer of encoded pictures, an enhancement layer exists at a higher position in the base layer to achieve adaptability, the enhancement layer includes meta-information based on image statistical information, etc. The decoding side performs super-resolution on the pictures of the base layer based on the meta-information, and thereby it is also possible to generate high-quality content. Super-resolution can also refer to either an increase in the signal-to-noise ratio of the same resolution or an expansion of the resolution. The meta-information includes information for specifying linear or non-linear filter coefficients used in the super-resolution process, or information for specifying parameter values of filtering processes, machine learning, or least squares operations used in the super-resolution process, etc. Alternatively, it can also be configured as follows: according to the semantics of objects in the image, etc., the picture is divided into blocks, etc., and the decoding side selects the blocks to be decoded, thereby only decoding a part of the area. Also, object attributes (such as people, cars, balls, etc.) and positions within the image (coordinate positions within the same image, etc.) are stored as meta-information. Thereby, the decoding side can, based on the meta-information, specify the position of the required object and determine the block containing the object. For example, as shown in FIG. 24, the meta-information is stored using a data storage structure different from the pixel data, such as the SEI message of HEVC. The meta-information represents, for example, the position, size, or color of the main object, etc. Also, it is also possible to store the meta-information in units composed of a plurality of pictures, such as in a stream, sequence, random access unit, etc. Thereby, the decoding side can obtain the moment when a specific person appears in the image, etc. By matching the information of the picture unit, the picture in which the object exists and the position of the object within the picture can be specified. [Web Optimization] FIG. 25 is a diagram showing an example of a display screen of a web page of a computer ex111, etc. FIG. 26 is a diagram showing an example of a display screen of a web page of a smartphone ex115, etc. As shown in FIGS. 25 and 26, a web page sometimes includes a plurality of linked images linked to the image content, and the viewing method varies depending on the device being browsed. When a plurality of linked images can be seen on the screen, the display device (decoding device) displays a still image or I picture included in each content as a linked image, or displays an image such as a gif animation composed of a plurality of still images or I pictures, or only receives the base layer, decodes and displays the image until the user clearly selects a linked image, or the linked image approaches near the center of the image, or the entire linked image enters the screen. When an image is selected by the user, the display device decodes the base layer with the highest priority. Furthermore, when the HTML constituting the web page has information indicating adaptable content, the display device can also decode up to the enhancement layer. Also, in order to ensure immediacy, when the communication bandwidth is very strict before selection or during selection, the display device only decodes and displays the pictures for forward reference (I pictures, P pictures, B pictures for only forward reference), thereby reducing the delay between the decoding time and the display time of the starting picture (the delay from the start of content decoding to the start of display). Also, the display device can also deliberately ignore the reference relationship of the pictures, perform forward reference, roughly decode all B pictures and P pictures, and perform normal decoding as time passes and the received pictures increase. [Autopilot] Also, when receiving still image or video data such as two-dimensional or three-dimensional map information for vehicle autonomous driving or driving support, the receiving terminal device can also receive, in addition to the image data belonging to layer 1 or above, weather or construction information, etc. as meta-information, and decode them in correspondence. Furthermore, the meta-information can belong to a layer or be multiplexed simply with the image data. At this time, since vehicles, drones, airplanes, etc. that include the receiving terminal device move, when the receiving terminal device receives a request, it sends the location information of the receiving terminal device, thereby enabling seamless reception and decoding while switching base stations ex106 to ex110. Also, the receiving terminal device can dynamically switch the reception level of metadata or the update level of map information according to the user's selection, the user's situation, or the state of the communication band. As described above, in the content supply system ex100, the customer can immediately receive the encoded information sent by the user, decode it, and play it. [Personal content publishing] Also, in the content supply system ex100, not only can high-quality and long-duration content from video publishers be published, but also low-quality and short-duration unicast or multicast publishing from individuals can be performed. Also, this type of personal content should increase in the future. In order to make personal content higher-quality content, the server can also perform an editing process and then an encoding process. This can be achieved by, for example, the following configuration. During photography, the server performs recognition processes such as photography error, scene exploration, semantic analysis, and object detection on the original image or the encoded data immediately or after accumulation after photography. Then, based on the recognition results, the server performs the following editing: manually or automatically corrects out-of-focus or camera shake, deletes scenes with low importance such as scenes with lower brightness or out-of-focus than other pictures, emphasizes the edges of objects, or changes the color tone. The server encodes the edited data based on the editing results. Also, it is known that if the photography time is too long, the viewing rate will decrease. The server can also automatically edit scenes with low importance as described above and scenes with little movement according to the photography time based on the image processing results, so as to maintain the content within a specific time range. Also, the server can generate and encode a summary based on the result of the semantic analysis of the scene. Furthermore, in the case of personal content, there are also cases where content that directly captures infringement of copyright, moral rights of the author, or portrait rights, etc. appears, and there are also cases where the scope of sharing exceeds the intended scope, etc., which are inconvenient for individuals. Therefore, for example, the user can also encode the image by deliberately defocusing the face of a person or the inside of a house, etc. at the periphery of the screen. Also, the server can recognize whether the face of a person different from the pre-registered person is captured in the image to be encoded, and when it is captured, it can also perform processing such as adding a mosaic to the face part. Or, as pre-processing or post-processing before encoding, from the perspective of copyright, etc., the user can specify the person or background area of the image processing to be performed, and the server performs processing such as replacing the specified area with other images or blurring the focus. If it is a person, the image of the face part can also be replaced while tracking the person in the moving image. In addition, for viewing personal content with a small amount of data, immediacy is strongly required. Therefore, although it varies depending on the bandwidth, the decoding device first receives the base layer with the highest priority, decodes it, and plays it. During this period, the decoding device receives the enhancement layer. When playing the loop more than twice, for example, it is also possible to play a high-definition image including the enhancement layer. In this way, for a stream with scalable coding, the following experience can be provided: when not selected or at the start of viewing, the animation is rough, but as the stream becomes finer, the image also improves. In addition to scalable coding, the rough stream played for the first time and the second stream encoded with reference to the first animation can be combined into one stream to provide the same experience. [Other usage examples] In addition, these encoding or decoding processes are generally processed by the LSIex500 in each terminal device. The LSIex500 can be a single-chip or a configuration composed of multiple chips. Furthermore, software for encoding or decoding dynamic images can be incorporated into a certain recording medium (such as a CD-ROM, floppy disk, or hard disk) readable by a computer ex111, etc., and the encoding or decoding process can be performed using this software. Furthermore, when a smartphone ex115 is equipped with a camera, it is also possible to transmit the video data obtained by the camera. The video data at this time is data encoded by the LSIex500 in the smartphone ex115. Furthermore, the LSIex500 can also be configured to download and enable application software. At this time, the terminal device first determines whether the terminal device supports the encoding method of the content or has the execution ability of a specific service. When the terminal device does not support the encoding method of the content or does not have the execution ability of a specific service, the terminal device downloads the content or application software, and then acquires and plays the content. In addition, not limited to the content supply system ex100 via the Internet ex101, in a digital playback system, at least any one of the dynamic image encoding devices (image encoding devices) or dynamic image decoding devices (image decoding devices) of the above embodiments can also be incorporated. Since the playback radio wave carries multiplexed video and audio through satellites, etc., for a unicast-friendly configuration compared to the content supply system ex100, the difference is that it is suitable for multicast, but the same applications can be made for the encoding process and the decoding process. [Hardware configuration] FIG. 27 is a diagram showing an example of the smartphone ex115. Further, FIG. 28 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes: an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110; a camera unit ex465 capable of shooting images and still pictures; and a display unit ex458 for displaying decoded data such as the images shot by the camera unit ex465 and the images received by the antenna ex450. The smartphone ex115 further includes: an operation unit ex466, which is a touch panel or the like; a sound output unit ex457, such as a speaker for outputting sound or audio; a sound input unit ex456, such as a microphone for performing sound input; a memory unit ex467 for storing encoded data such as the shot images or still pictures, the recorded sounds, the received images or still pictures, emails, or decoded data; and a slot unit ex464, which is an interface portion with the SIM ex468, and the SIM ex468 is used to identify the user and perform authentication for accessing various data such as the network. Further, an external memory may be used instead of the memory unit ex467. Further, the main control unit ex460 that overall controls the display unit ex458, the operation unit ex466, etc. is connected via a bus ex470 to a power supply circuit unit ex461, an operation input control unit ex462, an image signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, a sound signal processing unit ex454, a slot unit ex464, and a memory unit ex467. When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack and activates the smartphone ex115 to an operable state. The smartphone ex115 performs processes such as calls and data communication under the control of the main control unit ex460 composed of a CPU, ROM, RAM, etc. During a call, the voice signal processing unit ex454 converts the voice signal received by the voice input unit ex456 into a digital voice signal, and the modulation / demodulation unit ex452 performs spread spectrum processing on it. After performing digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, the received data is amplified, frequency conversion processing and analog-to-digital conversion processing are performed, and the modulation / demodulation unit ex452 performs despread spectrum processing. After being converted into analog voice data by the voice signal processing unit ex454, it is output from the voice output unit ex457. In the data communication mode, by operating the operation unit ex466 of the main body, etc., text, still image, or image data is sent to the main control unit ex460 via the operation input control unit ex462, and the sending and receiving processes are performed in the same way. When sending an image, still image, or image and voice in the data communication mode, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 by the dynamic image encoding method shown in the above embodiments and various modification examples, and sends the encoded image data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453, where the voice signal is the signal received by the voice input unit ex456 during the process of shooting an image or still image, etc. with the camera unit ex465. The multiplexing / demultiplexing unit ex453 multiplexes the encoded image data and the encoded voice data in a predetermined manner, and performs modulation processing and conversion processing with the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450. When receiving an image attached to an email or chat, or an image linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data to distinguish the multiplexed data into a bitstream of video data and a bitstream of audio data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronous bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the dynamic image encoding method shown in the above embodiments and each modification example, and displays the image or still image contained in the linked dynamic image file from the display unit ex458 through the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal and outputs the audio from the audio output unit ex457. Furthermore, since real-time streaming has become popular, depending on the user's situation, it may also occur that the playback of the audio is inappropriate from a social perspective. Therefore, as an initial value, it is advisable to adopt a configuration that does not play the audio signal and only plays the video signal. The audio can also be played synchronously only when the user performs an operation such as clicking on the video data, etc. Also, here, the smartphone ex115 is taken as an example for explanation, but in addition to the transceiver terminal device having both an encoder and a decoder as a terminal device, three implementation forms such as a transmitting terminal device having only an encoder and a receiving terminal device having only a decoder can also be considered. Furthermore, it is explained that in a digital playback system, multiplexed data such as audio data is multiplexed with video data for reception or transmission, but in the multiplexed data, in addition to the audio data, text data related to the video, etc. can also be multiplexed, or the video data itself can be received or transmitted without receiving or transmitting the multiplexed data. Furthermore, it is explained that the main control unit ex460 including the CPU controls the encoding or decoding process, but the terminal device also often has a GPU. Therefore, it can also be configured as follows: by using the performance of the GPU to uniformly process a large area through the memory in which the CPU and the GPU are shared, or through the memory whose address is managed so as to be commonly used. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is very efficient to use the GPU instead of the CPU to uniformly perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing in units of pictures, etc. Industrial Applicability This disclosure can be applied to a television receiver, a digital video recorder, a car navigation device, a mobile phone, a digital camera, or a digital video camera, etc. 10~23: Block 100: Encoding device 102: Splitting section 104: Subtraction section 106: Conversion section 108: Quantization section 110: Entropy encoding section 112, 204: Inverse quantization section 114, 206: Inverse conversion section 116, 208: Addition section 118, 210: Block memory 120, 212: Loop filtering section 122, 214: Frame memory 124, 216: Intra-frame prediction section 126, 218: Inter-frame prediction section 128, 220: Prediction control section 200: Decoding device 202: Entropy decoding section ALF: Adaptive loop filter AMT: Adaptive multiple transform AR: Augmented reality AST: Adaptive second transform BIO: Bidirectional optical flow CCLM: Cross-component linear mode CABAC: Context-adaptive binary arithmetic coding method CDN: Content delivery network CTU: Coding tree unit CU: Coding unit Cur block: Current block DCT: Discrete cosine transform DF: Deblocking filter DST: Discrete sine transform EMT: Explicit multiple core transform ex100: Content supply system ex101: Internet ex102: Internet service provider ex103: Streaming server ex104: Communication network ex106~ex110: Base station ex111: Computer ex112: Game console ex113: Camera ex114: Home appliance ex115: Smartphone ex116: Satellite ex117: Aircraft ex450: Antenna ex451: Transmitting / receiving section ex452: Modulating / demodulating section ex453: Multiplexing / demultiplexing section ex454: Audio signal processing section ex455: Video signal processing section ex456: Audio input section ex457: Audio output section ex458: Display section ex459: Display control section ex460: Main control section ex461: Power circuit section ex462: Operation input control section ex463: Camera interface section ex464: Slot section ex465: Camera section ex466: Operation section ex467: Memory section ex468: SIM ex470: Bus, synchronous bus ex500: LSI FRUC: Frame rate up-conversion GOP: Group of pictures HEVC: High efficiency video coding HyGT: Hypercube Givens transform JCT-VC: Joint Collaborative Team on Video Coding JEM: Joint Exploration Model JVET: Joint Video Exploration Team KLT: K-L transform MBT: Polymorphic tree MV, MV0, MV1, MV_L,MV_U: Motion Vector, MVC: Multi-View Coding, NSST: Non-Separable Second Transform, OBMC: Overlapped Block Motion Compensation, PDPC: Prediction Combination in Intra-Prediction with Independent Positions, PMMVD: Pattern-Matching Motion Vector Derivation, PPS: Picture Parameter Set, Pred, Pred_L, Pred_U: Predicted Picture, PU: Prediction Unit, QP: Quantization Parameter, QTBT: Quadtree plus Binary Tree, Ref0, Ref1: Reference Picture, S101~S106, S201~S204, S401~S408, S501~S601~S606, S701~S703, S801, S901~S908: Steps, SAO: Sample Adaptive Offset, SPS: Sequence Parameter Set, TU: Transform Unit, v, 0 ,v 1 ,v x ,v y : Motion Vector, VLC: Variable Length Coding, VR: Virtual Reality FIG. 1 is a block diagram showing the functional configuration of the encoding device according to Embodiment 1. FIG. 2 is a diagram showing an example of block division according to Embodiment 1. FIG. 3 is a table showing conversion basis functions corresponding to respective conversion types. FIG. 4A is a diagram showing an example of the shape of the filter used for ALF. FIG. 4B is another example of the shape of the filter used for ALF. FIG. 4C is another example of the shape of the filter used for ALF. FIG. 5A is a diagram showing 67 in-frame prediction modes of in-frame prediction. FIG. 5B is a flowchart for explaining the outline of prediction image correction processing for OBMC processing. FIG. 5C is a conceptual diagram for explaining the outline of prediction image correction processing for OBMC processing. FIG. 5D is a diagram showing an example of FRUC. FIG. 6 is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory. FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in the current picture and a block in a reference picture. FIG. 8 is a diagram for explaining a model assuming uniform linear motion. FIG. 9A is a diagram for explaining derivation of a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks. FIG. 9B is a diagram for explaining the outline of motion vector derivation processing using a merge mode. FIG. 9C is a conceptual diagram for explaining the outline of DMVR processing. FIG. 9D is a diagram for explaining the outline of a method for generating a prediction image of luminance correction processing using LIC processing. FIG. 10 is a block diagram showing the functional configuration of the decoding device according to Embodiment 1. FIG. 11 is a flowchart showing the conversion and quantization processing of the encoding device according to Embodiment 2. FIG. 12 is a flowchart showing the inverse quantization and inverse conversion processing of the decoding device according to Embodiment 2. FIG. 13 is a flowchart showing the conversion and quantization processing of the encoding device according to Embodiment 3. FIG. 14 is a flowchart showing the inverse quantization and inverse conversion processing of the decoding device according to Embodiment 3. FIG. 15 is a flowchart showing the conversion and quantization processing of the encoding device according to Embodiment 4. FIG. 16 is a flowchart showing the inverse quantization and inverse conversion processing of the decoding device according to Embodiment 4. FIG. 17 is a flowchart showing the encoding processing of the encoding device according to Embodiment 5. FIG. 18 is a diagram showing a specific example of the syntax according to Embodiment 5. FIG. 19 is a flowchart showing the decoding processing of the decoding device according to Embodiment 5. FIG. 20 is a flowchart showing the encoding processing of the encoding device according to Embodiment 6. FIG. 21 is a flowchart showing the decoding processing of the decoding device according to Embodiment 6. FIG. 22 is an overall configuration diagram of a content supply system that implements a content distribution service. FIG. 23 is a diagram showing an example of an encoding structure in the case of adaptive encoding. FIG. 24 is a diagram showing an example of an encoding structure in the case of adaptive encoding. FIG. 25 is a diagram showing an example of a display screen of a web page. FIG. 26 is a diagram showing an example of a display screen of a web page. FIG. 27 is a diagram showing an example of a smartphone. FIG. 28 is a block diagram showing a configuration example of a smartphone. S101~S106: Steps

Claims

1. An encoding apparatus for encoding a first current block and a second current block in an image, the encoding apparatus comprising: a circuit; and a memory, wherein the circuit, during operation, performs the following processing: applying a first transformation to a first residual signal of the first current block and a second residual signal of the second current block using a transformation basis selected from candidate transformation bases to generate first transformation coefficients; and (i) when the transformation basis used for the first transformation applied to the first residual signal is different from a predetermined transformation basis, applying a first quantization to the first transformation coefficients of the first current block without performing a second transformation, and (ii) when the transformation basis used for the first transformation applied to the second residual signal is the same as the predetermined transformation basis, applying a second transformation to the first transformation coefficients of the second current block to generate second transformation coefficients, and applying a second quantization to the second transformation coefficients; wherein, The aforementioned second transformation is performed based on the transformation basis selected from the candidate transformation basis.

2. An encoding method for encoding a first current block and a second current block in an image, the aforementioned encoding method comprising: Using a conversion base selected from candidate conversion bases, a first conversion is performed on the first residual signal of the first current block and the second residual signal of the second current block to generate first conversion coefficients; and (i) when the conversion base used for the first conversion applied to the first residual signal is different from the predetermined conversion base, a second conversion is not performed and the first conversion coefficients of the first current block are first quantized; and (ii) when the conversion base used for the first conversion applied to the second residual signal is the same as the predetermined conversion base, a second conversion is performed on the first conversion coefficients of the second current block to generate second conversion coefficients, and the second conversion coefficients are second quantized; wherein the second conversion is performed based on a conversion base selected from candidate conversion bases.

3. A decoding apparatus for generating a first residual signal of a first current block and a second residual signal of a second current block in an image, the decoding apparatus comprising: a circuit; and a memory, wherein the circuit, during operation, performs the following processing: (i) applying a first inverse quantization to a first quantization coefficient of the first current block to generate a first conversion coefficient, and, when the inverse conversion substrate used to apply the first inverse conversion to the first conversion coefficient is different from a predetermined inverse conversion substrate, not performing a second inverse conversion, but using the inverse conversion substrate selected from candidate inverse conversion substrates to apply the first inverse conversion to the first conversion coefficient to generate the first residual signal, and (ii) applying a second inverse quantization to a second quantization coefficient of the second current block to generate a second conversion coefficient; The second inverse transformation is applied to the aforementioned second transformation coefficients to generate the third transformation coefficients; and when the inverse transformation basis used to apply the aforementioned first inverse transformation to the aforementioned third transformation coefficients is the same as the aforementioned predetermined inverse transformation basis, the aforementioned inverse transformation basis included in the aforementioned candidate inverse transformation basis is used to apply the aforementioned first inverse transformation to the aforementioned third transformation coefficients to generate the aforementioned second residual signal, wherein, The aforementioned second inverse transformation is performed based on the inverse transformation basis selected from the candidate inverse transformation basis.

4. A decoding method for generating a first residual signal of a first current block and a second residual signal of a second current block in an image, the aforementioned decoding method comprising: (i) A first inverse quantization is performed on the first quantization coefficient of the first current block to generate a first conversion coefficient. When the inverse conversion base used to apply the first inverse conversion to the first conversion coefficient is different from the predetermined inverse conversion base, a second inverse conversion is not performed. Instead, the first inverse conversion is performed on the first conversion coefficient using the inverse conversion base selected from the candidate inverse conversion base to generate the first residual signal. (ii) A second inverse quantization is performed on the second quantization coefficient of the second current block to generate a second conversion coefficient. A second inverse conversion is performed on the second conversion coefficient to generate a third conversion coefficient. When the inverse conversion base used to apply the first inverse conversion to the third conversion coefficient is the same as the predetermined inverse conversion base, the first inverse conversion is performed on the third conversion coefficient using the inverse conversion base included in the candidate inverse conversion base to generate the second residual signal. The second inverse conversion is performed based on the inverse conversion base selected from the candidate inverse conversion base.

5. A method for transmitting a bitstream, comprising: performing a first conversion on a first residual signal of a first current block and a second residual signal of a second current block using a conversion basis selected from candidate conversion bases to generate first conversion coefficients; and (i) when the conversion basis used for the first conversion applied to the first residual signal is different from a predetermined conversion base, performing a first quantization on the first conversion coefficients of the first current block without performing a second conversion, and (ii) when the conversion basis used for the first conversion applied to the second residual signal is the same as the predetermined conversion base, performing a second conversion on the first conversion coefficients of the second current block to generate second conversion coefficients, and performing a second quantization on the second conversion coefficients, wherein the method for transmitting the bitstream encodes the results of the first quantization and the second quantization to generate a bitstream, and transmits the bitstream, wherein the second conversion is performed based on the conversion basis selected from candidate conversion bases.

Citation Information

Patent Citations

  • Hybrid video compression method

    US20060251330A1