Encoder, decoder, encoding method, and decoding method
The encoding and decoding device optimizes conversion basis selection based on block size to enhance compression efficiency and reduce processing load in video coding, addressing the limitations of existing technologies like HEVC.
Patent Information
- Application Number
- JP2025073147
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-28
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2038-12-19
AI Technical Summary
Existing video coding technologies, such as HEVC, require further improvements in compression efficiency and reduction in processing load.
An encoding and decoding device that determines the validity of a conversion basis selection mode based on the size of the encoding/decoding target block, selecting a first conversion basis when the vertical size exceeds a threshold and a fixed basis when it does not, and performs corresponding conversions to generate bitstreams.
This approach enhances compression efficiency and reduces processing load by optimizing conversion basis selection, thereby improving overall performance.
Smart Images

Figure 2025107227000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method.
Background Art
[0002] A video coding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding).
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] In such encoding and decoding technologies, further improvement in compression efficiency and reduction in processing load are required.
[0005] Therefore, the present disclosure provides an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement in compression efficiency and reduction in processing load.
Means for Solving the Problem
[0006] An encoding device according to an aspect of the present disclosure is an encoding device including a circuit and a memory. The circuit determines whether a mode of selecting a conversion basis according to the size of an encoding target block is valid by using the memory. When the mode is valid, when the vertical size of the encoding target block is larger than a threshold size, a first conversion basis is selected as the vertical conversion basis from among a plurality of conversion basis candidates, and when the vertical size of the encoding target block is smaller than the threshold size, a second conversion basis that is a fixed conversion basis is selected as the vertical conversion basis. A first conversion coefficient is generated by performing a first conversion in the vertical direction on the residual of the encoding target block by using the selected vertical conversion basis, a second conversion coefficient is generated by performing a second conversion on the first conversion coefficient, and a bitstream including information indicating whether the mode is valid is generated.
[0007] A decoding device according to an aspect of the present disclosure is a decoding device including a circuit and a memory. The circuit generates a conversion coefficient by performing a second inverse conversion on the coefficients of a decoding target block by using the memory, determines whether a mode of selecting a conversion basis according to the size of the decoding target block is valid, and when the mode is valid, when the vertical size of the decoding target block is larger than a threshold size, a first inverse conversion basis is selected as the vertical inverse conversion basis from among a plurality of inverse conversion basis candidates, and when the vertical size of the decoding target block is smaller than the threshold size, a second inverse conversion basis that is a fixed inverse conversion basis is selected as the vertical inverse conversion basis. A prediction residual is generated by performing a first inverse conversion in the vertical direction on the conversion coefficient of the decoding target block by using the selected vertical inverse conversion basis.
[0008] Incidentally, all or specific aspects of these may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
Advantages of the Invention
[0009] The present disclosure can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement in compression efficiency and reduction in processing load.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 10
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 13A
Figure 13B
Figure 14
Figure 15
Figure 16
Figure 17A
Figure 17B
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22A
Figure 22B
Figure 23
Figure 24A
Figure 24B
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
MODE FOR CARRYING OUT THE INVENTION
[0011] Hereinafter, the embodiments will be specifically described with reference to the drawings.
[0012] Note that all of the embodiments described below show comprehensive or specific examples. Numerical values, shapes, materials, components, the arrangement positions and connection forms of the components, steps, the order of steps, etc. shown in the following embodiments are merely examples, and are not intended to limit the scope of the claims. In addition, among the components in the following embodiments, the components not described in the independent claims indicating the most superior concept are described as optional components.
[0013] (First Embodiment) First, as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure to be described later are applicable, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.
[0014] When applying the processes and / or configurations described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.
[0015] (1) With respect to the encoding device or the decoding device of Embodiment 1, among the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. (2) With respect to the encoding device or the decoding device of Embodiment 1, after performing any change such as addition, replacement, or deletion of functions or processes to be performed on some of the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. (3) With respect to the method performed by the encoding device or the decoding device of Embodiment 1, after performing addition of processes and / or any change such as replacement or deletion of some of the plurality of processes included in the method, the process corresponding to the process described in each aspect of the present disclosure is replaced with the process described in each aspect of the present disclosure. (4) Some of the plurality of components constituting the encoding device or the decoding device of Embodiment 1 are combined and implemented with the component described in each aspect of the present disclosure, a component having a part of the functions provided by the component described in each aspect of the present disclosure, or a component that performs a part of the processes performed by the component described in each aspect of the present disclosure. (5) A component that includes a part of the functions provided by some of the components that make up the encoding device or decoding device of Embodiment 1, or a component that performs a part of the processes performed by some of the components that make up the encoding device or decoding device of Embodiment 1, is combined with the components described in each aspect of the present disclosure, a component that includes a part of the functions provided by the components described in each aspect of the present disclosure, or a component that performs a part of the processes performed by the components described in each aspect of the present disclosure for implementation. (6) For the method performed by the encoding device or decoding device of Embodiment 1, among the plurality of processes included in the method, the process corresponding to the process described in each aspect of the present disclosure is replaced with the process described in each aspect of the present disclosure. (7) A part of the processes included in the method performed by the encoding device or decoding device of Embodiment 1 is combined with the processes described in each aspect of the present disclosure for implementation. Note that the ways of implementing the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Also, the processes and / or configurations described in different aspects may be combined for implementation.
[0016] [Overview of Encoding Device] First, the overview of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.
[0017] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0018] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Also, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0019] Each component included in the encoding device 100 will be described below.
[0020] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These blocks of fixed size are sometimes called coding tree units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into variable-size blocks (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be the processing units of CUs, PUs, and TUs.
[0021] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, solid lines represent block boundaries by quadtree block splitting, and dashed lines represent block boundaries by binary tree block splitting.
[0022] Here, block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).
[0023] The upper-left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper-left 64x64 block is split into two 16x64 blocks 11, 12 and a 32x64 block 13.
[0024] The upper-right 64x64 block is horizontally split into two rectangular 64x32 blocks 14, 15 (binary tree block splitting).
[0025] The 64x64 block in the lower left is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the 64x64 block in the lower left is divided into 16 16x32 blocks, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.
[0026] The 64x64 block 23 in the lower right is not divided.
[0027] As described above, in FIG. 2, the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.
[0028] Note that in FIG. 2, one block is divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary-tree block division). Such division including ternary-tree block division is sometimes called MBT (multi type tree) division.
[0029] [Subtraction unit] The subtraction unit 104 subtracts the prediction signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.
[0030] The original signal is the input signal of the encoding device 100, and is a signal representing the images of each picture constituting a moving image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing an image may also be referred to as a sample.
[0031] [Transformation unit] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0032] Note that the transformation unit 106 may adaptively select a transformation type from a plurality of transformation types and transform the prediction error into transformation coefficients using a transform basis function corresponding to the selected transformation type. Such a transformation may be referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0033] The plurality of transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the transform basis functions corresponding to each transformation type. In FIG. 3, N indicates the number of input pixels. The selection of the transformation type from among these plurality of transformation types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.
[0034] Information indicating whether to apply such EMT or AMT (for example, called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that the signaling of this information need not be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0035] Further, the conversion unit 106 may re-convert the conversion coefficient (conversion result). Such re-conversion may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (e.g., 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of these pieces of information does not have to be limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0036] Here, separable conversion is a method in which conversion is performed multiple times by separating for each direction by the number of dimensions of the input, and non-separable conversion is a method in which when the input is multi-dimensional, two or more dimensions are regarded as one dimension and conversion is performed together.
[0037] For example, as an example of non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.
[0038] Similarly, after regarding a 4×4 input block as an array having 16 elements, a method in which Givens rotation is performed multiple times on the array (Hypercube Givens Transform) is also an example of non-separable conversion.
[0039] [Quantization unit] The quantization unit 108 quantizes the transform coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the scanned transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0040] The predetermined order is the order for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).
[0041] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0042] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable length encoding on the quantization coefficients that are input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.
[0043] [Inverse Quantization Unit] The inverse quantization unit 112 inverse quantizes the quantization coefficients that are input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse conversion unit 114.
[0044] [Inverse Conversion Unit] The inverse transform unit 114 restores the prediction error by inversely transforming the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0045] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.
[0046] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.
[0047] [Block memory] The block memory 118 is a storage unit for storing blocks within an encoded target picture (hereinafter referred to as the current picture) that are blocks referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.
[0048] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0049] In ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 sub-block within a current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.
[0050] Specifically, first, a sub-block (for example, a 2x2 sub-block) is classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-block is performed based on the direction and activity of the gradient. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-block is classified into a plurality of classes (for example, 15 or 25 classes).
[0051] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantizing the addition result.
[0052] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.
[0053] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. The information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (for example, sequence level, slice level, tile level, CTU level, or CU level).
[0054] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance difference, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or CU level. Note that the signaling of the information indicating the on / off of ALF does not necessarily have to be limited to the picture level or CU level, and it may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0055] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets does not necessarily have to be limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0056] [Frame Memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0057] [Intra Prediction Unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0058] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0059] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0060] The plurality of directional prediction modes includes, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.
[0061] In addition, in the intra prediction of a chrominance block, a luminance block may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0062] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. Intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of PDPC application (for example, called a PDPC flag) is signaled at, for example, the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0063] [Inter prediction unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. The inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0064] The motion information used for motion compensation is signaled. A motion vector predictor may be used for signaling the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0065] In addition to the motion information of the current block obtained by motion search, the motion information of adjacent blocks may also be used to generate an inter-prediction signal. Specifically, an inter-prediction signal may be generated for each sub-block in the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter-prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0066] In such an OBMC mode, information indicating the size of sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be other levels (for example, picture level, slice level, tile level, CTU level or sub-block level).
[0067] The OBMC mode will be described more specifically. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of prediction image correction processing by OBMC processing.
[0068] First, a prediction image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.
[0069] Next, the motion vector (MV_L) of the encoded left adjacent block is applied to the block to be coded to obtain a prediction image (Pred_L), and the first correction of the prediction image is performed by overlapping and weighting the prediction image and Pred_L.
[0070] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_U). The predicted image after the first correction and Pred_U are weighted and superimposed to perform the second correction of the predicted image, which is used as the final predicted image.
[0071] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to adopt a configuration in which more corrections than two stages are performed using the right adjacent block or the lower adjacent block.
[0072] Note that the area for superimposition may be only a partial area near the block boundary instead of the pixel area of the entire block.
[0073] Here, the prediction image correction process from a single reference picture has been described. However, the same applies to the case of correcting the prediction image from a plurality of reference pictures. After obtaining the prediction images corrected from each reference picture, the obtained prediction images are further superimposed to obtain the final prediction image.
[0074] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.
[0075] As a method for determining whether to apply the OBMC process, for example, there is a method using an obmc_flag, which is a signal indicating whether to apply the OBMC process. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as the obmc_flag and the OBMC process is applied for encoding. If it does not belong to a region with complex motion, the value 0 is set as the obmc_flag and encoding is performed without applying the OBMC process. On the other hand, in a decoding device, the obmc_flag described in the stream is decoded, and decoding is performed by switching whether to apply the OBMC process according to the value.
[0076] Note that the motion information may be derived on the decoder side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also for example, the motion information may be derived by performing motion search on the decoder side. In this case, the motion search is performed without using the pixel values of the current block.
[0077] Here, a mode in which motion search is performed on the decoder side will be described. This mode in which motion search is performed on the decoder side may be referred to as the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.
[0078] An example of FRUC processing is shown in FIG. 5D. First, by referring to the motion vectors of the encoded blocks spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0079] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also for example, by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, a motion vector for the current block may be derived. That is, search is performed in the same manner for the region around the best candidate MV, and if there is an MV for which the evaluation value is a good value, the best candidate MV may be updated to the MV, and that may be used as the final MV of the current block. Note that a configuration in which the said processing is not performed is also possible.
[0080] When performing processing in sub-block units, the same processing may be performed.
[0081] The evaluation value is calculated by obtaining the difference value of the reconstructed image by pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that the evaluation value may be calculated using information other than the difference value.
[0082] As the pattern matching, first pattern matching or second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.
[0083] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the candidate described above, the region in another reference picture along the motion trajectory of the current block is used.
[0084] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is advisable to select the candidate MV with the best evaluation value among the plurality of candidate MVs as the final MV.
[0085] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0086] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidate.
[0087] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that most closely matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or either one of the left adjacent and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.
[0088] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level, and it may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0089] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called the BIO (bi - directional optical flow) mode.
[0090] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x , v yindicates the velocity vector, and τ0 and τ1 indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.
[0091] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), respectively, and the following optical flow equation (1) holds.
[0092]
Equation
[0093] Here, I (k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from the merge list or the like are corrected in pixel units.
[0094] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0095] Here, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.
[0096] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of the adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (v x , v y ) of each sub-block within the current block is derived by the following equation (2).
[0097] [Equation]
[0098] Here, x and y indicate the horizontal position and the vertical position of the sub-block, respectively, and w indicates a predetermined weighting coefficient.
[0099] Such an affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right control points are different. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode is not necessarily limited to the CU level, and it may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0100] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.
[0101] Here, an example of deriving the motion vector of the picture to be encoded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.
[0102] First, a prediction MV list in which candidates for the prediction MV are registered is generated. Examples of candidates for the prediction MV include a spatial adjacent prediction MV which is an MV possessed by a plurality of encoded blocks located spatially adjacent to the block to be encoded, a temporal adjacent prediction MV which is an MV possessed by a nearby block obtained by projecting the position of the block to be encoded in the encoded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV which is an MV with a value of zero.
[0103] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be encoded.
[0104] Furthermore, in the variable length encoding unit, a merge_idx which is a signal indicating which prediction MV has been selected is described in the stream and encoded.
[0105] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are merely examples, and the number may be different from that in the figure, the configuration may not include some of the types of prediction MVs in the figure, or prediction MVs other than the types of prediction MVs in the figure may be added.
[0106] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be encoded derived in the merge mode.
[0107] Here, an example of determining the MV using the DMVR process will be described.
[0108] Figure 9C is a conceptual diagram for explaining the outline of the DMVR process.
[0109] First, using the optimal MVP set for the processing target block as a candidate MV, reference pixels are respectively obtained from the first reference picture which is the processed picture in the L0 direction and the second reference picture which is the processed picture in the L1 direction according to the candidate MV, and a template is generated by taking the average of each reference pixel.
[0110] Next, using the template, the peripheral areas of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, the MV value, etc.
[0111] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.
[0112] Note that even if it is not the processing itself described here, any other processing may be used as long as it is a process capable of searching the periphery of the candidate MV to derive the final MV.
[0113] Here, the mode of generating a predicted image using the LIC process will be described.
[0114] Figure 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction process by the LIC process.
[0115] First, an MV for obtaining a reference image corresponding to the block to be encoded from the reference picture which is the encoded picture is derived.
[0116] Next, for the block to be encoded, using the luminance pixel values of the left adjacent and upper adjacent encoded peripheral reference areas and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, information indicating how the luminance values change between the reference picture and the block to be encoded is extracted to calculate the luminance correction parameter.
[0117] By performing a luminance correction process on the reference image in the reference picture specified by MV using the luminance correction parameter, a predicted image for the block to be encoded is generated.
[0118] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.
[0119] Also, although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from a plurality of reference pictures. After performing a luminance correction process on the reference images obtained from each reference picture in the same manner, a predicted image is generated.
[0120] As a method for determining whether to apply the LIC process, for example, there is a method using a lic_flag which is a signal indicating whether to apply the LIC process. As a specific example, in an encoding apparatus, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, a value 1 is set as the lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, a value 0 is set as the lic_flag and encoding is performed without applying the LIC process. On the other hand, in a decoding apparatus, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply the LIC process according to the value.
[0121] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process has been applied to peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the peripheral encoded blocks selected during the derivation of the MV in the merge mode process have been encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. Note that in this example, the process in decoding is exactly the same.
[0122] [Overview of the Decoder] Next, an overview of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing the functional configuration of a decoder 200 according to Embodiment 1. The decoder 200 is a moving image / image decoder that decodes moving images / images in block units.
[0123] As shown in FIG. 10, the decoder 200 includes an entropy decoder 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0124] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoder 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoder 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0125] Hereinafter, each component included in the decoder 200 will be described.
[0126] [Entropy Decoder] The entropy decoder 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoder 202, for example, arithmetically decodes the encoded bit stream into a binary signal. Then, the entropy decoder 202 debinarizes the binary signal. As a result, the entropy decoder 202 outputs quantization coefficients in block units to the inverse quantization unit 204.
[0127] [Inverse quantization unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0128] [Inverse transform unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.
[0129] For example, when the information decoded from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.
[0130] Also, for example, when the information decoded from the encoded bitstream indicates that NSST is to be applied, the inverse transform unit 206 applies inverse-retransformation to the transform coefficients.
[0131] [Addition unit] The addition unit 208 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 206 and the prediction sample that is the input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0132] [Block memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture) that are blocks referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the addition unit 208.
[0133] [Loop Filter Section] The loop filter section 212 applies a loop filter to the block reconstructed by the adder section 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.
[0134] When the information read from the encoded bitstream indicating the on / off of the ALF indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.
[0135] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used for inter prediction and may also be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter section 212.
[0136] [Intra Prediction Section] The intra prediction section 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on the intra prediction mode read from the encoded bitstream. Specifically, the intra prediction section 216 generates an intra prediction signal by performing intra prediction with reference to samples (for example, luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control section 220.
[0137] In addition, when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance difference block is selected, the intra prediction section 216 may predict the chrominance difference component of the current block based on the luminance component of the current block.
[0138] Also, when the information read from the encoded bitstream indicates the application of PDPC, the intra prediction section 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction.
[0139] [Inter Prediction Unit] The inter prediction unit 218 predicts the current block by referring to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using motion information (e.g., motion vectors) decoded from the encoded bit stream to generate an inter prediction signal for the current block or sub-block, and outputs the inter prediction signal to the prediction control unit 220.
[0140] Note that when the information decoded from the encoded bit stream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.
[0141] Also, when the information decoded from the encoded bit stream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.
[0142] Also, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion when the BIO mode is applied. Also, when the information decoded from the encoded bit stream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks.
[0143] [Prediction Control Unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208.
[0144] (First Aspect of Embodiment 1) Next, the first aspect of Embodiment 1 will be specifically described with reference to the drawings.
[0145] [Internal Configuration of the Conversion Unit of the Encoding Device] First, the internal configuration of the conversion unit 106 of the encoding device 100 according to this aspect will be described with reference to FIG. 11A. FIG. 11A is a block diagram showing the internal configuration of the conversion unit 106 of the encoding device 100 according to the first aspect of Embodiment 1.
[0146] As shown in FIG. 11A, the conversion unit 106 according to this aspect includes a conversion mode determination unit 1061, a size determination unit 1062, a first conversion basis selection unit 1063, a first conversion unit 1064, a second conversion execution determination unit 1065, a second conversion basis selection unit 1066, and a second conversion unit 1067.
[0147] The conversion mode determination unit 1061 determines whether the adaptive conversion basis selection mode is valid for the encoding target block. The adaptive conversion basis selection mode is a mode in which the conversion basis is adaptively selected from one or more first conversion basis candidates. The determination of whether the adaptive conversion basis selection mode is valid is performed, for example, based on the first conversion basis or the identification information of the adaptive conversion basis selection mode.
[0148] The size determination unit 1062 determines whether the horizontal size of the encoding target block exceeds a first horizontal threshold size. Also, the size determination unit 1062 determines whether the vertical size of the encoding target block exceeds a first vertical threshold size. The first horizontal threshold size may be the same as or different from the first vertical threshold size. The first horizontal threshold size and the first vertical threshold size may be, for example, predefined in a standard specification. Also, for example, the first horizontal threshold size and the first vertical threshold size may be sizes determined based on an image and encoded in the bitstream.
[0149] The first conversion basis selection unit 1063 selects the first conversion basis. In the present disclosure, selecting a basis includes determining or setting at least one basis without candidates of a plurality of basis candidates in addition to selecting at least one basis from among a plurality of basis candidates.
[0150] When the adaptive conversion basis selection mode is not valid, the first conversion basis selection unit 1063 selects one basic conversion basis as the first conversion basis in the horizontal and vertical directions. Further, when the adaptive conversion basis selection mode is valid, the first conversion basis selection unit 1063 selects the first conversion basis in the horizontal and vertical directions as follows (1) to (4) based on the horizontal size and vertical size of the block to be encoded.
[0151] (1) When the horizontal size of the block to be encoded is larger than the first horizontal threshold size, the first conversion basis selection unit 1063 adaptively selects the first conversion basis in the horizontal direction from among one or more conversion basis candidates.
[0152] (2) When the horizontal size of the block to be encoded is less than or equal to the first horizontal threshold size, the first conversion basis selection unit 1063 selects a fixed conversion basis in the horizontal direction as the first conversion basis in the horizontal direction.
[0153] (3) When the vertical size of the block to be encoded is larger than the first vertical threshold size, the first conversion basis selection unit 1063 adaptively selects the first conversion basis in the vertical direction from among one or more conversion basis candidates.
[0154] (4) When the vertical size of the block to be encoded is less than or equal to the first vertical threshold size, the first conversion basis selection unit 1063 selects a fixed conversion basis in the vertical direction as the first conversion basis in the vertical direction.
[0155] The fixed conversion basis in the horizontal direction may be the same as or different from the fixed conversion basis in the vertical direction. As the fixed conversion basis in the horizontal and vertical directions, for example, a conversion basis of discrete sine transform type 7 (DST-VII) can be used.
[0156] The first conversion unit 1064 generates first conversion coefficients by performing a first conversion on the residual of the block to be encoded using the first conversion basis selected by the first conversion basis selection unit 1063. Specifically, the first conversion unit 1064 performs a first conversion in the horizontal direction using the first conversion basis in the horizontal direction and a first conversion in the vertical direction using the first conversion basis in the vertical direction.
[0157] The second conversion execution determination unit 1065 determines whether to perform a second conversion for further converting the first conversion coefficients based on whether the adaptive conversion basis selection mode is valid for the block to be encoded. Specifically, the second conversion execution determination unit 1065 determines to perform the second conversion when the adaptive conversion basis selection mode is not valid, and determines not to perform the second conversion when the adaptive conversion basis selection mode is valid.
[0158] The second conversion basis selection unit 1066 selects a second conversion basis when it is determined to perform the second conversion. That is, the second conversion basis selection unit 1066 selects the second conversion basis when the adaptive conversion basis selection mode is not valid. Conversely, when the adaptive conversion basis selection mode is valid, the second conversion basis selection unit 1066 does not select the second conversion basis. That is, the second conversion basis selection unit 1066 skips the selection of the second conversion basis when the adaptive conversion basis selection mode is valid.
[0159] The second conversion unit 1067 converts the first conversion coefficients using the second conversion basis selected by the second conversion basis selection unit 1066 when it is determined to perform the second conversion. That is, the second conversion unit 1067 generates second conversion coefficients by performing a second conversion on the first conversion coefficients using the second conversion basis when the adaptive conversion basis selection mode is not valid. Conversely, when the adaptive conversion basis selection mode is valid, the second conversion unit 1067 does not perform a second conversion on the first conversion coefficients. That is, the second conversion unit 1067 skips the second conversion when the adaptive conversion basis selection mode is valid.
[0160] [Internal Configuration of the Inverse Transformation Unit of the Symbolization Device] Next, the internal configuration of the inverse transformation unit 114 of the symbolization device 100 according to this embodiment will be described with reference to FIG. 11B. FIG. 11B is a block diagram showing the internal configuration of the inverse transformation unit 114 of the symbolization device 100 according to the first aspect of Embodiment 1.
[0161] As shown in FIG. 11B, the inverse transformation unit 114 according to this embodiment includes a second inverse transformation basis selection unit 1141, a second inverse transformation unit 1142, a first inverse transformation basis selection unit 1143, and a first inverse transformation unit 1144.
[0162] When the adaptive transformation basis selection mode is not valid for the block to be encoded, the second inverse transformation basis selection unit 1141 selects the inverse transformation basis of the second transformation basis selected by the second transformation basis selection unit 1066 as the second inverse transformation basis.
[0163] When the adaptive transformation basis selection mode is not valid for the block to be encoded, the second inverse transformation unit 1142 performs a second inverse transformation on the inverse quantized coefficients using the second inverse transformation basis selected by the second inverse transformation basis selection unit 1141 to generate second inverse transformation coefficients. The inverse quantized coefficients mean the coefficients inverse quantized by the inverse quantization unit 112.
[0164] The first inverse transformation basis selection unit 1143 selects the inverse transformation basis of the first transformation basis selected by the first transformation basis selection unit 1063 as the first inverse transformation basis.
[0165] When the adaptive transformation basis selection mode is not valid for the block to be encoded, the first inverse transformation unit 1144 reconstructs the residual of the block to be encoded by performing a first inverse transformation on the second inverse transformation coefficients using the first inverse transformation basis. On the other hand, when the adaptive transformation basis selection mode is valid for the block to be encoded, the first inverse transformation unit 1144 reconstructs the residual of the block to be encoded by performing a first inverse transformation on the inverse quantized coefficients using the first inverse transformation basis.
[0166] [Processing of the Conversion Unit and Quantization Unit of the Symbolization Device] Next, the processing of the conversion unit 106 configured as described above will be described with reference to FIG. 12A together with the processing of the quantization unit 108. FIG. 12A is a flowchart showing the processing of the conversion unit 106 and the quantization unit 108 of the symbolization device 100 according to the first aspect of the first embodiment.
[0167] The conversion mode determination unit 1061 determines whether the adaptive transform basis selection mode is valid for the encoding target block (S101).
[0168] When the adaptive transform basis selection mode is not valid (NO in S101), the first transform basis selection unit 1063 selects one basic transform basis as the first transform basis in the horizontal and vertical directions (S102).
[0169] When the adaptive transform basis selection mode is valid (YES in S101), the size determination unit 1062 determines whether the transform size in the horizontal direction exceeds a certain range (S103). That is, the size determination unit 1062 determines whether the horizontal size of the encoding target block is larger than the first horizontal threshold size.
[0170] When the transform size in the horizontal direction exceeds a certain range (YES in S103), the first transform basis selection unit 1063 selects a transform basis in the horizontal direction from a plurality of adaptive transform bases as the first transform basis in the horizontal direction (S104).
[0171] When the transform size in the horizontal direction is within a certain range (NO in S103), the first transform basis selection unit 1063 selects a fixed transform basis as the first transform basis in the horizontal direction (S105).
[0172] Next, the size determination unit 1062 determines whether the transform size in the vertical direction exceeds a certain range (S106). That is, the size determination unit 1062 determines whether the vertical size of the encoding target block is larger than the first vertical threshold size.
[0173] When the conversion size in the vertical direction exceeds a certain range (YES in S106), the first conversion basis selection unit 1063 selects a conversion basis in the vertical direction from a plurality of adaptive conversion bases as the first conversion basis in the vertical direction (S107).
[0174] When the conversion size in the vertical direction is within a certain range (NO in S106), the first conversion basis selection unit 1063 selects a fixed conversion basis as the first conversion basis in the vertical direction (S108).
[0175] Note that the selection order of the conversion bases in the horizontal and vertical directions may be in the order of the horizontal and vertical directions, or in the reverse order. Also, the conversion bases in the horizontal direction and the vertical direction may be selected simultaneously.
[0176] The first conversion unit 1064 performs a first conversion on the prediction residual using the first conversion basis selected in step S102, S107 or step S108 to generate first conversion coefficients (S109).
[0177] Next, the second conversion execution determination unit 1065 determines whether to perform a second conversion on the first conversion coefficients (S110). Here, the second conversion execution determination unit 1065 determines whether to perform a second conversion based on whether the adaptive conversion basis selection mode is valid for the encoding target block.
[0178] When the adaptive conversion basis selection mode is valid (YES in S110), neither the selection of the second conversion basis nor the second conversion is performed, and the quantization unit 108 generates quantization coefficients by performing quantization of the first conversion coefficients (S113). That is, steps S111 and S112 in FIG. 12A are skipped.
[0179] When the adaptive transform basis selection mode is not valid (NO in S110), the second transform basis selection unit 1066 selects a second transform basis from among one or more candidates for the second transform basis (S111). Then, the second transform unit 1067 generates a second transform coefficient by performing a second transform on the first transform coefficient using the selected second transform basis (S112). Thereafter, the quantization unit 108 generates a quantization coefficient by quantizing the second transform coefficient (S113).
[0180] As the above-described basic transform basis, a predetermined transform basis can be used. In this case, it may be determined whether the adaptive transform basis selection mode is valid based on whether the first transform bases in the horizontal and vertical directions are the predetermined transform basis. Also, the predetermined transform basis may be one transform basis or two or more transform bases.
[0181] Also, when not performing (skipping) the second transform, it is not necessary to perform the second transform, or a transform equivalent to not performing the transform may be performed as the second transform. In the former case, information indicating that the second transform is not performed may be encoded in the bit stream. Also, in the latter case, information indicating a transform equivalent to not performing the transform may be encoded in the bit stream. The same can be said for the process of skipping each transform hereinafter.
[0182] Note that the steps and the order of the steps shown in FIG. 12A are merely examples and are not limited thereto. For example, as shown in FIG. 12B, the determination of the adaptive transform basis selection mode (S101) and the determination of performing the second transform (S110) in FIG. 12A may be integrated. FIG. 12B is a flowchart showing a modification of the processing of the transform unit 106 and the quantization unit 108 of the encoding apparatus 100 according to the first aspect of the first embodiment. The flowchart of FIG. 12B is substantially the same as the flowchart of FIG. 12A.
[0183] In FIG. 12B, the execution determination (S110) of the second conversion is deleted, and the first conversion (S109) is divided into two (S109A, S109B). In this case, the conversion unit 106 of the encoding device 100 may not include the second conversion execution determination unit 1065.
[0184] Regarding the selection of the second inverse conversion basis and the second inverse conversion in the inverse conversion unit 114, and the selection of the first inverse conversion basis and the first inverse conversion, they may be performed according to the conversion of the conversion unit 106 in FIG. 12A, so the description and illustration are omitted.
[0185] Note that the first conversion may be a frequency conversion that can adaptively select a conversion basis like the EMT described in Non-Patent Document 2, a frequency conversion that switches the conversion basis under certain conditions, or other general conversions. For example, a fixed conversion basis may be set instead of selecting the first conversion basis. Also, a first conversion basis equivalent to not performing the first conversion may be used. Further, in the first conversion, identification information indicating which of the adaptive conversion basis selection mode and the conversion basis fixed mode using a fixed basic conversion basis (for example, the conversion basis of type 2 discrete cosine transform (DCT-II)) is effective may be used to select either of the two modes. In this case, it is also possible to determine which of the adaptive conversion basis selection mode and the conversion basis fixed mode is effective for the encoding target block based on the identification information. For example, in the EMT described in Non-Patent Document 2, there is identification information (emt_cu_flag) indicating whether the adaptive conversion basis selection mode is effective in units such as a CU (Coding Unit), so it is possible to determine whether the adaptive conversion basis selection mode is effective for the encoding target block using that identification information.
[0186] Note that the second conversion may be a secondary conversion process such as NSST described in Non-Patent Document 2, a conversion that switches the conversion basis under certain conditions, or other general conversions. For example, a fixed conversion basis may be set instead of selecting the second conversion basis. Also, a second conversion basis equivalent to not performing the second conversion may be used. Further, NSST may be a frequency space conversion after DCT or DST. For example, it may be KLT (Karhunen Loveve Transform) for the conversion coefficients of DCT or DST obtained offline, or HyGT (Hypercube-Givens Transform) represented by a combination of a basis equivalent to KLT and rotation transformations.
[0187] Note that this process can be applied to both the luminance signal and the color difference signal. If the input signal is in RGB format, it may be applied to each of the R, G, and B signals. Further, the bases selectable in the first conversion or the second conversion may be different between the luminance signal and the color difference signal. For example, since the luminance signal has a wider frequency band than the color difference signal, more types of bases may be used as selection candidates in the first conversion or the second conversion of the luminance signal than for the color difference in order to perform an optimal conversion. Also, this process can be applied to both intra processing and inter processing.
[0188] [Effects, etc.] In the first conversion (primary conversion) and the second conversion (secondary conversion) described in Non-Patent Document 2, an optimal conversion basis or conversion coefficients (filters) are selected, and optimal encoding efficiency is realized in total. Therefore, it is necessary to perform the first conversion and the second conversion a large number of times in order to search for an optimal combination of candidates for the conversion basis and conversion coefficients (filters) used in the first conversion and the second conversion. That is, in the conversion method described in Non-Patent Document 2, it is necessary to calculate an evaluation value for all combinations of candidates for the conversion basis of the first conversion and candidates for the conversion basis of the second conversion, and select the combination that minimizes the evaluation value. Therefore, the inventors have found a problem that the processing amount becomes extremely large in the conversion method described in Non-Patent Document 2.
[0189] Therefore, the encoding device 100 according to this aspect does not always perform both the first conversion and the second conversion, but skips the second conversion based on whether the adaptive conversion basis selection mode is valid. As a result, the encoding device 100 can reduce the number of combinations of candidates for the conversion basis of the first conversion and candidates for the conversion basis of the second conversion, and can reduce the processing amount.
[0190] Further, according to the encoding device 100 according to this aspect, candidates for the first conversion basis can be limited based on the conditions of the conversion sizes in the horizontal and vertical directions. As a result, it is possible to reduce the processing amount of searching for the best first conversion basis by trial. Also, based on conditions such as the basis selected for the first conversion basis, it is possible to reduce the process of searching for the best second conversion basis by trial. Furthermore, it is possible to reduce the processing amount related to the trial of combinations of the first conversion and the second conversion.
[0191] As an example, as the basic conversion basis, a conversion basis of DCT-II can be used. DCT-II is highly likely to be adopted when the residual shape is flat or random. For example, when DCT-II is used as the first conversion basis, since the degree of aggregation in the low frequency band tends to increase, the effect of the second conversion may increase. On the other hand, with a conversion basis other than DCT-II, high frequency components tend to remain, and the effect of the second conversion may decrease.
[0192] Also, as an example, as the fixed conversion basis selected when the conversion size is within a certain range, a conversion basis of DST-VII can be used. DST-VII has a tendency to be selected with a very high probability especially in intra processing when the residual shape is inclined and the size is small.
[0193] Note that the basic conversion basis is not limited to one predetermined conversion basis, and a plurality of predetermined conversion bases may be used.
[0194] Also, whether to perform the selection of the second conversion basis and the second conversion may be switched according to the conversion size. Also, candidates for the second conversion basis may be switched according to the conversion size.
[0195] Also, it may be configured to only switch whether to perform the second conversion based on whether the adaptive conversion basis selection mode is valid, and not to switch the first conversion basis according to the conversion size. That is, in FIG. 12A, steps S103, S105, S106, and step S108 may be deleted. Here, whether the adaptive conversion basis selection mode is valid may be determined based on identification information indicating the use of the mode, or based on the type of the first conversion basis.
[0196] Similarly, it may be configured to only switch the first conversion basis according to the conversion size without switching whether to perform the second conversion based on whether the adaptive conversion basis selection mode is valid. That is, in FIG. 12A, step S110 may be deleted.
[0197] Note that, regardless of whether the adaptive conversion basis selection mode is valid, the selection of the second conversion basis and the second conversion may not be skipped. Also, regardless of the method of selecting the first conversion basis, when the adaptive conversion basis selection mode is not valid, the selection of the second conversion basis and the second conversion may be performed, and when the adaptive conversion basis selection mode is valid, the selection of the second conversion basis and the second conversion may be skipped.
[0198] Also, as specific conversion size thresholds in the horizontal or vertical direction for selecting candidates from a plurality of adaptive conversion bases or selecting a fixed conversion basis as the first conversion basis (that is, the first horizontal threshold size and the first vertical threshold size), 4, 8, 16, 32, or 64 pixels, etc. may be used.
[0199] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, some processes described in the flowchart of this aspect, some configurations of the apparatus, some syntax, etc. may be implemented in combination with other aspects.
[0200] (Second Aspect of Embodiment 1) Next, the second aspect of Embodiment 1 will be described. In this aspect, an example of encoding various signals related to the first conversion and the second conversion in the first aspect will be described. Hereinafter, this aspect will be specifically described with reference to the drawings, centering on the differences from the first aspect.
[0201] Note that the internal configurations of the conversion unit 106 and the inverse conversion unit 114 of the encoding apparatus 100 according to this aspect are the same as those in the first aspect, so the illustration is omitted.
[0202] [Processes of the Conversion Unit, Quantization Unit, and Entropy Encoding Unit of the Encoding Apparatus] The processes of the conversion unit 106, the quantization unit 108, and the entropy encoding unit 110 of the encoding apparatus 100 according to this aspect will be described with reference to FIGS. 13A and 13B. FIG. 13A is a flowchart showing the processes of the conversion unit 106 and the quantization unit 108 of the encoding apparatus 100 according to the second aspect of Embodiment 1. FIG. 13B is a flowchart showing the process of the entropy encoding unit 110 of the encoding apparatus 100 according to the second aspect of Embodiment 1. In FIGS. 13A and 13B, for the processes common to the first aspect, the same reference numerals are given and the description is omitted.
[0203] After quantization (S113), the entropy encoding unit 110 encodes an adaptive transform basis selection mode signal (S201). The adaptive transform basis selection mode signal is an example of identification information of the adaptive transform basis selection mode.
[0204] If the adaptive transform basis selection mode is enabled (YES in S202), when the horizontal transform size exceeds a certain range (YES in S203), the entropy encoding unit 110 encodes the first basis selection signal in the horizontal direction (S204). On the other hand, when the horizontal transform size is within a certain range (NO in S203), the entropy encoding unit 110 does not encode the first basis selection signal in the horizontal direction. Further, when the vertical transform size exceeds a certain range (YES in S205), the entropy encoding unit 110 encodes the first basis selection signal in the vertical direction (S206). On the other hand, when the vertical transform size is within a certain range (NO in S205), the entropy encoding unit 110 does not encode the first basis selection signal in the vertical direction.
[0205] When the adaptive transform basis selection mode is not enabled (NO in S202), the encoding of the first basis selection signal (S204, S206) is skipped.
[0206] Next, the entropy encoding unit 110 encodes the quantization coefficients (S207).
[0207] Here, when the adaptive transform basis selection mode is not enabled (NO in S208), the entropy encoding unit 110 encodes the second basis selection signal (S209). On the other hand, when the adaptive transform basis selection mode is enabled (YES in S208), the encoding of the second basis selection signal (S209) is skipped.
[0208] Note that the order of each encoding may be determined in advance, and various signals may be encoded so as to be different from the above encoding order.
[0209] When the second transform is not performed (skipped), a signal indicating that the second transform is not performed may be encoded, or a signal selecting a second basis equivalent to not performing the transform may be encoded.
[0210] [Syntax] Here, the syntax in this aspect will be described. FIG. 14 shows a specific example of the syntax in the second aspect of Embodiment 1.
[0211] In FIG. 14, for example, when the adaptive transform basis selection mode signal (emt_cu_flag) is set (the 4th line), if the horizontal transform size (horizontal_tu_size) is larger than the first horizontal threshold size (horizontal_tu_size_th) (the 5th line), the first horizontal basis selection signal (emt_horizontal_tridx) is encoded (the 6th line). Also, if the vertical transform size (vertical_tu_size) is larger than the first vertical threshold size (vertical_tu_size_th) (the 11th line), the first vertical basis selection signal (emt_vertical_tridx) is encoded (the 12th line). Under other conditions (the 8th line and the 14th line), the encoding of the first basis selection signal is skipped (the 9th line and the 15th line).
[0212] Also, when the adaptive transform basis selection mode signal (emt_cu_flag) is not set (the 19th line), the second basis selection signal (secondary_tridx) is encoded (the 20th line). Conversely, when the adaptive transform basis selection mode signal (emt_cu_flag) is set (the 22nd line), the encoding of the second basis selection signal (secondary_tridx) is skipped (the 23rd line).
[0213] [Specific Examples of Transform Bases and Encoded Signals] Next, specific examples of the transform bases and encoded signals will be described. FIG. 15 shows specific examples of the transform bases used in the second aspect of Embodiment 1 and the presence or absence of signal encoding.
[0214] In FIG. 15, when the adaptive transform basis selection mode is not valid, regardless of the size of the encoding block, the transform basis of DCT-II is used as the first transform basis in the horizontal and vertical directions. That is, the transform basis of DCT-II is used as the basic transform basis. Also, while the second transform is being performed (ON), a second basis selection signal (secondary_tridx) indicating the second transform basis used in the second transform is encoded in the bit stream.
[0215] On the other hand, when the adaptive transform basis selection mode is valid, depending on the horizontal size H and vertical size V of the block to be encoded, a combination of the transform basis of DST-VII and other transform bases (index0~index3) is used as candidates for the first transform basis in the horizontal and vertical directions. Also, regardless of the size of the block to be encoded, the second transform is not performed (OFF). Also, the second basis selection signal (secondary_tridx) is not encoded, but an adaptive transform basis selection mode signal (emt_cu_flag) is encoded in the bit stream. Further, when the horizontal size H of the block to be encoded is larger than 4 pixels, a first basis selection signal (emt_horizontal_tridx) in the horizontal direction is encoded in the bit stream. Also, when the vertical size V of the block to be encoded is larger than 4 pixels, a first basis selection signal (emt_vertical_tridx) in the vertical direction is encoded in the bit stream.
[0216] For example, when the horizontal size H is 4 pixels or less and the vertical size V is 4 pixels or less, only the transform basis of DST-VII is used as a candidate for the first transform basis in the horizontal and vertical directions. At this time, the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) in the horizontal and vertical directions are not encoded.
[0217] For example, when the horizontal size H is 4 pixels or less and the vertical size V is greater than 4 pixels, only the transform basis of DST-VII is used as a candidate for the first transform basis in the horizontal direction, and the transform basis of DST-VII and other transform bases are used as candidates for the first transform basis in the vertical direction. At this time, the first basis selection signal (emt_horizontal_tridx) in the horizontal direction is not encoded, but the first basis selection signal (emt_vertical_tridx) in the vertical direction is encoded.
[0218] For example, when the horizontal size H is greater than 4 pixels and the vertical size V is 4 pixels or less, the transform basis of DST-VII and other transform bases are used as candidates for the first transform basis in the horizontal direction, and only the transform basis of DST-VII is used as a candidate for the first transform basis in the vertical direction. At this time, the first basis selection signal (emt_horizontal_tridx) in the horizontal direction is encoded, but the first basis selection signal (emt_vertical_tridx) in the vertical direction is not encoded.
[0219] For example, when the horizontal size H is greater than 4 pixels and the vertical size V is greater than 4 pixels, the transform basis of DST-VII and other transform bases are used as candidates for the first transform basis in each of the horizontal and vertical directions. At this time, the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) in the horizontal and vertical directions are encoded.
[0220] [Effects, etc.] As described above, according to the encoding device 100 according to this aspect, when the adaptive transform basis selection mode is valid and the transform size exceeds a certain range, information indicating the first transform basis (the first basis selection signal) can be encoded, and there is a possibility of reducing the amount of code required for signaling the first transform basis. Further, when the adaptive transform basis selection mode is not valid, information indicating the second transform basis (the second basis selection signal) can be encoded, and there is a possibility of reducing the amount of code required for signaling the second transform basis. Further, by encoding information (such as an adaptive transform basis selection mode signal) for determining whether to skip the second transform before the information indicating the second transform basis, it is possible to determine at the time of decoding whether the information indicating the second transform basis has been encoded.
[0221] Note that, regardless of the adaptive transform basis selection mode, the encoding of the second basis selection signal may always be performed. Further, regardless of the transform size, if it is the adaptive transform basis selection mode, the encoding of the first basis selection signal may always be performed. Further, the presence or absence of the encoding of the first basis selection signal may be determined independently for the horizontal size and the vertical size, or may be determined in combination.
[0222] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Further, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.
[0223] (Third aspect of Embodiment 1) Next, the third aspect of Embodiment 1 will be described. In this aspect, when the adaptive transform basis selection mode is not valid, a different basic transform basis is used as the first transform basis according to the size of the encoding target block, which is different from the first aspect above. Hereinafter, this aspect will be specifically described with reference to the drawings, centering on the differences from the first and second aspects.
[0224] Note that since the internal configurations of the conversion unit 106 and the inverse conversion unit 114 of the encoding device 100 according to this aspect are the same as those of the first aspect, illustration thereof is omitted.
[0225] [Processing of the conversion unit and quantization unit of the encoding device] The processing of the conversion unit 106 and the quantization unit 108 of the encoding device 100 according to this aspect will be described with reference to FIG. 16. FIG. 16 is a flowchart showing the processing of the conversion unit 106 and the quantization unit 108 of the encoding device 100 according to the third aspect of the first embodiment. In FIG. 16, for the processing common to the first aspect, the same reference numerals are given and the description thereof is omitted.
[0226] When the adaptive conversion basis selection mode is not valid (NO in S101), the size determination unit 1062 determines whether the conversion size is within a certain range (S301). That is, the size determination unit 1062 determines whether the size of the block to be encoded is less than or equal to the second threshold size. For example, the size determination unit 1062 determines whether the size of the block to be encoded is less than or equal to the second threshold size by determining whether the product of the horizontal size and the vertical size of the block to be encoded is less than or equal to the threshold.
[0227] Here, when the conversion size is within a certain range (YES in S301), the first conversion basis selection unit 1063 selects the second basic conversion basis as the first conversion basis in the horizontal and vertical directions (S302). On the other hand, when the conversion size exceeds a certain range (NO in S301), the first conversion basis selection unit 1063 selects the first basic conversion basis as the first conversion basis in the horizontal and vertical directions (S303).
[0228] As an example, as the first basic conversion basis, the conversion basis of DCT-II can be used, and as the second basic conversion basis, the conversion basis of DST-VII can be used.
[0229] Note that the basic conversion basis may be selected from among a plurality of basic conversion basis candidates.
[0230] Note that, regardless of whether the adaptive transform basis selection mode is effective or not, the selection of the second transform basis and the second transform may not be skipped. Also, regardless of the method for selecting the first transform basis, when the adaptive transform basis selection mode is not effective, the selection of the second transform basis and the second transform are performed, and when the adaptive transform basis selection mode is effective, the selection of the second transform basis and the second transform may be skipped.
[0231] Also, when the adaptive transform basis selection mode is not effective, as the second threshold size for selecting one of the first basic transform basis and the second basic transform basis, for example, pixel sizes such as 4x4, 4x8, 8x4, 8x8 can be used. Also, as the transform size to be compared with the threshold, the product of the horizontal size and the vertical size of the encoding target block may be used as in this embodiment, or each of the horizontal size and the vertical size may be used.
[0232] Note that when the adaptive transform basis selection mode is not effective, if the product of the horizontal size and the vertical size is within a certain range, the second basic transform basis is selected as the first transform basis in the horizontal and vertical directions, and the selection of the second transform basis and the second transform may be skipped.
[0233] [Effects, etc.] As described above, according to the encoding device 100 according to this embodiment, when the adaptive transform basis selection mode is not effective, the first transform basis can be switched between the first basic transform basis and the second basic transform basis according to the transform size. Therefore, the first transform can be performed using the first transform basis corresponding to the transform size, and the amount of code can be reduced.
[0234] [Combination with Other Embodiments] This embodiment may be implemented in combination with at least a part of other embodiments in the present disclosure. Also, a part of the processing described in the flowchart of this embodiment, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other embodiments.
[0235] (Fourth Aspect of Embodiment 1) Next, a fourth aspect of Embodiment 1 will be described. In this aspect, an example of encoding various signals related to the first conversion and the second conversion according to the third aspect will be described. Hereinafter, this aspect will be specifically described with reference to the drawings, centering on the differences from the first to the third aspects.
[0236] Note that the internal configurations of the conversion unit 106 and the inverse conversion unit 114 of the encoding device 100 according to this aspect are the same as those of the first aspect, so the illustration thereof is omitted.
[0237] [Processing of the conversion unit, quantization unit, and entropy encoding unit of the encoding device] The processing of the conversion unit 106, the quantization unit 108, and the entropy encoding unit 110 of the encoding device 100 according to this aspect will be described with reference to FIGS. 17A and 17B. FIG. 17A is a flowchart showing the processing of the conversion unit 106 and the quantization unit 108 of the encoding device 100 according to the fourth aspect of Embodiment 1. FIG. 17B is a flowchart showing the processing of the entropy encoding unit 110 of the encoding device 100 according to the fourth aspect of Embodiment 1. In FIGS. 17A and 17B, for the processing common to any of the first to third aspects, the same reference numerals are given and the description thereof is omitted.
[0238] After the quantization is performed (S113), the entropy encoding unit 110 determines whether to skip the encoding of the adaptive conversion basis selection mode signal (S401). For example, the entropy encoding unit 110 determines to skip the encoding of the adaptive conversion basis selection mode signal when any of the following conditions (A) and (B) is satisfied, and determines not to skip the encoding of the adaptive conversion basis selection mode signal otherwise.
[0239] (A) The adaptive conversion basis selection mode is not valid.
[0240] (B) The adaptive conversion basis selection mode is valid, and all of the following conditions (B1) to (B4) are satisfied.
[0241] (B1) The conversion size is equal to or less than the second threshold size W1xH1 used in step S301.
[0242] (B2) The conversion size in the horizontal direction is equal to or less than the first horizontal threshold size W2 used in step S103.
[0243] (B3) The conversion size in the vertical direction is equal to or less than the first vertical threshold size H2 used in step S106.
[0244] (B4) The second basic conversion basis and the fixed conversion bases in the horizontal and vertical directions are the same conversion basis.
[0245] As a specific example, when the second threshold size W1xH1 is 4x4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and both the second basic conversion basis and the fixed conversion basis are the DST-VII conversion basis, if the conversion size is 4x4 pixels or less, the entropy encoding unit 110 determines to skip the encoding of the adaptive conversion basis selection mode signal.
[0246] Conversely, if neither of the above conditions (A) and (B) is satisfied, the entropy encoding unit 110 determines not to skip the encoding of the adaptive conversion basis selection mode signal.
[0247] Here, when it is determined to skip the encoding of the adaptive conversion basis selection mode signal (YES in S401), the entropy encoding unit 110 skips steps S201 to S206 and encodes the quantization coefficients (S207). On the other hand, when it is determined not to skip the encoding of the adaptive conversion basis selection mode signal (NO in S401), the entropy encoding unit 110 executes steps S201 to S206 in the same manner as in the second aspect and then encodes the quantization coefficients (S207).
[0248] Note that the order of each encoding may be determined in advance, and various signals may be encoded in a different order from the above encoding order.
[0249] [Syntax] Here, the syntax in this aspect will be described. FIG. 18 shows a specific example of the syntax in the fourth aspect of Embodiment 1.
[0250] In FIG. 18, for example, when encoding of the adaptive transform base selection mode signal is skipped (line 20), encoding of the adaptive transform base selection mode signal (emt_cu_flag) and the first base selection signals (emt_horizontal_tridx and emt_vertical_tridx) is skipped (line 21). Here, when the horizontal transform size (horizontal_tu_size) is less than or equal to the first horizontal threshold size (horizontal_tu_size_th) and the vertical transform size (vertical_tu_size) is less than or equal to the first vertical threshold size (vertical_tu_size_th), encoding of the adaptive transform base selection mode signal is skipped. When encoding of the adaptive transform base selection mode signal is not skipped (lines 3 - 4), the adaptive transform base selection mode signal (emt_cu_flag) is encoded (line 5), and the first base selection signals (emt_horizontal_tridx and emt_vertical_tridx) are encoded as needed in the same manner as in the second aspect (lines 7 - 16).
[0251] Note that when encoding of the adaptive transform base selection mode signal is skipped, selection of the second transform base and the second transform may be skipped.
[0252] [Specific Examples of Transform Base and Encoding Signals] Next, specific examples of the transform base and encoding signals will be described. FIG. 19 shows a specific example of the transform base used in the fourth aspect of Embodiment 1 and the presence or absence of encoding of the signal. In FIG. 19, when both the horizontal size and the vertical size of the block to be encoded are 4 pixels or less, the transform base and the presence or absence of encoding are different from those in FIG. 15. FIG. 19 will be described centering on the differences from FIG. 15.
[0253] In FIG. 19, when the adaptive transform basis selection mode is not valid, if both the horizontal size H and the vertical size V of the block to be encoded are 4 pixels or less, the transform basis of DST-VII is used as the first transform basis in the horizontal and vertical directions instead of the transform basis of DCT-II.
[0254] Also, when the adaptive transform basis selection mode is valid, if both the horizontal size H and the vertical size V of the block to be encoded are 4 pixels or less, the adaptive transform basis selection mode signal (emt_cu_flag) is not encoded.
[0255] [Effects, etc.] As described above, according to the encoding apparatus 100 according to the present aspect, when the condition for skipping the encoding of the adaptive transform basis selection mode signal is satisfied, all encoding of the adaptive transform basis selection mode signal and the first basis selection signal can be omitted, and there is a possibility of reducing the amount of code.
[0256] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.
[0257] (Fifth aspect of Embodiment 1) Next, the fifth aspect of Embodiment 1 will be described. In this aspect, a decoding apparatus will be described. Note that the decoding apparatus according to this aspect corresponds to the encoding apparatus according to the first aspect. That is, the decoding apparatus according to this aspect can decode the bitstream encoded by the encoding apparatus according to the first aspect. Hereinafter, this aspect will be specifically described with reference to the drawings.
[0258] [Internal configuration of the transform unit and inverse transform unit of the decoding apparatus] First, the internal configuration of the inverse transform unit 206 of the decoding apparatus 200 according to this aspect will be described. FIG. 20 is a block diagram showing the internal configuration of the inverse transform unit 206 of the decoding apparatus 200 according to the fifth aspect of Embodiment 1.
[0259] As shown in FIG. 20, the inverse conversion unit 206 according to this aspect includes a second inverse conversion execution determination unit 2061, a second inverse conversion basis selection unit 2062, a second inverse conversion unit 2063, a conversion mode determination unit 2064, a size determination unit 2065, a first inverse conversion basis selection unit 2066, and a first inverse conversion unit 2067.
[0260] The second inverse conversion execution determination unit 2061 determines whether to perform a second inverse conversion on the inverse quantization coefficients of the block to be decoded output from the inverse quantization unit 204 based on whether the adaptive conversion basis selection mode is valid for the block to be decoded. Specifically, the second inverse conversion execution determination unit 2061 performs the second inverse conversion when the adaptive conversion basis selection mode is not valid, and determines not to perform the second inverse conversion when the adaptive conversion basis selection mode is valid.
[0261] The second inverse conversion basis selection unit 2062 selects a second inverse conversion basis when it is determined to perform the second inverse conversion. Specifically, when the adaptive conversion basis selection mode is not valid, the second inverse conversion basis selection unit 2062 obtains a second basis selection signal 2062S indicating the second inverse conversion basis decoded from the bit stream by the entropy decoding unit 202. Then, the second inverse conversion basis selection unit 2062 selects the second inverse conversion basis based on the second basis selection signal 2062S. Conversely, when the adaptive conversion basis selection mode is valid, the second inverse conversion basis selection unit 2062 does not select the second inverse conversion basis. That is, when the adaptive conversion basis selection mode is valid, the second inverse conversion basis selection unit 2062 skips the selection of the second inverse conversion basis.
[0262] When it is determined that the second inverse transform is to be performed, the second inverse transform unit 2063 performs a second inverse transform on the inverse quantization coefficients of the block to be decoded using the second inverse transform basis selected by the second inverse transform basis selection unit 2062. That is, when the adaptive transform basis selection mode is not valid, the second inverse transform unit 2063 generates second inverse transform coefficients by performing a second inverse transform on the inverse quantization coefficients using the second inverse transform basis. Conversely, when the adaptive transform basis selection mode is valid, the second inverse transform unit 2063 does not perform a second inverse transform on the inverse quantization coefficients. That is, when the adaptive transform basis selection mode is valid, the second inverse transform unit 2063 skips the second inverse transform.
[0263] The transform mode determination unit 2064 determines whether the adaptive transform basis selection mode is valid for the block to be decoded. The determination of whether the adaptive transform basis selection mode is valid is made based on the first basis selection signal 2066S or the adaptive transform basis selection mode signal 2064S decoded from the bit stream by the entropy decoding unit 202. That is, the determination is made based on the identification information of the first inverse transform basis or the adaptive transform basis selection mode.
[0264] The size determination unit 2065 determines whether the horizontal size of the block to be decoded exceeds a first horizontal threshold size. Also, the size determination unit 1062 determines whether the vertical size of the block to be decoded exceeds a first vertical threshold size. The determination of the horizontal size and the vertical size is made based on the size signal 2065S decoded from the bit stream by the entropy decoding unit 202.
[0265] The first inverse transform basis selection unit 2066 selects the first inverse transform basis. Specifically, when the adaptive transform basis selection mode is not valid, the first inverse transform basis selection unit 2066 selects one basic transform basis as the first inverse transform basis in the horizontal and vertical directions. Also, when the adaptive transform basis selection mode is valid, the first inverse transform basis selection unit 2066 selects the first inverse transform basis in the horizontal and vertical directions as follows (1) to (4) based on the horizontal size and vertical size of the block to be decoded.
[0266] (1) When the horizontal size of the block to be decoded is larger than the first horizontal threshold size, the first inverse transform basis selection unit 2066 obtains a first basis selection signal 2066S indicating the first inverse transform basis, which is decoded from the bit stream by the entropy decoding unit 202. Then, the first inverse transform basis selection unit 2066 selects the first inverse transform basis in the horizontal direction based on the first basis selection signal 2066S.
[0267] (2) When the horizontal size of the block to be decoded is less than or equal to the first horizontal threshold size, the first inverse transform basis selection unit 2066 selects a fixed transform basis in the horizontal direction as the first inverse transform basis in the horizontal direction.
[0268] (3) When the vertical size of the block to be decoded is larger than the first vertical threshold size, the first inverse transform basis selection unit 2066 obtains the first basis selection signal 2066S. Then, the first inverse transform basis selection unit 2066 selects the first inverse transform basis in the vertical direction based on the first basis selection signal 2066S.
[0269] (4) When the vertical size of the block to be decoded is less than or equal to the first vertical threshold size, the first inverse transform basis selection unit 2066 selects a fixed transform basis in the vertical direction as the first inverse transform basis in the vertical direction.
[0270] The first inverse transformation unit 2067 restores the residual of the block to be decoded by performing a first inverse transformation on the inverse quantization coefficients of the block to be decoded using the first inverse transformation basis selected by the first inverse transformation basis selection unit 2066. Specifically, the first inverse transformation unit 2067 performs a first inverse transformation in the horizontal direction using the first inverse transformation basis in the horizontal direction, and performs a first inverse transformation in the vertical direction using the first inverse transformation basis in the vertical direction.
[0271] [Processing of the Inverse Quantization Unit and Inverse Transformation Unit of the Decoding Device] Next, the processing of the inverse transformation unit 206 configured as described above will be described with reference to FIG. 21 together with the processing of the inverse quantization unit 204. FIG. 21 is a flowchart showing the processing of the inverse quantization unit 204 and the inverse transformation unit 206 of the decoding device 200 according to the fifth aspect of the first embodiment.
[0272] The inverse quantization unit 204 generates inverse quantization coefficients by inverse quantizing the quantization coefficients of the block to be decoded decoded by the entropy decoding unit 202 (S501).
[0273] The second inverse transformation execution determination unit 2061 determines whether to perform a second inverse transformation on the inverse quantization coefficients (S502). Here, the second inverse transformation execution determination unit 2061 determines whether to perform the second inverse transformation based on whether the adaptive transformation basis selection mode is valid for the block to be decoded.
[0274] Here, when the adaptive transformation basis selection mode is valid (YES in S502), neither the selection of the second inverse transformation basis nor the second inverse transformation is performed. That is, steps S503 and S504 are skipped.
[0275] On the other hand, when the adaptive transformation basis selection mode is not valid (NO in S502), the second inverse transformation basis selection unit 2062 selects a second inverse transformation basis based on the second basis selection signal 2062S (S503). Further, the second inverse transformation unit 2063 performs a second inverse transformation on the inverse quantization coefficients using the selected second inverse transformation basis (S504).
[0276] Next, the conversion mode determination unit 2064 determines whether the adaptive conversion basis selection mode is valid for the block to be decoded (S505). For example, the conversion mode determination unit 2064 determines whether the adaptive conversion basis selection mode is valid based on the adaptive conversion basis selection mode signal 2064S.
[0277] If the adaptive conversion basis selection mode is not valid (NO in S505), the first inverse conversion basis selection unit 2066 selects one basic conversion basis as the first inverse conversion basis in the horizontal and vertical directions (S512). On the other hand, if the adaptive conversion basis selection mode is valid (YES in S505), the size determination unit 2065 determines whether the conversion size in the horizontal direction exceeds a certain range (S506). That is, the size determination unit 2065 determines whether the horizontal size of the block to be decoded is larger than the first horizontal threshold size.
[0278] If the conversion size in the horizontal direction exceeds a certain range (YES in S506), the first inverse conversion basis selection unit 2066 selects a conversion basis in the horizontal direction from a plurality of adaptive conversion bases as the first inverse conversion basis in the horizontal direction (S507). On the other hand, if the conversion size in the horizontal direction is within a certain range (NO in S506), the first inverse conversion basis selection unit 2066 selects a fixed conversion basis as the first inverse conversion basis in the horizontal direction (S508).
[0279] The size determination unit 2065 determines whether the conversion size in the vertical direction exceeds a certain range (S509). That is, the size determination unit 2065 determines whether the vertical size of the block to be decoded is larger than the first vertical threshold size.
[0280] When the conversion size in the vertical direction exceeds a certain range (YES in S509), the first inverse transform basis selection unit 2066 selects a transform basis from a plurality of adaptive transform bases as the first inverse transform basis in the vertical direction (S510). When the conversion size in the vertical direction is within a certain range (NO in S509), the first inverse transform basis selection unit 2066 selects a fixed transform basis as the first inverse transform basis in the vertical direction (S511).
[0281] The first inverse transform unit 2067 restores the residual of the block to be decoded by performing a first inverse transform on the inverse quantized coefficients or the second inverse transform coefficients using the first inverse transform basis selected as described above (S513).
[0282] Note that the selection order of the inverse transform bases in the horizontal and vertical directions may be in the horizontal and vertical order, or in the reverse order. Also, the inverse transform bases in the horizontal and vertical directions may be selected simultaneously.
[0283] Note that selecting an inverse transform basis in the decoding apparatus 200 means decoding information indicating the basis to be used for inverse transform included in the encoded bit stream and determining the inverse transform basis based on the decoded information, or determining an inverse transform basis uniquely indicated based on information such as the intra prediction mode, the block size to be decoded, or the basis in the first inverse transform.
[0284] Note that a decoding method adapted to the encoding method of the first aspect shown in FIG. 12A or FIG. 12B may be employed.
[0285] [Effects, etc.] As described above, according to the decoding apparatus 200 according to the present aspect, the same effects as those of the encoding apparatus 100 according to the first aspect can be achieved.
[0286] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, some processes described in the flowchart of this aspect, some configurations of the apparatus, some syntax, etc. may be implemented in combination with other aspects.
[0287] (Sixth Aspect of Embodiment 1) Next, the sixth aspect of Embodiment 1 will be described. In this aspect, an example of decoding various signals related to the first conversion and the second conversion in the fifth aspect will be described. Note that the decoding apparatus according to this aspect corresponds to the encoding apparatus according to the second aspect. Hereinafter, this aspect will be specifically described with reference to the drawings, centering on the differences from the fifth aspect.
[0288] Note that the internal configuration of the inverse conversion unit 206 of the decoding apparatus 200 according to this aspect is the same as that in the fifth aspect, so the illustration is omitted.
[0289] [Processing of Entropy Decoding Unit, Inverse Quantization Unit, and Inverse Conversion Unit of Decoding Apparatus] The processing of the entropy decoding unit 202, the inverse quantization unit 204, and the inverse conversion unit 206 of the decoding apparatus 200 according to this aspect will be described with reference to FIGS. 22A and 22B. In FIGS. 22A and 22B, for the processes common to the fifth aspect, the same reference numerals are given and the description is omitted.
[0290] First, the entropy decoding unit 202 decodes an adaptive transform basis selection mode signal from the bitstream (S601). Then, the transform mode determination unit 2064 determines whether the adaptive transform basis selection mode is valid for the decoding target block based on the adaptive transform basis selection mode signal (S602).
[0291] If the adaptive transform basis selection mode is valid (YES in S602), when the horizontal transform size exceeds a certain range (YES in S603), the entropy decoding unit 202 decodes the first basis selection signal in the horizontal direction from the bit stream (S604). On the other hand, when the horizontal transform size is within a certain range (NO in S603), the entropy decoding unit 202 does not decode the first basis selection signal in the horizontal direction. Further, when the vertical transform size exceeds a certain range (YES in S605), the entropy decoding unit 202 decodes the first basis selection signal in the vertical direction from the bit stream (S606). On the other hand, when the vertical transform size is within a certain range (NO in S605), the entropy decoding unit 202 does not decode the first basis selection signal in the vertical direction.
[0292] When the adaptive transform basis selection mode is not valid (NO in S602), the decoding of the first basis selection signal (S604, S606) is skipped.
[0293] Next, the entropy decoding unit 202 decodes the quantization coefficients (S607).
[0294] Here, when the adaptive transform basis selection mode is not valid (NO in S608), the entropy decoding unit 202 decodes the second basis selection signal from the bit stream (S609). On the other hand, when the adaptive transform basis selection mode is valid (YES in S608), the decoding of the second basis selection signal (S609) is skipped.
[0295] Also, in conjunction with the encoding method, the order of each decoding may be determined in advance, and various signals may be decoded so as to be different from the above decoding order. Further, when the second inverse transform is not performed (skipped), the entropy decoding unit 202 may decode a signal indicating that the second inverse transform is not performed from the bit stream, or may decode a signal for selecting a second inverse transform basis equivalent to not performing the transform from the bit stream.
[0296] Note that a decoding method corresponding to the encoding method of the second aspect shown in FIGS. 13A, 13B, and 14 may be employed.
[0297] [Effects, etc.] As described above, according to the decoding apparatus 200 according to the present aspect, the same effects as those of the encoding apparatus 100 according to the second aspect can be achieved.
[0298] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Further, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.
[0299] (Seventh aspect of Embodiment 1) Next, the seventh aspect of Embodiment 1 will be described. In this aspect, when the adaptive transform basis selection mode is not valid, different basic transform bases are used as the first inverse transform basis according to the size of the block to be encoded, which is different from the fifth aspect described above. Note that the decoding apparatus according to this aspect corresponds to the encoding apparatus according to the third aspect. Hereinafter, this aspect will be specifically described with reference to the drawings, centering on the differences from the fifth and sixth aspects.
[0300] Note that the internal configuration of the inverse transform unit 206 of the decoding apparatus 200 according to this aspect is the same as that of the fifth aspect, and thus the illustration thereof is omitted.
[0301] [Processing of the inverse quantization unit and inverse transform unit of the decoding apparatus] The processing of the transform unit 106 and quantization unit 108 of the encoding apparatus 100 according to this aspect will be described with reference to FIG. 23. FIG. 23 is a flowchart showing the processing of the inverse quantization unit 204 and inverse transform unit 206 of the decoding apparatus 200 according to the seventh aspect of Embodiment 1. In FIG. 23, for the processing common to the fifth aspect, the same reference numerals are given and the description thereof is omitted.
[0302] When the adaptive transform basis selection mode is not valid (NO in S505), the size determination unit 2065 determines whether the transform size is within a certain range (S701). That is, the size determination unit 2065 determines whether the horizontal size and the vertical size of the block to be decoded are equal to or less than the second threshold size. Specifically, the size determination unit 2065 determines, for example, whether the product of the horizontal size and the vertical size of the block to be decoded is equal to or less than the threshold.
[0303] Here, when the transform size is within a certain range (YES in S701), the first inverse transform basis selection unit 2066 selects the second basic transform basis as the first inverse transform basis in the horizontal direction and the vertical direction (S702). On the other hand, when the transform size exceeds the certain range (NO in S701), the first inverse transform basis selection unit 2066 selects the first basic transform basis as the first inverse transform basis in the horizontal direction and the vertical direction (S703).
[0304] Note that a decoding method according to the third mode shown in FIG. 16 may be adopted.
[0305] [Effects, etc.] As described above, according to the decoding apparatus 200 according to the present mode, the same effects as those of the encoding apparatus 100 according to the third mode can be achieved.
[0306] [Combination with other modes] This mode may be implemented in combination with at least a part of other modes in the present disclosure. Also, a part of the processing described in the flowchart of this mode, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other modes.
[0307] (Eighth mode of Embodiment 1) Next, the eighth mode of Embodiment 1 will be described. In this mode, an example of decoding various signals related to the first transform and the second transform in the seventh mode will be described. Note that the decoding apparatus according to this mode corresponds to the encoding apparatus according to the fourth mode. Hereinafter, this mode will be specifically described with reference to the drawings, centering on the differences from the fifth to seventh modes.
[0308] Note that the internal configuration of the inverse conversion unit 206 of the decoding device 200 according to this aspect is the same as that of the fifth aspect, and thus the illustration thereof is omitted.
[0309] [Processing of the Entropy Decoding Unit, Inverse Quantization Unit, and Inverse Conversion Unit of the Decoding Device] The processing of the entropy decoding unit 202, inverse quantization unit 204, and inverse conversion unit 206 of the decoding device 200 according to this aspect will be described with reference to FIGS. 24A and 24B. In FIGS. 24A and 24B, for the processing common to any of the fifth to seventh aspects, the same reference numerals are given and the description thereof is omitted.
[0310] The entropy decoding unit 202 determines whether to skip the decoding of the adaptive transform basis selection mode signal (S801). For example, the entropy decoding unit 202 determines to skip the decoding of the adaptive transform basis selection mode signal when any of the following conditions (A) and (B) is satisfied, and determines not to skip the decoding of the adaptive transform basis selection mode signal when not.
[0311] (A) The adaptive transform basis selection mode is not valid.
[0312] (B) The adaptive transform basis selection mode is valid, and all of the following conditions (B1) to (B4) are satisfied.
[0313] (B1) The transform size is equal to or less than the second threshold size W1xH1 used in step S701.
[0314] (B2) The horizontal transform size is equal to or less than the first horizontal threshold size W2 used in step S506.
[0315] (B3) The vertical transform size is equal to or less than the first vertical threshold size H2 used in step S509.
[0316] (B4) The second basic transform basis and the fixed transform bases in the horizontal and vertical directions are the same transform basis.
[0317] As a specific example, when the second threshold size W1xH1 is 4x4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and both the second basic conversion basis and the fixed conversion basis are DST-VII conversion bases, if the conversion size is 4x4 pixels or less, the entropy decoding unit 202 determines that decoding of the adaptive conversion basis selection mode signal is skipped.
[0318] Conversely, when neither of the above conditions (A) and (B) is satisfied, the entropy decoding unit 202 determines that decoding of the adaptive conversion basis selection mode signal is not skipped.
[0319] Here, when it is determined that decoding of the adaptive conversion basis selection mode signal is skipped (YES in S801), the entropy decoding unit 202 skips steps S601 to S606 and decodes the quantization coefficients (S607). On the other hand, when it is determined that decoding of the adaptive conversion basis selection mode signal is not skipped (NO in S801), the entropy decoding unit 202 executes steps S601 to S606 in the same manner as in the sixth aspect and then decodes the quantization coefficients (S207).
[0320] Note that a decoding method corresponding to the encoding method of the fourth aspect shown in FIGS. 17A, 17B, and 18 may be adopted.
[0321] [Effects, etc.] As described above, according to the decoding apparatus 200 according to the present aspect, the same effects as those of the encoding apparatus 100 according to the fourth aspect can be achieved.
[0322] [Combination with other aspects] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.
[0323] (Modification Examples of Each Aspect of Embodiment 1) Note that a signal indicating whether to enable some or all of the processing described in any of the first to eighth aspects may be encoded and decoded. Such a signal may be encoded in units of CU (Coding Unit) or CTU (Coding Tree Unit), or may be encoded in SPS (Sequence Parameter Set), PPS (Picture Parameter Set), or slice units corresponding to the H.265 / HEVC standard.
[0324] Based on the picture type (I, P, B), slice type (I, P, B), conversion size (4x4 pixels, 8x8 pixels, or others), number of non-zero coefficients, quantization parameter, Temporal_id (layer of hierarchical coding), or any combination thereof, the selection of the first conversion basis and the first conversion may be skipped, and the selection of the second conversion basis and the second conversion may be skipped.
[0325] When the encoding device according to the first to fourth aspects performs the above operations, the decoding device according to the fifth to eighth aspects also performs corresponding operations. For example, when the encoding device encodes information indicating whether to enable the process of skipping the first conversion or the second conversion, the decoding device decodes the information and determines whether the first conversion or the second conversion is valid and whether the information indicating the first conversion or the second conversion is encoded.
[0326] (Embodiment 2) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, and the like. Also, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (dedicated circuit).
[0327] Also, the processes described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing may be performed, or distributed processing may be performed.
[0328] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible, and these are also included within the scope of the aspects of the present disclosure.
[0329] Furthermore, here, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, and an image encoding / decoding device having both. Other configurations in the system can be appropriately changed as the case may be.
[0330] [Usage Example] FIG. 25 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, ex110, which are fixed radio stations, are installed in each cell.
[0331] In this content supply system ex100, devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via Internet service provider ex102 or communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be configured to connect by combining any of the above elements. The devices may be directly or indirectly connected to each other via a telephone network or short-range wireless communication without going through base stations ex106 to ex110 which are fixed radio stations. Also, streaming server ex103 is connected to devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via Internet ex101 or the like. Further, streaming server ex103 is connected to terminals within a hotspot in an airplane ex117 via satellite ex116.
[0332] Note that a wireless access point or a hotspot etc. may be used instead of base stations ex106 to ex110. Also, streaming server ex103 may be directly connected to communication network ex104 without going through Internet ex101 or Internet service provider ex102, or may be directly connected to airplane ex117 without going through satellite ex116.
[0333] Camera ex113 is a device capable of taking still images and videos such as a digital camera. Also, smartphone ex115 is generally a smartphone device, mobile phone, or PHS (Personal Handyphone System) etc. that supports the mobile communication system methods generally called 2G, 3G, 3.9G, 4G, and in the future 5G.
[0334] Home appliance ex118 is a device such as a refrigerator or a device included in a household fuel cell cogeneration system.
[0335] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 through a base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) performs the encoding process described in each of the above embodiments on the still image or moving image content photographed by the user using the terminal, multiplexes the video data obtained by encoding with the audio data obtained by encoding the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.
[0336] On the other hand, the streaming server ex103 streams the content data transmitted to the requested client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117 that can decode the encoded data. Each device that has received the distributed data decodes and reproduces the received data. That is, each device functions as an image decoding device according to one aspect of the present disclosure.
[0337] [Distributed Processing] Further, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world and the edge servers. In a CDN, a physically closer edge server is dynamically assigned according to the client. Then, by caching and delivering the content to the edge server, the delay can be reduced. Also, when some error occurs or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the delivery entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the delivery, so high-speed and stable delivery can be realized.
[0338] Moreover, not only the distributed processing of the delivery itself, but also the encoding process of the captured data may be performed on each terminal, on the server side, or they may be shared. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image in units of frames or scenes, or the amount of code is detected. Also, in the second loop, processing is performed to improve the encoding efficiency while maintaining the image quality. For example, by having the terminal perform the first encoding process and the server side that receives the content perform the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode in almost real-time, since the first encoded data performed by the terminal can also be received and played back by other terminals, more flexible real-time delivery becomes possible.
[0339] As another example, cameras such as ex113 perform feature extraction from images, compress data related to the features as metadata, and transmit it to the server. The server performs compression according to the meaning of the image, for example, determines the importance of the object from the features and switches the quantization accuracy. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.
[0340] As yet another example, in a stadium, shopping mall, factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using a plurality of terminals that have performed shooting, and other terminals and servers that have not performed shooting as necessary, encoding processes are respectively assigned and distributed processing is performed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing a picture. This can reduce the delay and achieve more real-time performance.
[0341] Also, since the plurality of video data are substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can refer to each other. Alternatively, the encoded data from each terminal may be received by the server, and the reference relationship may be changed between the plurality of data, or the picture itself may be corrected or replaced and re-encoded. This can generate a stream with improved quality and efficiency for each piece of data.
[0342] Also, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based method, or convert H.264 to H.265.
[0343] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the entity performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.
[0344] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and utilize different scenes photographed by a plurality of terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or images or videos obtained by photographing the same scene from different angles. The videos photographed by each terminal are integrated based on the relative positional relationship between the terminals obtained separately or the regions where the feature points included in the videos match.
[0345] The server may not only encode a two-dimensional moving image, but also automatically or at a time specified by the user, encode a still image based on scene analysis of the moving image and transmit it to the receiving terminal. If the server can further obtain the relative positional relationship between the photographing terminals, it can generate the three-dimensional shape of the scene based not only on the two-dimensional moving image but also on videos of the same scene photographed from different angles. Note that the server may separately encode the three-dimensional data generated by a point cloud or the like, or select or reconstruct the video transmitted to the receiving terminal from the videos photographed by a plurality of terminals based on the results of recognizing or tracking a person or an object using the three-dimensional data.
[0346] In this way, the user can arbitrarily select each video corresponding to each photographing terminal to enjoy the scene, or can also enjoy the content obtained by cutting out the video from an arbitrary viewpoint from the three-dimensional data reconstructed using a plurality of images or videos. Furthermore, similar to the video, sound is also collected from a plurality of different angles, and the server may multiplex and transmit the sound from a specific angle or space according to the video.
[0347] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create right-eye and left-eye viewpoint images respectively, and perform encoding that allows references between each viewpoint video by means of Multi-View Coding (MVC) or the like, or may perform encoding as separate streams without referring to each other. At the time of decoding the separate streams, they may be played back in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.
[0348] In the case of AR images, the server superimposes virtual object information in the virtual space on the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold the virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting the virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may encode it with the α value of the portion other than the object created from the three-dimensional data set to 0 or the like so that the portion is in a transparent state. Or, the server may generate data with an RGB value of a predetermined value set as the background like chroma key and set the portion other than the object to the background color.
[0349] The decryption process of the data delivered in the same way may be performed on each client terminal, on the server side, or may be shared between them. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request may be received by another terminal and decrypted, and the decrypted signal may be transmitted to a device having a display. By dispersing the processing regardless of the performance of the communicable terminals themselves and selecting appropriate content, it is possible to reproduce high-quality data. As another example, while receiving large-size image data on a TV or the like, only a part of the area such as tiles in which the picture is divided may be decrypted and displayed on the viewer's personal terminal. Thereby, while sharing the overall image, it is possible to confirm at hand the area of one's own field of responsibility or the area to be confirmed in more detail.
[0350] In the future, it is expected that, regardless of indoors or outdoors, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used, by using a delivery system standard such as MPEG-DASH, appropriate data can be switched for the ongoing communication and content can be received seamlessly. As a result, the user can be switched in real time while freely selecting not only their own terminal but also a decryption device or display device such as a display installed indoors or outdoors. Also, based on their own location information or the like, decryption can be performed while switching the terminal to be decrypted and the terminal to be displayed. Thereby, it is also possible to move while displaying map information on a part of the wall surface or ground of the adjacent building in which the displayable device is embedded during the movement to the destination. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or being copied to an edge server in a content delivery service, it is also possible to switch the bit rate of the received data.
[0351] [Scalable Encoding] Regarding content switching, an explanation will be given using a scalable stream compressed and encoded by applying the moving image encoding method shown in each of the above embodiments, as shown in FIG. 26. The server may have a plurality of streams with the same content but different qualities as individual streams, but it may also be configured to switch content by taking advantage of the characteristics of a temporally / spatially scalable stream realized by encoding in layers as shown in the figure. That is, by determining up to which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth on the decoding side, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on your smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0352] Furthermore, as described above, in addition to the configuration that realizes scalability in which pictures are encoded for each layer and an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of the image, and the decoding side may generate high-quality content by super-resolving the pictures of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in the resolution. The meta information includes information for specifying linear or non-linear filter coefficients used in super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used in super-resolution processing.
[0353] Alternatively, the picture may be divided into tiles or the like according to the meaning such as an object in the image, and the decoding side may be configured to decode only a part of the area by selecting the tile to be decoded. Further, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 27, the meta information is stored using a data storage structure different from the pixel data such as the SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.
[0354] Further, the meta information may be stored in a unit composed of a plurality of pictures such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining it with the information in picture units, the picture in which the object exists and the position of the object in the picture can be specified.
[0355] [Optimization of Web Page] FIG. 28 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 29 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 28 and 29, a web page may include a plurality of link images that are links to image contents, and the appearance thereof may be different depending on the device to be browsed. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) displays a still image or an I picture that each content has as a link image, displays a video like a gif animation with a plurality of still images or I pictures, etc., or receives only the base layer and decodes and displays the video.
[0356] When a user selects a linked image, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Also, in order to ensure real-time performance, before selection or when the communication bandwidth is very strict, the display device can reduce the delay (the delay from the start of content decoding to the start of display) between the decoding time and the display time of the leading picture by decoding and displaying only forward reference pictures (I pictures, P pictures, B pictures with only forward reference). Further, the display device may deliberately ignore the reference relationship of the pictures, roughly decode all B pictures and P pictures as forward references, and perform normal decoding as the received pictures increase over time.
[0357] [Autonomous driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for the autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.
[0358] In this case, since a vehicle, drone, airplane, etc. including the receiving terminal moves, the receiving terminal can realize seamless reception and decoding by transmitting the position information of the receiving terminal at the time of a reception request while switching between base stations ex106 to ex110. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication bandwidth.
[0359] In the above manner, in the content supply system ex100, the client can receive, decode, and play back the encoded information transmitted by the user in real time.
[0360] [Distribution of personal content] In addition, in the content supply system ex100, not only high-quality and long-duration content by video distributors but also unicast or multicast distribution of low-quality and short-duration content by individuals is possible. Also, it is considered that such individual content will increase in the future. In order to make individual content better, the server may perform encoding processing after performing editing processing. This can be realized, for example, with the following configuration.
[0361] At the time of shooting in real time or by accumulation and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection from the original image or encoded data. Then, based on the recognition result, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone, etc. The server encodes the edited data based on the editing result. Also, it is known that the viewing rate decreases if the shooting time is too long. The server may automatically clip not only less important scenes but also scenes with little movement, etc., so that the content is within a specific time range according to the shooting time, based on the image processing result. Or, the server may generate and encode a digest based on the result of semantic analysis of the scene.
[0362] Note that in some cases, personal content may contain elements that, as they are, would infringe copyright, moral rights of the author, or the right of portrait, etc., and may be inconvenient for individuals, such as when the sharing range exceeds the intended range. Therefore, for example, the server may deliberately change the image so that the faces of people in the peripheral part of the screen or the inside of a house are out of focus and then encode it. Also, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Alternatively, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user designates a person or background area that the user wants to process the image, and the server can perform processing such as replacing the designated area with another video or blurring the focus. In the case of a person, the video of the face part can be replaced while tracking the person in the moving image.
[0363] In addition, since the viewing of personal content with a small data volume has a strong requirement for real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and decodes and plays it. During this period, the decoding device receives the enhancement layer, and when the playback is looped or played more than twice, it may play a high-quality video including the enhancement layer. For a stream with scalable encoding like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played for the first time and a second stream encoded with reference to the first video are configured as one stream.
[0364] [Other Usage Examples] Also, these encoding or decoding processes are generally processed by the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. In addition, software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processing may be performed using the software. Further, when the smartphone ex115 has a camera, video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.
[0365] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays the content.
[0366] Also, not limited to the content supply system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above embodiments can be incorporated into a digital broadcast system. In order to transmit and receive multiplexed data in which video and audio are multiplexed on a broadcast radio wave using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the unicast-friendly configuration of the content supply system ex100, but the same application is possible for encoding and decoding processing.
[0367] [Hardware Configuration] FIG. 30 is a diagram showing smartphone ex115. Further, FIG. 31 is a diagram showing a configuration example of smartphone ex115. Smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by camera unit ex465 and video received by antenna ex450. Smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with SIM ex468 for identifying the user and authenticating access to various data including the network. Note that an external memory may be used instead of memory unit ex467.
[0368] Also, a main control unit ex460 for comprehensively controlling display unit ex458, operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, slot unit ex464, and memory unit ex467 are connected via bus ex470.
[0369] When the power key is turned on by the user's operation, power supply circuit unit ex461 supplies power to each unit from the battery pack to activate smartphone ex115 to an operable state.
[0370] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified, frequency conversion processing and analog-to-digital conversion processing are performed, spectral inverse spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. In the data communication mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body unit, and the same transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while the camera unit ex465 is imaging video or still images, etc., and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.
[0371] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a moving image decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and a video or a still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming has become widespread, there may be situations where it is not socially appropriate to play audio depending on the user's circumstances. Therefore, as an initial value, it is desirable to have a configuration that plays only the video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.
[0372] Also, although the smartphone ex115 has been described as an example here, as the terminal, in addition to a transceiver type terminal having both an encoder and a decoder, three implementation forms are conceivable: a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcast system, although it has been described as receiving or transmitting multiplexed data in which audio data and the like are multiplexed in video data, in the multiplexed data, character data related to the video or the like may be multiplexed in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.
[0373] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, many terminals also have a GPU. Therefore, a configuration in which a wide area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU, or a memory whose address is managed so that it can be used in common, may be used. As a result, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform quantization in units such as pictures using the GPU instead of the CPU.
Industrial Applicability
[0374] The present disclosure can be used, for example, in a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, or a digital video camera.
Explanation of Signs
[0375] 100 Encoding device 102 Splitting unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 200 Decoding device 202 Entropy decoding unit 1061, 2064 Transformation mode determination unit 1062, 2065 Size determination unit 1063 First transformation basis selection unit 1064 First conversion unit 1065 Second conversion execution determination unit 1066 Second conversion basis selection unit 1067 Second conversion unit 1141, 2062 Second inverse conversion basis selection unit 1142, 2063 Second inverse conversion unit 1143, 2066 First inverse conversion basis selection unit 1144, 2067 First inverse conversion unit 2061 Second inverse conversion execution determination unit 2062S Second basis selection signal 2064S Adaptive conversion basis selection mode signal 2065S Size signal 2066S First basis selection signal
Claims
1. An encoding apparatus comprising a circuit and a memory, wherein the circuit uses the memory to determine whether a mode of selecting a transform basis according to the size of a block to be encoded is effective, and when the mode is effective, when the vertical size of the block to be encoded is larger than a threshold size, select a first transform basis from among a plurality of transform basis candidates as the vertical transform basis, when the vertical size of the block to be encoded is smaller than the threshold size, select a second transform basis, which is a fixed transform basis, as the vertical transform basis, generate first transform coefficients by performing a first vertical transform on the residual of the block to be encoded using the selected vertical transform basis, generate second transform coefficients by performing a second transform on the first transform coefficients, generate a bit stream including information indicating whether the mode is effective, an encoding apparatus.
2. A decoding apparatus comprising a circuit and a memory, wherein the circuit uses the memory to generate transform coefficients by performing a second inverse transform on the coefficients of a block to be decoded, determine whether a mode of selecting a transform basis according to the size of the block to be decoded is effective, and when the mode is effective, when the vertical size of the block to be decoded is larger than a threshold size, select a first inverse transform basis from among a plurality of inverse transform basis candidates as the vertical inverse transform basis, when the vertical size of the block to be decoded is smaller than the threshold size, select a second inverse transform basis, which is a fixed inverse transform basis, as the vertical inverse transform basis, generate a prediction residual by performing a first vertical inverse transform on the transform coefficients of the block to be decoded using the selected vertical inverse transform basis, a decoding apparatus.
3. Determine whether a mode of selecting a transform basis according to the size of a block to be encoded is effective, and when the mode is effective, when the vertical size of the block to be encoded is larger than a threshold size, select a first transform basis from among a plurality of transform basis candidates as the vertical transform basis, when the vertical size of the block to be encoded is smaller than the threshold size, select a second transform basis, which is a fixed transform basis, as the vertical transform basis, generate first transform coefficients by performing a first vertical transform on the residual of the block to be encoded using the selected vertical transform basis, Generating a second conversion coefficient by performing a second conversion on the first conversion coefficient, Generating a bit stream including information indicating whether the mode is valid, Encoding method.
4. Generating a conversion coefficient by performing a second inverse conversion on the coefficients of the block to be decoded, Determining whether a mode of selecting a conversion basis according to the size of the block to be decoded is valid, When the mode is valid, When the vertical size of the block to be decoded is larger than a threshold size, selecting a first inverse conversion basis as the vertical inverse conversion basis from among a plurality of inverse conversion basis candidates, When the vertical size of the block to be decoded is smaller than the threshold size, selecting a second inverse conversion basis, which is a fixed inverse conversion basis, as the vertical inverse conversion basis, Generating a prediction residual by performing a first inverse conversion in the vertical direction on the conversion coefficient of the block to be decoded using the selected vertical inverse conversion basis, Decoding method.
Citation Information
Patent Citations
Enhanced multiple transforms for prediction residual
US20160219290A1
Non-separable secondary transform for video coding
US20170094313A1
Method and Apparatus of Adaptive Multiple Transforms for Video Coding
US20180332289A1
Image decoding device and image encoding device
WO2017195555A1