Coding device, decoding device and storage medium
By adopting the adaptive transform base selection mode in the video encoding and decoding device, and selecting the transform base according to the object block size for transformation, the problems of insufficient compression efficiency and large processing load in the prior art are solved, and more efficient video encoding and decoding are achieved.
Patent Information
- Application Number
- CN202211053435.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-28
- Filing Date
- 2018-12-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2038-12-19
AI Technical Summary
In the existing video encoding technology, the compression efficiency is insufficient and the processing load is large, making it difficult to meet the needs of higher compression efficiency and processing efficiency.
By implementing the adaptive transform base selection mode in the encoding and decoding device, an appropriate transform base is selected according to the size of the encoded or decoded target block, the first transform and (or) the second transform are performed to generate a transform coefficient or a restore residual.
It improves the compression efficiency of video encoding, reduces processing load, and realizes a more efficient encoding and decoding process.
Smart Images

Figure CN115278239B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an encoding device, a decoding device and a storage medium. Background Art
[0002] The video coding standard specification called HEVC (High-Efficiency Video Coding) is standardized by JCT-VC (Joint Collaborative Team on Video Coding).
[0003] Prior art literature
[0004] Non-patent literature
[0005] Non-patent document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding))
[0006] Non-patent literature 2: Jianle Chen et al., Algorithm Description of JointExploration Test Model 5 (JEM5), Joint Video Exploration Team (JVET) of ITU-T SG16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 5th Meeting: Geneva, CH, Document: JVET-E1001, January 2017 Summary of the invention
[0007] Problems to be solved by the invention
[0008] In such encoding and decoding technologies, further improvement of compression efficiency and reduction of processing load are required.
[0009] Therefore, the present invention provides an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement in compression efficiency and reduction in processing load.
[0010] Means for solving problems
[0011] A coding device according to a technical solution of the present invention comprises a circuit and a memory, wherein the circuit uses the memory to perform the following processing: determining whether a mode for selecting a transform base according to the size of an encoding object block is valid; if the mode is valid, when the horizontal size of the encoding object block is larger than a threshold size, selecting a first transform base from a plurality of transform base candidates as a transform base in the horizontal direction; when the horizontal size of the encoding object block is smaller than the threshold size, selecting a second transform base as a transform base in the horizontal direction, wherein the second transform base is a fixed transform base; and generating a first transform coefficient by performing a first transform on a residual of the encoding object block using the selected transform base in the horizontal direction.
[0012] A decoding device according to a technical solution of the present invention comprises a circuit and a memory, wherein the circuit uses the memory to perform the following processing: determining whether a mode for selecting an inverse transform base according to the size of a decoding object block is valid; if the mode is valid, when the horizontal size of the decoding object block is larger than a threshold size, selecting a first inverse transform base from multiple inverse transform base candidates as an inverse transform base in the horizontal direction; when the horizontal size of the decoding object block is smaller than the threshold size, selecting a second inverse transform base as an inverse transform base in the horizontal direction, wherein the second inverse transform base is a fixed inverse transform base; and generating a prediction residual by performing a first inverse transform on coefficients of the decoding object block using the selected inverse transform base in the horizontal direction.
[0013] In addition, these general or specific technical solutions may also be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0014] Effects of the Invention
[0015] The present invention can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement in compression efficiency and reduction in processing load. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a block diagram showing the functional structure of the encoding device related to embodiment 1.
[0017] Figure 2 This is a diagram showing an example of block division according to the first embodiment.
[0018] Figure 3 This is a table showing the transformation basis functions corresponding to each transformation type.
[0019] Figure 4A This is a diagram showing an example of the shape of a filter used in ALF.
[0020] Figure 4B FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0021] Figure 4C FIG. 1 is a diagram showing another example of the shape of the filter used in ALF.
[0022] Figure 5A This is a diagram showing 67 intra prediction modes of intra prediction.
[0023] Figure 5B This is a flowchart for explaining the outline of the predicted image correction process based on the OBMC process.
[0024] Figure 5C This is a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.
[0025] Figure 5D This is a diagram showing an example of FRUC.
[0026] Figure 6 This is a diagram for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0027] Figure 7 This is a diagram used to explain pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0028] Figure 8 This is a diagram for explaining a model assuming uniform linear motion.
[0029] Fig. 9A This is a diagram for explaining the derivation of a motion vector in a sub-block unit based on motion vectors of a plurality of adjacent blocks.
[0030] Fig. 9B This is a diagram for explaining the outline of motion vector derivation processing based on the merge mode.
[0031] Fig. 9C This is a conceptual diagram used to explain the outline of DMVR processing.
[0032] Fig.9D This is a diagram for explaining an overview of a method for generating a predicted image using a brightness correction process based on an LIC process.
[0033] Fig.10 This is a block diagram showing the functional structure of a decoding device according to Embodiment 1.
[0034] Fig.11A This is a block diagram showing the internal structure of the transformation unit of the encoding device according to the first mode of implementation mode 1.
[0035] Fig. 11B This is a block diagram showing the internal structure of the inverse transform unit of the encoding device according to the first mode of implementation mode 1.
[0036] Fig. 12A This is a flowchart showing the processing of the transform unit and the quantization unit of the encoding device according to the first mode of implementation mode 1.
[0037] Fig. 12B This is a flowchart showing a modified example of the processing of the transform unit and the quantization unit of the encoding device according to the first aspect of the first embodiment.
[0038] Fig.13A This is a flowchart showing the processing of the transform unit and the quantization unit of the encoding device according to the second embodiment of the first embodiment.
[0039] Fig. 13B This is a flowchart showing the processing of the entropy coding unit of the coding device according to the second mode of implementation mode 1.
[0040] Fig.14 This is a diagram showing a specific example of syntax in the second aspect related to Implementation Method 1.
[0041] Fig.15 This is a table showing specific examples of the transform base used in the second mode of implementation mode 1 and the presence or absence of encoding of the signal.
[0042] Fig.16 This is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the third aspect of implementation mode 1.
[0043] Fig.17A This is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the fourth embodiment of the first embodiment.
[0044] Fig. 17B This is a flowchart showing the processing of the entropy coding unit of the coding device according to the fourth mode of implementation mode 1.
[0045] Fig.18 This is a diagram showing a specific example of syntax in the fourth aspect of implementation mode 1.
[0046] Fig.19 This is a table showing specific examples of the transform base used in the fourth mode of implementation mode 1 and the presence or absence of encoding of the signal.
[0047] Fig. 20 This is a block diagram showing the internal structure of the inverse transformation unit of the decoding device according to the fifth mode of implementation mode 1.
[0048] Fig.21 This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the fifth mode of implementation mode 1.
[0049] Fig.22A This is a flowchart showing the processing of the entropy decoding unit of the decoding device according to the sixth mode of implementation mode 1.
[0050] Fig. 22B This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the sixth mode of implementation mode 1.
[0051] Fig.23 This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the seventh mode of implementation mode 1.
[0052] Fig.24A This is a flowchart showing the processing of the entropy decoding unit of the decoding device according to the eighth mode of implementation mode 1.
[0053] Fig. 24B This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the eighth mode of implementation mode 1.
[0054] Fig.25 It is an overall structural diagram of the content supply system that realizes content distribution services.
[0055] Fig.26 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0056] Fig. 27 This is a diagram showing an example of a coding structure in the case of scalable coding.
[0057] Fig.28 This is a diagram showing an example of a display screen of a web page.
[0058] Fig.29 This is a diagram showing an example of a display screen of a web page.
[0059] Fig.30 The diagram shows an example of a smart phone.
[0060] Fig.31 This is a block diagram showing a structural example of a smart phone. DETAILED DESCRIPTION
[0061] Hereinafter, the embodiments will be described in detail with reference to the drawings.
[0062] In addition, the embodiments described below are all inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the claims. In addition, among the components of the following embodiments, components that are not described in the independent claims representing the highest concept are described as arbitrary components.
[0063] (Implementation Method 1)
[0064] First, the outline of Embodiment 1 is described as an example of an encoding device and a decoding device to which the processing and / or configuration described in each embodiment of the present invention described later can be applied. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processing and / or configuration described in each embodiment of the present invention can be applied, and the processing and / or configuration described in each embodiment of the present invention can also be implemented in an encoding device and a decoding device different from Embodiment 1.
[0065] When the processing and / or configuration described in each aspect of the present invention is applied to Embodiment 1, for example, any of the following may be performed.
[0066] (1) In the encoding device or decoding device of Embodiment 1, components corresponding to the components described in each aspect of the present invention among a plurality of components constituting the encoding device or decoding device are replaced with the components described in each aspect of the present invention;
[0067] (2) In the encoding device or decoding device of Embodiment 1, after making any changes such as addition, replacement, deletion, etc. to the functions or the processes to be performed on some of the multiple components constituting the encoding device or decoding device, the components corresponding to the components described in each aspect of the present invention are replaced with the components described in each aspect of the present invention;
[0068] (3) With respect to the method implemented by the encoding device or decoding device of Embodiment 1, after adding a process and / or replacing or deleting a part of the multiple processes included in the method, the process corresponding to the process described in each aspect of the present invention is replaced with the process described in each aspect of the present invention;
[0069] (4) implementing the method by combining some of the plurality of components constituting the encoding device or decoding device of Embodiment 1 with the components described in the various aspects of the present invention, components having some of the functions possessed by the components described in the various aspects of the present invention, or components implementing some of the processing implemented by the components described in the various aspects of the present invention;
[0070] (5) A component having a part of the functions possessed by some of the multiple components constituting the encoding device or decoding device of Embodiment 1, or a component implementing a part of the processing implemented by some of the multiple components constituting the encoding device or decoding device of Embodiment 1, is implemented in combination with a component described in each aspect of the present invention, a component having a part of the functions possessed by the component described in each aspect of the present invention, or a component implementing a part of the processing implemented by the component described in each aspect of the present invention;
[0071] (6) In a method implemented by the encoding device or decoding device of Embodiment 1, among a plurality of processes included in the method, a process corresponding to the process described in each aspect of the present invention is replaced with the process described in each aspect of the present invention;
[0072] (7) Some of the multiple processes included in the method implemented by the encoding device or decoding device of Embodiment 1 are combined with the processes described in each aspect of the present invention and implemented.
[0073] In addition, the implementation method of the processing and / or structure described in each embodiment of the present invention is not limited to the above-mentioned example. For example, it can also be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, and the processing and / or structure described in each embodiment can be implemented separately. In addition, the processing and / or structure described in different embodiments can also be combined and implemented.
[0074] [Overview of Encoding Device]
[0075] First, an overview of the encoding device according to Embodiment 1 will be described. Figure 1 1 is a block diagram showing a functional structure of a coding device 100 according to Embodiment 1. The coding device 100 is a moving picture / image coding device that encodes a moving picture / image in units of blocks.
[0076] like Figure 1 As shown, the encoding device 100 is a device that encodes an image in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy coding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126 and a prediction control unit 128.
[0077] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. In addition, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128.
[0078] Hereinafter, each component included in the encoding device 100 will be described.
[0079] [Division]
[0080] The segmentation unit 102 segments each picture included in the input motion image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the picture into blocks of a fixed size (e.g., 128×128). The fixed-size blocks are sometimes called coding tree units (CTUs). Furthermore, the segmentation unit 102 segments each fixed-size block into blocks of a variable size (e.g., less than 64×64) based on recursive quadtree and / or binary tree block segmentation. The variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in the present embodiment, there is no need to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture may be used as processing units of CUs, PUs, or TUs.
[0081] Figure 2 FIG. 1 is a diagram showing an example of block division in Implementation 1. Figure 2 In FIG, the solid line represents the block boundary based on the quadtree block partition, and the dotted line represents the block boundary based on the binary tree block partition.
[0082] Here, the block 10 is a square block of 128×128 pixels (128×128 block). The 128×128 block 10 is first divided into four square 64×64 blocks (quadtree block division).
[0083] The upper left 64×64 block is further divided vertically into two rectangular 32×64 blocks, and the left 32×64 block is further divided vertically into two rectangular 16×64 blocks (binary tree block division). As a result, the upper left 64×64 block is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.
[0084] The upper right 64×64 block is horizontally partitioned into two rectangular 64×32 blocks 14 and 15 (binary tree block partitioning).
[0085] The 64×64 block at the bottom left is divided into four square 32×32 blocks (quadtree block division). The upper left block and the lower right block of the four 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the 16×32 block on the right is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the lower left 64×64 block is divided into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.
[0086] The lower right 64×64 block 23 is not split.
[0087] As above, in Figure 2 In FIG. 1 , block 10 is divided into 13 blocks 11 to 23 of variable size based on recursive quad-tree and binary tree block division. This type of division is sometimes called QTBT (quad-tree plus binary tree) division.
[0088] In addition, Figure 2 In the example, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary tree division). The division including such ternary tree division is called MBT (multi type tree) division.
[0089] [Subtraction Department]
[0090] The subtracting unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in units of blocks divided by the dividing unit 102. That is, the subtracting unit 104 calculates the prediction error (also referred to as residual) of the encoding target block (hereinafter referred to as the current block). Then, the subtracting unit 104 outputs the calculated prediction error to the transforming unit 106.
[0091] The original signal is an input signal of the encoding device 100 and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing an image may also be referred to as a sample.
[0092] [Conversion Department]
[0093] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.
[0094] In addition, the transform unit 106 may adaptively select a transform type from a plurality of transform types, and transform the prediction error into a transform coefficient using a transform basis function corresponding to the selected transform type. Such a transform is sometimes called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0095] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. Figure 3 In , N represents the number of input pixels. The selection of a transform type from among these multiple transform types may depend on the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode, for example.
[0096] Information indicating whether such EMT or AMT is applied (e.g., called an AMT flag) and information indicating the selected transform type are signaled at the CU level. In addition, the signaling of such information does not need to be limited to the CU level, but can also be other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0097] In addition, the transformation unit 106 may also re-transform the transformation coefficients (transformation results). Such re-transformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 re-transforms each sub-block (for example, 4×4 sub-blocks) contained in the block of transformation coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transformation matrix used in NSST are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but can also be other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0098] Here, separable transformation refers to a method of performing multiple transformations in each direction according to the number of dimensions of the input, and non-separable transformation refers to a method of treating two or more dimensions as one dimension and performing transformations together when the input is multi-dimensional.
[0099] For example, as an example of non-separable transformation, when a 4×4 block is input, it is considered as an array of 16 elements, and the array is transformed using a 16×16 transformation matrix.
[0100] Similarly, a method (Hypercube Givens Transform) of treating a 4×4 input block as an array of 16 elements and performing multiple Givens rotations on the array is also an example of a Non-Separable transformation.
[0101] [Quantitative Department]
[0102] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on the quantization parameters (QP) corresponding to the scanned transform coefficients. The quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.
[0103] The predetermined order is the order for quantization / inverse quantization of transform coefficients. For example, the predetermined scanning order is defined in ascending order (order from low frequency to high frequency) or descending order (order from high frequency to low frequency) of frequency.
[0104] The quantization parameter refers to a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0105] [Entropy coding unit]
[0106] The entropy coding unit 110 generates a coded signal (coded bit stream) by performing variable length coding on the quantized coefficients that are input from the quantization unit 108. Specifically, the entropy coding unit 110 binarizes the quantized coefficients and performs arithmetic coding on the binary signal, for example.
[0107] [Inverse Quantization Unit]
[0108] The inverse quantization unit 112 inversely quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantized coefficients of the current block in a predetermined scanning order. The inverse quantization unit 112 outputs the inversely quantized transform coefficients of the current block to the inverse transform unit 114.
[0109] [Inverse transformation unit]
[0110] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform performed by the transform unit 106. The inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0111] In addition, since the restored prediction error loses information due to quantization, it does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error includes a quantization error.
[0112] [Addition Department]
[0113] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. Then, the adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.
[0114] [Block Memory]
[0115] The block memory 118 is a storage unit for storing blocks in a picture to be coded (hereinafter referred to as a current picture) to be referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the adding unit 116 .
[0116] [Loop filter section]
[0117] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within a coding loop (in-loop filtering), such as deblocking filtering (DF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).
[0118] In ALF, a least square error filter is used to remove coding distortion. For example, for each 2×2 sub-block in the current block, one filter is selected from a plurality of filters based on the direction and activity of the local gradient.
[0119] Specifically, first, the sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C (e.g., C=5D+A) is calculated. And, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).
[0120] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (eg, horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantizing the added result.
[0121] Based on the result of such classification, a filter to be used for a sub-block is determined from among a plurality of filters.
[0122] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figure 4A to Figure 4C The diagrams show a plurality of examples of filter shapes used in ALF. Figure 4A represents a 5×5 diamond-shaped filter, Figure 4B represents a 7×7 diamond shape filter, Figure 4C Indicates a 9×9 diamond filter. The information indicating the shape of the filter is signaled at the picture level. In addition, the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, but can also be other levels (for example, sequence level, slice level, tile level, CTU level or CU level).
[0123] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, regarding brightness, whether to use ALF is determined at the CU level, and regarding color difference, whether to use ALF is determined at the picture level. Information indicating whether ALF is on / off is signaled at the picture level or the CU level. In addition, the signaling of information indicating whether ALF is on / off does not need to be limited to the picture level or the CU level, and can also be other levels (for example, the sequence level, the slice level, the tile level, or the CTU level).
[0124] The coefficient sets of the selectable multiple filters (e.g., filters up to 15 or 25) are signaled at the picture level. In addition, the signaling of the coefficient set does not need to be limited to the picture level, and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0125] [Frame Memory]
[0126] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120 .
[0127] [Intra-frame prediction unit]
[0128] The intra prediction unit 124 performs intra prediction (also referred to as intra-screen prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 124 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., brightness values and color difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0129] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0130] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0131] The plurality of directional prediction modes include, for example, 33 directional prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 5A This is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows indicate 33 directions specified by the H.265 / HEVC standard, and dotted arrows indicate 32 additional directions.
[0132] In addition, in the intra-frame prediction of the chrominance block, the luminance block can also be referenced. That is, the chrominance component of the current block can also be predicted based on the luminance component of the current block. Such intra-frame prediction is called CCLM (cross-component linear model) prediction. Such an intra-frame prediction mode of the chrominance block with reference to the luminance block (for example, called CCLM mode) can also be added as one of the intra-frame prediction modes of the chrominance block.
[0133] The intra prediction unit 124 may also modify the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction. Intra prediction accompanied by such modification is called PDPC (position dependent intraprediction combination). Information indicating whether PDPC is adopted (e.g., called PDPC flag) is signaled, for example, at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, but may also be other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0134] [Inter-frame prediction unit]
[0135] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-screen prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 126 performs motion estimation in the reference picture for the current block or sub-block. Furthermore, the inter-frame prediction unit 126 performs motion compensation using motion information (e.g., motion vector) obtained by motion estimation, thereby generating an inter-frame prediction signal for the current block or sub-block. Furthermore, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.
[0136] The motion information used in motion compensation is signaled. A predicted motion vector (motion vector predictor) may be used to signalize the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0137] In addition, it is also possible to generate an inter-frame prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of the adjacent blocks. Specifically, a prediction signal based on the motion information obtained by motion estimation and a prediction signal based on the motion information of the adjacent blocks may be weighted and added to generate an inter-frame prediction signal in units of sub-blocks within the current block. Such inter-frame prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0138] In such an OBMC mode, information indicating the size of a sub-block used for OBMC (e.g., OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is adopted (e.g., OBMC flag) is signaled at the CU level. In addition, the signaling level of this information does not need to be limited to the sequence level and the CU level, and may be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0139] The OBMC mode will be described in more detail. Figure 5B and Figure 5C It is a flowchart and a conceptual diagram for explaining the outline of the predicted image correction processing based on the OBMC processing.
[0140] First, a predicted image (Pred) obtained by normal motion compensation is acquired using a motion vector (MV) assigned to a coding target block.
[0141] Next, a prediction image (Pred_L) is obtained by using the motion vector (MV_L) of the already encoded left adjacent block for the encoding target block, and the prediction image and Pred_L are weighted and superimposed to perform the first correction of the prediction image.
[0142] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is used to obtain the predicted image (Pred_U) for the encoding object block, and the predicted image after the first correction mentioned above and Pred_U are weighted and superimposed to perform the second correction on the predicted image, which is used as the final predicted image.
[0143] In addition, the method of performing two-stage correction using the left neighboring block and the upper neighboring block has been described here, but the method may be configured to perform correction more times than two stages using the right neighboring block and the lower neighboring block.
[0144] Furthermore, the area to be superimposed may not be the pixel area of the entire block, but may be only a partial area near the block boundary.
[0145] In addition, the predicted image correction process based on one reference picture is explained here, but the same is true when the predicted image is corrected based on multiple reference pictures. After obtaining the corrected predicted images based on each reference picture, the obtained predicted images are further superimposed to serve as the final predicted image.
[0146] Furthermore, the processing target block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.
[0147] As a method for determining whether to use the OBMC process, there is a method using, for example, obmc_flag, a signal indicating whether the OBMC process is used. As a specific example, in the encoding device, it is determined whether the encoding target block belongs to a complex motion region. If it belongs to a complex motion region, the obmc_flag is set to a value of 1 and the OBMC process is used for encoding. If it does not belong to a complex motion region, the obmc_flag is set to a value of 0 and the OBMC process is not used for encoding. On the other hand, in the decoding device, the obmc_flag described in the stream is decoded, and whether the OBMC process is used is switched according to the value thereof, thereby performing decoding.
[0148] In addition, the motion information may be derived on the decoding device side without being signaled. For example, a merge mode specified by the H.265 / HEVC specification may be used. In addition, the motion information may be derived, for example, by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.
[0149] Here, a mode for performing motion estimation on the decoding device side is described. The mode for performing motion estimation on the decoding device side is called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0150] exist Figure 5D An example of FRUC processing is shown in FIG. First, a list of multiple candidates each having a predicted motion vector is generated (which can also be shared with a merge list) with reference to the motion vectors of the coded blocks that are spatially or temporally adjacent to the current block. Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0151] And, based on the motion vector of the selected candidate, the motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as the motion vector for the current block. In addition, for example, the motion vector for the current block can be derived by performing pattern matching in the surrounding area of the position in the reference picture corresponding to the selected candidate motion vector. That is, the surrounding area of the best candidate MV can also be searched by the same method, and when there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and it is used as the final MV of the current block. In addition, a structure that does not implement this process can also be made.
[0152] Completely the same processing can be performed even when processing is performed in sub-block units.
[0153] The evaluation value is calculated by obtaining a difference value of the reconstructed image through pattern matching between a region in the reference picture corresponding to the motion vector and a predetermined region. In addition to the difference value, other information may be used to calculate the evaluation value.
[0154] As pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are sometimes called bidirectional matching and template matching, respectively.
[0155] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures along the motion trajectory of the current block. Therefore, in the first pattern matching, as the prescribed area for calculating the candidate evaluation value, an area in another reference picture along the motion trajectory of the current block is used.
[0156] Figure 6 is a diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory. Figure 6 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among the pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV with the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs.
[0157] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0158] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the prescribed area for calculating the candidate evaluation value described above.
[0159] Figure 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 7 As shown, in the second pattern matching, the motion vector of the current block is derived by searching the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded area of both or one of the left adjacent and upper adjacent areas and the reconstructed image at the same position in the coded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value is selected as the best candidate MV among multiple candidate MVs.
[0160] Such information indicating whether the FRUC mode is adopted (for example, called a FRUC flag) is signaled at the CU level. In addition, when the FRUC mode is adopted (for example, when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and can also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level or sub-block level).
[0161] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion is described. This mode is sometimes called BIO (bi-directional optical flow).
[0162] Figure 8 This is a diagram used to illustrate a model that assumes uniform linear motion. Figure 8 In (v x , v y ) represents the velocity vector, τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.
[0163] At this time, the velocity vector (v x , v y ) is assumed to be a uniform linear motion, (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1) respectively, and the following optical flow equation (1) holds.
[0164] [Formula 1]
[0165]
[0166] Here, I (k) Represents the brightness value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation indicates that the sum of (i) the temporal differential of the brightness value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of the block unit obtained from the merge list, etc. is corrected in pixel units.
[0167] In addition, the motion vector may be derived on the decoding device side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0168] Here, a mode for deriving a motion vector in sub-block units based on motion vectors of a plurality of adjacent blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.
[0169] Fig. 9A FIG. 1 is a diagram for explaining the derivation of a motion vector for a sub-block based on motion vectors of a plurality of adjacent blocks. Fig. 9AIn the example, the current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v0 of the upper left corner control point of the current block is derived, and based on the motion vectors of the adjacent sub-blocks, the motion vector v1 of the upper right corner control point of the current block is derived. Furthermore, using the two motion vectors v0 and v1, the motion vectors (v x , v y ).
[0170] [Formula 2]
[0171]
[0172] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a preset weight coefficient.
[0173] In such an affine motion compensation prediction mode, several modes in which the motion vectors of the upper left and upper right control points are derived in different ways may be included. Information indicating such an affine motion compensation prediction mode (e.g., called an affine flag) is signaled at the CU level. In addition, the signaling of the information indicating the affine motion compensation prediction mode does not need to be limited to the CU level, but may be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0174] [Prediction Control Department]
[0175] The prediction control unit 128 selects one of the intra prediction signal and the inter prediction signal, and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.
[0176] Here, an example of deriving a motion vector of a picture to be encoded by using the merge mode is described. Fig. 9B This is a diagram for explaining the outline of motion vector derivation processing based on the merge mode.
[0177] First, a prediction MV list is generated in which candidates for prediction MVs are registered. Candidates for prediction MVs include MVs of multiple coded blocks located spatially around the target block, that is, spatial neighbor prediction MVs; MVs of blocks near the position of the target block in the coded reference picture, that is, temporal neighbor prediction MVs; MVs generated by combining the MV values of the spatial neighbor prediction MV and the temporal neighbor prediction MV, that is, combined prediction MVs; and MVs with a value of zero, that is, zero prediction MVs.
[0178] Next, one predicted MV is selected from among a plurality of predicted MVs registered in the predicted MV list, thereby determining the MV of the encoding target block.
[0179] Furthermore, in the variable length coding unit, merge_idx, which is a signal indicating which predicted MV is selected, is described in the stream and encoded.
[0180] In addition, Fig. 9B The predicted MVs registered in the predicted MV list described in the figure are just examples, and may be a number different from that in the figure, or a structure that does not include some types of predicted MVs in the figure, or a structure that adds predicted MVs other than the types of predicted MVs in the figure.
[0181] In addition, the MV of the encoding target block derived by the merge mode may be used to perform the DMVR process described later, thereby determining the final MV.
[0182] Here, an example of determining MV using DMVR processing is described.
[0183] Fig. 9C This is a conceptual diagram used to explain the outline of DMVR processing.
[0184] First, the optimal MVP set for the processing object block is used as the candidate MV. According to the above candidate MV, reference pixels are obtained from the first reference picture which is the processed picture in the L0 direction and the second reference picture which is the processed picture in the L1 direction, and a template is generated by taking the average of each reference pixel.
[0185] Next, the template is used to search the surrounding areas of the candidate MVs of the first reference image and the second reference image, and the MV with the minimum cost is determined as the final MV. In addition, the cost value is calculated using the difference between the pixel values of the template and the pixel values of the search area and the MV value.
[0186] Note that the outline of the processing described here is basically common to the encoding device and the decoding device.
[0187] Furthermore, even if it is not the process described here, other processes may be used as long as they can search the vicinity of the candidate MVs and derive the final MV.
[0188] Here, a mode for generating a predicted image using the LIC process is described.
[0189] Fig.9D This is a diagram for explaining an overview of a method for generating a predicted image using a brightness correction process based on an LIC process.
[0190] First, an MV for acquiring a reference image corresponding to a coding target block from a reference picture that is an already coded picture is derived.
[0191] Next, for the encoding target block, the brightness pixel values of the encoded surrounding reference areas adjacent to the left and above and the brightness pixel values at the same position in the reference picture specified by MV are used to extract information indicating how the brightness values in the reference picture and the encoding target picture change, and calculate the brightness correction parameters.
[0192] By performing brightness correction processing on the reference image in the reference picture specified by the MV using the brightness correction parameters, a predicted image for the encoding target block is generated.
[0193] in addition, Fig.9D The shape of the peripheral reference area described above is just an example, and shapes other than this may also be used.
[0194] In addition, the process of generating a predicted image based on one reference picture is described here, but the same is true when generating a predicted image based on multiple reference pictures. The predicted image is generated after brightness correction processing is performed on the reference images obtained from the respective reference pictures in the same way.
[0195] As a method for determining whether to use the LIC process, for example, there is a method using lic_flag, which is a signal indicating whether the LIC process is used. As a specific example, in the encoding device, it is determined whether the encoding target block belongs to an area where the brightness change occurs. If it belongs to an area where the brightness change occurs, the value 1 is set as lic_flag, and the encoding is performed using the LIC process. If it does not belong to an area where the brightness change occurs, the value 0 is set as lic_flag, and the encoding is performed without using the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, whether to use the LIC process is switched according to the value thereof, and decoding is performed.
[0196] As another method for determining whether to use the LIC process, there is a method of determining whether the LIC process is used in the surrounding blocks. As a specific example, when the encoding target block is in the merge mode, it is determined whether the surrounding coded blocks selected when deriving the MV in the merge mode process are coded using the LIC process, and based on the result, it is switched whether to use the LIC process for coding. In addition, in the case of this example, the processing in decoding is completely the same.
[0197] [Overview of Decoding Device]
[0198] Next, an overview of a decoding device that can decode the coded signal (coded bit stream) output from the coding device 100 will be described. Fig.102 is a block diagram showing a functional structure of a decoding device 200 according to Embodiment 1. The decoding device 200 is a moving picture / image decoding device that decodes a moving picture / image in units of blocks.
[0199] like Fig.10 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218 and a prediction control unit 220.
[0200] The decoding device 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when the processor executes a software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. In addition, the decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0201] Hereinafter, each component included in the decoding device 200 will be described.
[0202] [Entropy decoding unit]
[0203] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. Next, the entropy decoding unit 202 debinarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks.
[0204] [Inverse Quantization Unit]
[0205] The inverse quantization unit 204 inversely quantizes the quantized coefficients of the decoding target block (hereinafter referred to as the current block) as the input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. In addition, the inverse quantization unit 204 outputs the inversely quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0206] [Inverse transformation unit]
[0207] The inverse transform unit 206 restores the prediction error by performing inverse transform on the transform coefficients which are input from the inverse quantization unit 204 .
[0208] For example, when the information read from the encoded bit stream indicates that EMT or AMT is adopted (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the read information indicating the transform type.
[0209] Furthermore, for example, when the information read out from the encoded bit stream indicates that NSST is adopted, the inverse transform unit 206 applies an inverse re-transform to the transform coefficients.
[0210] [Addition Department]
[0211] The adding unit 208 reconstructs the current block by adding the prediction error input from the inverse transform unit 206 and the prediction sample input from the prediction control unit 220 . The adding unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212 .
[0212] [Block Memory]
[0213] The block memory 210 is a storage unit for storing blocks in a decoding target picture (hereinafter referred to as a current picture) to be referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the adding unit 208 .
[0214] [Loop filter section]
[0215] The loop filter unit 212 performs loop filtering on the block reconstructed by the adder unit 208 , and outputs the reconstructed block after filtering to the frame memory 214 , a display device, and the like.
[0216] When the information indicating the on / off of ALF read from the coded bit stream indicates that ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0217] [Frame Memory]
[0218] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212 .
[0219] [Intra-frame prediction unit]
[0220] The intra prediction unit 216 performs intra prediction based on the intra prediction mode read from the coded bit stream, and refers to the block in the current picture stored in the block memory 210, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 performs intra prediction by referring to samples (e.g., luminance values, color difference values) of blocks adjacent to the current block, thereby generating an intra prediction signal, and outputs the intra prediction signal to the prediction control unit 220.
[0221] Furthermore, when an intra prediction mode that refers to a luminance block is selected in intra prediction of a chrominance block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0222] Furthermore, when the information read from the coded bit stream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal and vertical directions.
[0223] [Inter-frame prediction unit]
[0224] The inter-frame prediction unit 218 predicts the current block with reference to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or a sub-block (e.g., a 4×4 block) in the current block. For example, the inter-frame prediction unit 218 performs motion compensation using motion information (e.g., a motion vector) read from the coded bit stream, thereby generating an inter-frame prediction signal of the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0225] When the information decoded from the encoded bit stream indicates that the OBMC mode is adopted, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of the neighboring blocks.
[0226] Furthermore, when the information read from the coded bit stream indicates that the FRUC mode is adopted, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the coded stream, thereby deriving motion information. Furthermore, the inter-frame prediction unit 218 performs motion compensation using the derived motion information.
[0227] In addition, when the BIO mode is adopted, the inter-frame prediction unit 218 derives the motion vector based on the model assuming uniform linear motion. In addition, when the information read from the encoded bit stream indicates that the affine motion compensation prediction mode is adopted, the inter-frame prediction unit 218 derives the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks.
[0228] [Prediction Control Department]
[0229] The prediction control unit 220 selects one of the intra prediction signal and the inter prediction signal, and outputs the selected signal to the adding unit 208 as a prediction signal.
[0230] (First embodiment of implementation method 1)
[0231] Next, the first aspect of the first embodiment will be described in detail with reference to the drawings.
[0232] [Internal Structure of the Transformation Unit of the Encoding Device]
[0233] First, refer to Fig.11A The internal structure of the transformation unit 106 of the encoding device 100 according to this embodiment will be described. Fig.11A This is a block diagram showing the internal structure of the transformation unit 106 of the encoding device 100 according to the first mode of embodiment 1.
[0234] like Fig.11A As shown, the transformation unit 106 related to this method includes a transformation mode determination unit 1061, a size determination unit 1062, a first transformation base selection unit 1063, a first transformation unit 1064, a second transformation implementation determination unit 1065, a second transformation base selection unit 1066, and a second transformation unit 1067.
[0235] The transform mode determination unit 1061 determines whether the adaptive transform base selection mode is effective in the encoding target block. The adaptive transform base selection mode refers to a mode for adaptively selecting a transform base from one or more first transform base candidates. The determination of whether the adaptive transform base selection mode is effective is performed based on, for example, identification information of the first transform base or the adaptive transform base selection mode.
[0236] The size determination unit 1062 determines whether the horizontal size of the encoding target block exceeds the first horizontal threshold size. In addition, the size determination unit 1062 determines whether the vertical size of the encoding target block exceeds the first vertical threshold size. The first horizontal threshold size may be the same as or different from the first vertical threshold size. The first horizontal threshold size and the first vertical threshold size may be predefined in a standard specification, for example. In addition, for example, the first horizontal threshold size and the first vertical threshold size may be sizes determined based on an image, or may be encoded in a bitstream.
[0237] The first transformation base selection unit 1063 selects a first transformation base. In the present invention, selecting a base includes not only selecting at least one base from a plurality of base candidates, but also determining or setting at least one base without a plurality of base candidates.
[0238] When the adaptive transform base selection mode is not valid, the first transform base selection unit 1063 selects one basic transform base as the first transform base in the horizontal direction and the vertical direction. In addition, when the adaptive transform base selection mode is valid, the first transform base selection unit 1063 selects the first transform base in the horizontal direction and the vertical direction as follows (1) to (4) based on the horizontal size and the vertical size of the encoding target block.
[0239] (1) When the horizontal size of the encoding target block is larger than the first horizontal threshold size, the first transform base selection unit 1063 adaptively selects a first transform base in the horizontal direction from one or more transform base candidates.
[0240] (2) When the horizontal size of the encoding target block is equal to or smaller than the first horizontal threshold size, the first transform base selection unit 1063 selects a fixed transform base in the horizontal direction as the first transform base in the horizontal direction.
[0241] (3) When the vertical size of the encoding target block is larger than the first vertical threshold size, the first transform base selection unit 1063 adaptively selects a first transform base in the vertical direction from one or more transform base candidates.
[0242] (4) When the vertical size of the encoding target block is equal to or smaller than the first vertical threshold size, the first transform base selection unit 1063 selects a fixed transform base in the vertical direction as the first transform base in the vertical direction.
[0243] The fixed transformation basis in the horizontal direction may be the same as or different from the fixed transformation basis in the vertical direction. As the fixed transformation basis in the horizontal direction and the vertical direction, for example, the transformation basis of discrete sine transform of type 7 (DST-VII) can be used.
[0244] The first transform unit 1064 performs a first transform on the residual of the encoding target block using the first transform base selected by the first transform base selection unit 1063, thereby generating a first transform coefficient. Specifically, the first transform unit 1064 performs a first transform in the horizontal direction using the first transform base in the horizontal direction, and performs a first transform in the vertical direction using the first transform base in the vertical direction.
[0245] The second transform implementation determination unit 1065 determines whether to implement the second transform for further transforming the first transform coefficients based on whether the adaptive transform base selection mode is valid in the encoding target block. Specifically, the second transform implementation determination unit 1065 implements the second transform when the adaptive transform base selection mode is not valid, and determines not to implement the second transform when the adaptive transform base selection mode is valid.
[0246] When determining that the second transform is to be performed, the second transform base selection unit 1066 selects the second transform base. That is, when the adaptive transform base selection mode is not valid, the second transform base selection unit 1066 selects the second transform base. On the contrary, when the adaptive transform base selection mode is valid, the second transform base selection unit 1066 does not select the second transform base. That is, when the adaptive transform base selection mode is valid, the second transform base selection unit 1066 skips the selection of the second transform base.
[0247] When it is determined that the second transform is to be performed, the second transform unit 1067 transforms the first transform coefficient using the second transform base selected by the second transform base selection unit 1066. That is, when the adaptive transform base selection mode is not valid, the second transform unit 1067 generates the second transform coefficient by performing the second transform on the first transform coefficient using the second transform base. On the contrary, when the adaptive transform base selection mode is valid, the second transform unit 1067 does not perform the second transform on the first transform coefficient. That is, when the adaptive transform base selection mode is valid, the second transform unit 1067 skips the second transform.
[0248] [Internal Structure of the Inverse Transformation Unit of the Encoding Device]
[0249] Next, refer to Fig. 11B The internal structure of the inverse transform unit 114 of the encoding device 100 according to this embodiment will be described. Fig. 11B This is a block diagram showing the internal structure of the inverse transform unit 114 of the encoding device 100 according to the first mode of embodiment 1.
[0250] like Fig. 11B As shown, the inverse transformation unit 114 according to the present embodiment includes a second inverse transformation base selection unit 1141 , a second inverse transformation unit 1142 , a first inverse transformation base selection unit 1143 , and a first inverse transformation unit 1144 .
[0251] When the adaptive transform base selection mode is not valid in the encoding target block, the second inverse transform base selection unit 1141 selects the inverse transform base of the second transform base selected by the second transform base selection unit 1066 as the second inverse transform base.
[0252] When the adaptive transform base selection mode is not effective in the encoding target block, the second inverse transform unit 1142 generates second inverse transform coefficients by performing a second inverse transform on the inverse quantized coefficients using the second inverse transform base selected by the second inverse transform base selection unit 1141. The inverse quantized coefficients are coefficients inversely quantized by the inverse quantization unit 112.
[0253] The first inverse transformation base selection unit 1143 selects an inverse transformation base of the first transformation base selected by the first transformation base selection unit 1063 as the first inverse transformation base.
[0254] When the adaptive transform base selection mode is not valid for the encoding target block, the first inverse transform unit 1144 reconstructs the residual of the encoding target block by performing the first inverse transform on the second inverse transform coefficient using the first inverse transform base. On the other hand, when the adaptive transform base selection mode is valid for the encoding target block, the residual of the encoding target block is reconstructed by performing the first inverse transform on the inverse quantized coefficient using the first inverse transform base.
[0255] [Processing of the Transformation Unit and Quantization Unit of the Coding Device]
[0256] Next, together with the processing of the quantization unit 108, refer to Fig. 12A The processing of the conversion unit 106 configured as above will be described. Fig. 12A This is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to the first mode of embodiment 1.
[0257] The transform mode determination unit 1061 determines whether the adaptive transform base selection mode is valid in the encoding target block ( S101 ).
[0258] When the adaptive transform base selection mode is not valid (No in S101 ), the first transform base selection unit 1063 selects one basic transform base as the first transform base in the horizontal direction and the vertical direction ( S102 ).
[0259] When the adaptive transform base selection mode is valid (Yes in S101), the size determination unit 1062 determines whether the transform size in the horizontal direction exceeds a certain range (S103). That is, the size determination unit 1062 determines whether the horizontal size of the encoding target block is larger than the first horizontal threshold size.
[0260] When the transform size in the horizontal direction exceeds a certain range (Yes in S103 ), the first transform base selection unit 1063 selects a transform base in the horizontal direction from a plurality of adaptive transform bases as a first transform base in the horizontal direction ( S104 ).
[0261] When the transformation size in the horizontal direction is within a certain range (No in S103 ), the first transformation base selection unit 1063 selects a fixed transformation base as the first transformation base in the horizontal direction ( S105 ).
[0262] Next, the size determination unit 1062 determines whether the transformation size in the vertical direction exceeds a certain range (S106). In other words, the size determination unit 1062 determines whether the vertical size of the encoding target block is larger than a first vertical threshold size.
[0263] When the transform size in the vertical direction exceeds a certain range (Yes in S106 ), the first transform base selection unit 1063 selects a transform base in the vertical direction from a plurality of adaptive transform bases as a first transform base in the vertical direction ( S107 ).
[0264] When the transformation size in the vertical direction is within a certain range (No in S106 ), the first transformation base selection unit 1063 selects a fixed transformation base as the first transformation base in the vertical direction ( S108 ).
[0265] In addition, the order of selecting the transformation bases in the horizontal direction and the vertical direction may be the order of the horizontal direction and the vertical direction, or the reverse order. In addition, the transformation base in the horizontal direction and the transformation base in the vertical direction may be selected at the same time.
[0266] The first transform unit 1064 performs a first transform on the prediction residual using the first transform base selected in step S102, S107 or step S108, and generates a first transform coefficient (S109).
[0267] Next, the second transform execution determination unit 1065 determines whether to execute the second transform on the first transform coefficients (S110). Here, the second transform execution determination unit 1065 determines whether to execute the second transform based on whether the adaptive transform base selection mode is valid in the encoding target block.
[0268] When the adaptive transform base selection mode is valid (Yes in S110), neither the selection of the second transform base nor the second transform is performed, and the quantization unit 108 generates quantized coefficients by quantizing the first transform coefficients (S113). Fig. 12A Step S111 and step S112.
[0269] When the adaptive transform base selection mode is not valid (No in S110), the second transform base selection unit 1066 selects a second transform base from one or more second transform base candidates (S111). Then, the second transform unit 1067 performs a second transform on the first transform coefficient using the selected second transform base to generate a second transform coefficient (S112). Then, the quantization unit 108 performs quantization on the second transform coefficient to generate a quantized coefficient (S113).
[0270] As the basic transformation basis, a predetermined transformation basis can be used. In this case, whether the adaptive transformation basis selection mode is effective can be determined based on whether the first transformation basis in the horizontal direction and the first transformation basis in the vertical direction is a predetermined transformation basis. In addition, the predetermined transformation basis can be one transformation basis or two or more transformation bases.
[0271] In addition, when the second transformation is not implemented (skipped), the second transformation may not be implemented, or a transformation equivalent to not implementing the transformation may be implemented as the second transformation. In the former, information indicating that the second transformation is not implemented may also be encoded in the bit stream. In the latter, information indicating that the transformation equivalent to not implementing the transformation may also be encoded in the bit stream. The same can be said about the process of skipping each transformation below.
[0272] also, Fig. 12A The steps and the order of the steps shown are examples and are not limited thereto. Fig. 12B As shown, it is also possible to merge Fig. 12A The determination of the adaptive transformation base selection mode (S101) and the implementation determination of the second transformation (S110). Fig. 12B This is a flowchart showing a modification of the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to the first aspect of the first embodiment. Fig. 12B The flowchart is with Fig. 12A The flowchart of is substantially equivalent to the flowchart of .
[0273] exist Fig. 12B In this case, the transformation unit 106 of the encoding device 100 may not include the second transformation implementation determination unit 1065.
[0274] Regarding the selection of the second inverse transformation base and the second inverse transformation in the inverse transformation unit 114, and the selection of the first inverse transformation base and the first inverse transformation, as long as the Fig. 12A It can be implemented by the transformation of the transformation unit 106, so the description and illustration are omitted.
[0275] In addition, the first transform may be a frequency transform capable of adaptively selecting a transform base, such as the EMT described in non-patent document 2, or a frequency transform that switches the transform base under certain conditions, or other general transforms. For example, a fixed transform base may be set instead of selecting the first transform base. In addition, a first transform base equivalent to not implementing the first transform may be used. In addition, in the first transform, identification information indicating which of the adaptive transform base selection mode and the transform base fixed mode using a fixed basic transform base (for example, the transform base of the discrete cosine transform (DCT-II) of type 2) is effective may be used to select any of the two modes. In this case, it is also possible to determine which of the adaptive transform base selection mode and the transform base fixed mode is effective in the encoding object block based on the identification information. For example, in the EMT described in non-patent document 2, since there is identification information (emt_cu_flag) indicating whether the adaptive transform base selection mode is effective in units of CU (Coding Unit) or the like, it is possible to use this identification information to determine whether the adaptive transform base selection mode is effective in the encoding object block.
[0276] In addition, the second transform may be a secondary transform process such as the NSST described in non-patent document 2, or a transform in which the transform base is switched under certain conditions, or other general transforms. For example, instead of selecting the second transform base, a fixed transform base may be set. In addition, a second transform base equivalent to not implementing the second transform may also be used. In addition, NSST may be a frequency-space transform after DCT or DST, for example, KLT (Karhunen Loveve Transform) or a basis equivalent to KLT that represents the transform coefficients of the DCT or DST obtained offline, or a HyGT (Hypercube-Givens Transform) represented by a combination of rotation transforms.
[0277] In addition, the present processing can also be applied to any one of the luminance signal and the color difference signal, and can also be applied to each signal of R, G, and B as long as the input signal is in the RGB format. Furthermore, the bases selectable in the first conversion or the second conversion may be different for the luminance signal and the color difference signal. For example, since the frequency band of the luminance signal is wider than that of the color difference signal, in order to perform the best conversion, in the first conversion or the second conversion of the luminance signal, more types of bases than those of the color difference can be used as selection candidates. In addition, the present processing can be applied to any one of the intra-frame processing and the inter-frame processing.
[0278] [Effects, etc.]
[0279] In the first transform (first transform) and the second transform (second transform) described in non-patent document 2, the best transform base or transform coefficient (filter) is selected to achieve the best overall coding efficiency. Therefore, in order to search for the best combination of candidates for the transform base and transform coefficient (filter) used in the first transform and the second transform, it is necessary to try the first transform and the second transform multiple times. That is, in the transform method described in non-patent document 2, it is necessary to calculate evaluation values for all combinations of candidates for the transform base of the first transform and candidates for the transform base of the second transform, and select the combination with the smallest evaluation value. Therefore, the inventors have found that in the transform method described in non-patent document 2, the amount of processing becomes huge.
[0280] Therefore, the encoding device 100 according to the present embodiment does not always perform both the first transform and the second transform, but skips the second transform based on whether the adaptive transform base selection mode is effective. As a result, the encoding device 100 can reduce the number of combinations of candidates for the transform base of the first transform and candidates for the transform base of the second transform, and can reduce the amount of processing.
[0281] In addition, according to the encoding device 100 related to the present embodiment, the candidates of the first transform base can be limited based on the conditions of the transform size in the horizontal direction and the vertical direction. As a result, the amount of processing for searching the best first transform base by trial can be reduced. In addition, based on the conditions such as the base selected in the first transform base, the processing for searching the best second transform base by trial can be reduced. In addition, the amount of processing for trial of the combination of the first transform and the second transform can be reduced.
[0282] As an example, as a basic transform basis, a transform basis of DCT-II can be used. DCT-II is highly likely to be adopted when the residual shape is flat or randomly adopted. For example, if DCT-II is used as the first transform basis, there is a possibility that the effect of the second transform is improved because there is a tendency to increase the convergence degree toward low frequencies. On the other hand, in transform bases other than DCT-II, high-frequency components are likely to remain, and there is a possibility that the effect of the second transform is reduced.
[0283] In addition, as an example, the DST-VII transform base can be used as a fixed transform base selected when the transform size is within a certain range. In particular, if it is intra-frame processing, DST-VII has a tendency to be selected with a very high probability when the residual shape is inclined and the size is small.
[0284] In addition, the basic transformation base is not limited to one predetermined transformation base, and multiple predetermined transformation bases may be used.
[0285] In addition, whether to perform the selection of the second transform base and the second transform can be switched according to the transform size. In addition, the candidates of the second transform base can also be switched according to the transform size.
[0286] Alternatively, it may be configured such that whether or not to implement the second transform is switched based only on whether the adaptive transform base selection mode is effective, without switching the first transform base according to the transform size. Fig. 12A In the above method, steps S103, S105, S106 and S108 may be deleted. Here, whether the adaptive transform base selection mode is valid may be determined based on identification information indicating the use of the mode or the type of the first transform base.
[0287] Similarly, it is also possible to configure such that the switching of the second transform execution or not based on whether the adaptive transform base selection mode is effective is not performed, and only the switching of the first transform base according to the transform size is performed. Fig. 12A In the embodiment of the present invention, step S110 may also be deleted.
[0288] Furthermore, the selection of the second transform base and the second transform may not be skipped regardless of whether the adaptive transform base selection mode is effective. Furthermore, regardless of the selection method of the first transform base, the selection of the second transform base and the second transform may be performed when the adaptive transform base selection mode is not effective, and the selection of the second transform base and the second transform may be skipped when the adaptive transform base selection mode is effective.
[0289] In addition, as a threshold of a specific transformation size in the horizontal direction or vertical direction (i.e., the first horizontal threshold size and the first vertical threshold size) for selecting a candidate as the first transformation base from multiple adaptive transformation bases or selecting a fixed transformation base, 4, 8, 16, 32 or 64 pixels, etc. may be used.
[0290] [Combination with other methods]
[0291] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0292] (Second embodiment of implementation method 1)
[0293] Next, a second embodiment of the first embodiment will be described. In this embodiment, an example of encoding various signals related to the first transform and the second transform in the first embodiment will be described. Hereinafter, this embodiment will be specifically described with reference to the drawings, centering on the differences from the first embodiment.
[0294] In addition, since the internal structures of the transformation unit 106 and the inverse transformation unit 114 of the encoding device 100 according to this embodiment are the same as those in the first embodiment, their illustration is omitted.
[0295] [Processing by the Transformation Unit, Quantization Unit, and Entropy Coding Unit of the Coding Device]
[0296] Reference Fig.13A and Fig. 13B , the processing of the transformation unit 106, the quantization unit 108 and the entropy coding unit 110 of the encoding device 100 according to this embodiment is described. Fig.13A This is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to the second aspect of the first embodiment. Fig. 13B 1 is a flowchart showing the processing of the entropy coding unit 110 of the coding device 100 according to the second aspect of the first embodiment. Fig.13A and Fig. 13B In the present invention, the same reference numerals are used for the processes that are common to the first method, and the description thereof is omitted.
[0297] After quantization is performed (S113), the entropy coding unit 110 codes the adaptive transform base selection mode signal (S201). The adaptive transform base selection mode signal is an example of identification information of the adaptive transform base selection mode.
[0298] Then, if the adaptive transform base selection mode is valid ("Yes" in S202), when the transform size in the horizontal direction exceeds a certain range ("Yes" in S203), the entropy coding unit 110 encodes the first base selection signal in the horizontal direction (S204). On the other hand, when the transform size in the horizontal direction is within a certain range ("No" in S203), the entropy coding unit 110 does not encode the first base selection signal in the horizontal direction. Furthermore, when the transform size in the vertical direction exceeds a certain range ("Yes" in S205), the entropy coding unit 110 encodes the first base selection signal in the vertical direction (S206). On the other hand, when the transform size in the vertical direction is within a certain range ("No" in S205), the entropy coding unit 110 does not encode the first base selection signal in the vertical direction.
[0299] When the adaptive transform base selection mode is not valid (No in S202 ), the encoding of the first base selection signal is skipped ( S204 , S206 ).
[0300] Next, the entropy coding unit 110 encodes the quantized coefficients ( S207 ).
[0301] Here, when the adaptive transform base selection mode is not effective (No in S208), the entropy coding unit 110 encodes the second base selection signal (S209). On the other hand, when the adaptive transform base selection mode is effective (Yes in S208), the encoding of the second base selection signal (S209) is skipped.
[0302] Furthermore, the order of each encoding may be preset, and various signals may be encoded in a manner different from the above encoding order.
[0303] When the second transformation is not performed (skipped), a signal indicating that the second transformation is not performed may be encoded, and a signal indicating that the second basis equivalent to not performing the transformation is selected may also be encoded.
[0304] [grammar]
[0305] Here, the syntax in this method is explained. Fig.14 A specific example of the syntax in the second aspect related to the first implementation mode is shown.
[0306] exist Fig.14 In the example, when the adaptive transform base selection mode signal (emt_cu_flag) is set (line 4), if the transform size in the horizontal direction (horizontal_tu_size) is larger than the first horizontal threshold size (horizontal_tu_size_th) (line 5), the first base selection signal in the horizontal direction (emt_horizontal_tridx) is encoded (line 6). In addition, if the transform size in the vertical direction (vertical_tu_size) is larger than the first vertical threshold size (vertical_tu_size_th) (line 11), the first base selection signal in the vertical direction (emt_vertical_tridx) is encoded (line 12). Under other conditions (lines 8 and 14), the encoding of the first base selection signal is skipped (lines 9 and 15).
[0307] In addition, when the adaptive transform base selection mode signal (emt_cu_flag) is not set (line 19), the second base selection signal (secondary_tridx) is encoded (line 20). On the contrary, when the adaptive transform base selection mode signal (emt_cu_flag) is set (line 22), the encoding of the second base selection signal (secondary_tridx) is skipped (line 23).
[0308] [Specific examples of transform bases and coded signals]
[0309] Next, specific examples of transform bases and coded signals are described. Fig.15 A specific example showing the presence or absence of encoding of the transform basis and the signal used in the second mode of implementation mode 1.
[0310] exist Fig.15In the case where the adaptive transform base selection mode is not effective, the transform base of DCT-II is used as the first transform base in the horizontal direction and the vertical direction regardless of the size of the coding block. That is, the transform base of DCT-II is used as the basic transform base. In addition, when the second transform (ON) is performed, the second base selection signal (secondary_tridx) indicating the second transform base used in the second transform is encoded in the bit stream.
[0311] On the other hand, when the adaptive transform base selection mode is valid, according to the horizontal size H and vertical size V of the encoding object block, the combination of DST-VII transform base and other transform bases (index0 to index3) is used as a candidate for the first transform base in the horizontal direction and the vertical direction. In addition, regardless of the size of the encoding object block, the second transform (OFF) is not implemented. In addition, although the second base selection signal (secondary_tridx) is not encoded, the adaptive transform base selection mode signal (emt_cu_flag) is encoded in the bitstream. In addition, when the horizontal size H of the encoding object block is greater than 4 pixels, the first base selection signal in the horizontal direction (emt_horizontal_tridx) is encoded in the bitstream. In addition, when the vertical size V of the encoding object block is greater than 4 pixels, the first base selection signal in the vertical direction (emt_vertical_tridx) is encoded in the bitstream.
[0312] For example, when the horizontal size H is less than 4 pixels and the vertical size V is less than 4 pixels, only the DST-VII transformation basis is used as a candidate for the first transformation basis in the horizontal direction and the vertical direction. At this time, the first basis selection signals in the horizontal direction and the vertical direction (emt_horizontal_tridx and emt_vertical_tridx) are not encoded.
[0313] Furthermore, for example, when the horizontal size H is less than or equal to 4 pixels and the vertical size V is greater than 4 pixels, only the DST-VII transformation basis is used as a candidate for the first transformation basis in the horizontal direction, and the DST-VII transformation basis and other transformation bases are used as candidates for the first transformation basis in the vertical direction. In this case, although the first basis selection signal in the horizontal direction (emt_horizontal_tridx) is not encoded, the first basis selection signal in the vertical direction (emt_vertical_tridx) is encoded.
[0314] Furthermore, for example, when the horizontal size H is greater than 4 pixels and the vertical size V is less than or equal to 4 pixels, the DST-VII transformation basis and other transformation bases are used as candidates for the first transformation basis in the horizontal direction, and only the DST-VII transformation basis is used as a candidate for the first transformation basis in the vertical direction. In this case, the first basis selection signal in the horizontal direction (emt_horizontal_tridx) is encoded, but the first basis selection signal in the vertical direction (emt_vertical_tridx) is not encoded.
[0315] In addition, for example, when the horizontal size H is greater than 4 pixels and the vertical size V is greater than 4 pixels, the DST-VII transformation basis and other transformation bases are used as candidates for the first transformation bases in the horizontal direction and the vertical direction. At this time, the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) in the horizontal direction and the vertical direction are encoded.
[0316] [Effects, etc.]
[0317] As described above, according to the encoding device 100 related to the present embodiment, it is possible to encode information (first base selection signal) indicating the first transform base only when the adaptive transform base selection mode is valid and the transform size exceeds a certain range, and there is a possibility that the amount of code required for signaling of the first transform base can be reduced. In addition, it is possible to encode information (second base selection signal) indicating the second transform base only when the adaptive transform base selection mode is not valid, and there is a possibility that the amount of code required for signaling of the second transform base can be reduced. In addition, by encoding information for determining whether to skip the second transform (adaptive transform base selection mode signal, etc.) before the information indicating the second transform base, it is possible to determine whether the information indicating the second transform base is encoded during decoding.
[0318] In addition, the second base selection signal may be encoded regardless of the adaptive transform base selection mode. In addition, the first base selection signal may be encoded regardless of the transform size if the adaptive transform base selection mode is selected. In addition, the presence or absence of encoding of the first base selection signal may be determined independently using the size in the horizontal direction and the size in the vertical direction, or the determination may be made in combination.
[0319] [Combination with other methods]
[0320] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0321] (Third embodiment of implementation method 1)
[0322] Next, a third embodiment of the first embodiment is described. This embodiment is different from the first embodiment described above in that, when the adaptive transform base selection mode is not effective, a different basic transform base is used as the first transform base according to the size of the encoding target block. Hereinafter, this embodiment will be specifically described with reference to the accompanying drawings, centering on the differences from the first and second embodiments.
[0323] In addition, since the internal structures of the transformation unit 106 and the inverse transformation unit 114 of the encoding device 100 according to this embodiment are the same as those in the first embodiment, their illustration is omitted.
[0324] [Processing of the Transformation Unit and Quantization Unit of the Coding Device]
[0325] Reference Fig.16 The processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to this embodiment will be described. Fig.16 1 is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to the third aspect of the first embodiment. Fig.16 In the present invention, the same reference numerals are used for the processes that are common to the first method, and the description thereof is omitted.
[0326] When the adaptive transform base selection mode is not valid (No in S101), the size determination unit 1062 determines whether the transform size is within a certain range (S301). That is, the size determination unit 1062 determines whether the size of the encoding target block is less than or equal to the second threshold size. For example, the size determination unit 1062 determines whether the product of the horizontal size and the vertical size of the encoding target block is less than or equal to the threshold, thereby determining whether the size of the encoding target block is less than or equal to the second threshold size.
[0327] Here, when the transformation size is within a certain range ("Yes" in S301), the first transformation base selection unit 1063 selects the second basic transformation base as the first transformation base in the horizontal direction and the vertical direction (S302). On the other hand, when the transformation size exceeds the certain range ("No" in S301), the first transformation base selection unit 1063 selects the first basic transformation base as the first transformation base in the horizontal direction and the vertical direction (S303).
[0328] As an example, a DCT-II transform basis can be used as the first basic transform basis, and a DCT-VII transform basis can be used as the second basic transform basis.
[0329] In addition, the basic transformation basis can also be selected from multiple basic transformation basis candidates.
[0330] Furthermore, the selection of the second transform base and the second transform may not be skipped regardless of whether the adaptive transform base selection mode is effective. Furthermore, regardless of the selection method of the first transform base, the selection of the second transform base and the second transform may be performed when the adaptive transform base selection mode is not effective, and the selection of the second transform base and the second transform may be skipped when the adaptive transform base selection mode is effective.
[0331] In addition, when the adaptive transform base selection mode is not effective, as the second threshold size for selecting one of the first basic transform base and the second basic transform base, for example, 4x4, 4x8, 8x4, 8x8 pixel size, etc. can be used. In addition, as the transform size compared with the threshold, the product of the horizontal size and the vertical size of the encoding target block can be used as in the present embodiment, or the horizontal size and the vertical size can be used separately.
[0332] In addition, when the adaptive transformation base selection mode is not valid, if the product of the horizontal size and the vertical size is within a certain range, the second basic transformation base can also be selected as the first transformation base in the horizontal and vertical directions, skipping the selection of the second transformation base and the second transformation.
[0333] [Effects, etc.]
[0334] As described above, according to the encoding device 100 related to the present embodiment, when the adaptive transform base selection mode is not effective, the first transform base can be switched between the first basic transform base and the second basic transform base according to the transform size. Therefore, the first transform can be performed using the first transform base corresponding to the transform size, and the code amount can be reduced.
[0335] [Combination with other methods]
[0336] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0337] (Fourth embodiment of implementation method 1)
[0338] Next, a fourth embodiment of the first embodiment will be described. In this embodiment, an example of encoding various signals related to the first transform and the second transform of the third embodiment will be described. Hereinafter, this embodiment will be specifically described with reference to the drawings, centering on the differences from the first to third embodiments.
[0339] In addition, since the internal structures of the transformation unit 106 and the inverse transformation unit 114 of the encoding device 100 according to this embodiment are the same as those in the first embodiment, their illustration is omitted.
[0340] [Processing by the Transformation Unit, Quantization Unit, and Entropy Coding Unit of the Coding Device]
[0341] Reference Fig.17A and Fig. 17B , the processing of the transformation unit 106, the quantization unit 108 and the entropy coding unit 110 of the encoding device 100 according to this embodiment is described. Fig.17A This is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to the fourth aspect of the first embodiment. Fig. 17B 4 is a flowchart showing the processing of the entropy coding unit 110 of the coding device 100 according to the fourth aspect of the first embodiment. Fig.17A and Fig. 17B In the present invention, the same reference numerals are used for the processes that are common to any of the first to third embodiments, and the description thereof will be omitted.
[0342] After quantization is performed (S113), the entropy coding unit 110 determines whether to skip the coding of the adaptive transform base selection mode signal (S401). For example, when any one of the following conditions (A) and (B) is satisfied, the entropy coding unit 110 determines to skip the coding of the adaptive transform base selection mode signal, otherwise, the entropy coding unit 110 determines not to skip the coding of the adaptive transform base selection mode signal.
[0343] (A) The adaptive transform basis selection mode is not effective.
[0344] (B) The adaptive transform base selection mode is valid and satisfies all of the following conditions (B1) to (B4).
[0345] (B1) The transformation size is equal to or smaller than the second threshold size W1xH1 used in step S301.
[0346] (B2) The transformation size in the horizontal direction is equal to or smaller than the first horizontal threshold size W2 used in step S103.
[0347] (B3) The transformation size in the vertical direction is equal to or smaller than the first vertical threshold size H2 used in step S106.
[0348] (B4) The second basic transformation basis and the fixed transformation basis in the horizontal direction and the vertical direction are the same transformation basis.
[0349] As a specific example, when the second threshold size W1xH1 is 4×4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and the second basic transformation basis and the fixed transformation basis are both DST-VII transformation basis, if the transformation size is less than 4×4 pixels, the entropy coding unit 110 determines to skip the encoding of the adaptive transformation basis selection mode signal.
[0350] On the contrary, when any one of the above conditions (A) and (B) is not satisfied, the entropy coding unit 110 determines not to skip coding of the adaptive transform base selection mode signal.
[0351] Here, when it is determined that the encoding of the adaptive transform base selection mode signal is skipped ("Yes" in S401), the entropy coding unit 110 skips steps S201 to S206 and encodes the quantized coefficients (S207). On the other hand, when it is determined that the encoding of the adaptive transform base selection mode signal is not skipped ("No" in S401), the entropy coding unit 110 performs steps S201 to S206 in the same manner as in the second embodiment and then encodes the quantized coefficients (S207).
[0352] Furthermore, the order of each encoding may be preset, and various signals may be encoded in a manner different from the above encoding order.
[0353] [grammar]
[0354] Here, the syntax in this method is explained. Fig.18 A specific example of the syntax in the fourth aspect related to Implementation Method 1 is shown.
[0355] exist Fig.18 For example, in the case where the encoding of the adaptive transform base selection mode signal is skipped (line 20), the encoding of the adaptive transform base selection mode signal (emt_cu_flag) and the first base selection signal (emt_horizontal_tridx and emt_vertical_tridx) is skipped (line 21). Here, when the transform size in the horizontal direction (horizontal_tu_size) is less than the first horizontal threshold size (horizontal_tu_size_th) and the transform size in the vertical direction (vertical_tu_size) is less than the first vertical threshold size (vertical_tu_size_th), the encoding of the adaptive transform base selection mode signal is skipped. In the case where the encoding of the adaptive transform base selection mode signal is not skipped (lines 3-4), the adaptive transform base selection mode signal (emt_cu_flag) is encoded (line 5), and the first base selection signal (emt_horizontal_tridx and emt_vertical_tridx) is encoded as needed in the same manner as in the second method (lines 7-16).
[0356] Furthermore, when the encoding of the adaptive transform base selection mode signal is skipped, the selection of the second transform base and the second transform may be skipped.
[0357] [Specific examples of transform bases and coded signals]
[0358] Next, specific examples of transform bases and coded signals are described. Fig.19 A specific example of the transform base used in the fourth embodiment of the first embodiment and the presence or absence of coding of the signal is shown. Fig.19 In the case where both the horizontal size and the vertical size of the coding target block are less than 4 pixels, the transform base and the presence or absence of coding are different from Fig.15 Different. Fig.15 Different points as the center Fig.19 Provide explanation.
[0359] exist Fig.19 In the case where the adaptive transform base selection mode is not effective, if the horizontal size H and the vertical size V of the encoding object block are both less than 4 pixels, the DST-VII transform base rather than the DCT-II transform base is used as the first transform base in the horizontal and vertical directions.
[0360] Furthermore, when the adaptive transform base selection mode is valid, if the horizontal size H and the vertical size V of the encoding target block are both smaller than or equal to 4 pixels, the adaptive transform base selection mode signal (emt_cu_flag) is not encoded.
[0361] [Effects, etc.]
[0362] As described above, according to the encoding device 100 related to this method, when the conditions for skipping the encoding of the adaptive transform base selection mode signal are met, all encoding of the adaptive transform base selection mode signal and the first base selection signal can be omitted, and there is a possibility of reducing the amount of code.
[0363] [Combination with other methods]
[0364] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0365] (Fifth embodiment of implementation method 1)
[0366] Next, the fifth embodiment of the first embodiment is described. In this embodiment, a decoding device is described. In addition, the decoding device of this embodiment corresponds to the encoding device of the first embodiment. That is, the decoding device of this embodiment can decode the bit stream encoded by the encoding device of the first embodiment. Hereinafter, this embodiment is specifically described with reference to the accompanying drawings.
[0367] [Internal Configuration of Transformation Unit and Inverse Transformation Unit of Decoding Device]
[0368] First, the internal structure of the inverse transform unit 206 of the decoding device 200 according to this embodiment will be described. Fig. 20 This is a block diagram showing the internal structure of the inverse transform unit 206 of the decoding device 200 according to the fifth aspect of the first embodiment.
[0369] like Fig. 20 As shown, the inverse transformation unit 206 related to this method includes a second inverse transformation implementation determination unit 2061, a second inverse transformation base selection unit 2062, a second inverse transformation unit 2063, a transformation mode determination unit 2064, a size determination unit 2065, a first inverse transformation base selection unit 2066, and a first inverse transformation unit 2067.
[0370] The second inverse transform implementation determination unit 2061 determines whether to implement the second inverse transform on the inverse quantized coefficients of the decoding target block output from the inverse quantization unit 204, based on whether the adaptive transform base selection mode is valid in the decoding target block. Specifically, the second inverse transform implementation determination unit 2061 implements the second inverse transform when the adaptive transform base selection mode is not valid, and determines not to implement the second inverse transform when the adaptive transform base selection mode is valid.
[0371] When it is determined that the second inverse transform is to be performed, the second inverse transform base selection unit 2062 selects the second inverse transform base. Specifically, when the adaptive transform base selection mode is not valid, the second inverse transform base selection unit 2062 obtains the second base selection signal 2062S indicating the second inverse transform base decoded from the bit stream by the entropy decoding unit 202. Then, the second inverse transform base selection unit 2062 selects the second inverse transform base based on the second base selection signal 2062S. On the contrary, when the adaptive transform base selection mode is valid, the second inverse transform base selection unit 2062 does not select the second inverse transform base. That is, when the adaptive transform base selection mode is valid, the second inverse transform base selection unit 2062 skips the selection of the second inverse transform base.
[0372] When it is determined that the second inverse transform is to be performed, the second inverse transform unit 2063 performs the second inverse transform on the inverse quantized coefficients of the decoding target block using the second inverse transform base selected by the second inverse transform base selection unit 2062. That is, when the adaptive transform base selection mode is not effective, the second inverse transform unit 2063 generates the second inverse transform coefficients by performing the second inverse transform on the inverse quantized coefficients using the second inverse transform base. On the contrary, when the adaptive transform base selection mode is effective, the second inverse transform unit 2063 does not perform the second inverse transform on the inverse quantized coefficients. That is, when the adaptive transform base selection mode is effective, the second inverse transform unit 2063 skips the second inverse transform.
[0373] The transform mode determination unit 2064 determines whether the adaptive transform base selection mode is valid in the decoding target block. The determination of whether the adaptive transform base selection mode is valid is performed based on the first base selection signal 2066S or the adaptive transform base selection mode signal 2064S decoded from the bit stream by the entropy decoding unit 202. That is, the determination is performed based on the identification information of the first inverse transform base or the adaptive transform base selection mode.
[0374] The size determination unit 2065 determines whether the horizontal size of the decoding target block exceeds the first horizontal threshold size. In addition, the size determination unit 1062 determines whether the vertical size of the decoding target block exceeds the first vertical threshold size. The determination of the horizontal size and the vertical size is performed based on the size signal 2065S decoded from the bit stream by the entropy decoding unit 202.
[0375] The first inverse transform base selection unit 2066 selects the first inverse transform base. Specifically, when the adaptive transform base selection mode is not valid, the first inverse transform base selection unit 2066 selects one basic transform base as the first inverse transform base in the horizontal direction and the vertical direction. In addition, when the adaptive transform base selection mode is valid, the first inverse transform base selection unit 2066 selects the first inverse transform base in the horizontal direction and the vertical direction as follows (1) to (4) based on the horizontal size and the vertical size of the decoding target block.
[0376] (1) When the horizontal size of the decoding target block is larger than the first horizontal threshold size, the first inverse transform base selection unit 2066 obtains the first base selection signal 2066S indicating the first inverse transform base decoded from the bit stream by the entropy decoding unit 202. Then, the first inverse transform base selection unit 2066 selects the first inverse transform base in the horizontal direction based on the first base selection signal 2066S.
[0377] (2) When the horizontal size of the decoding target block is equal to or smaller than the first horizontal threshold size, the first inverse transform base selection unit 2066 selects a fixed transform base in the horizontal direction as the first inverse transform base in the horizontal direction.
[0378] (3) When the vertical size of the decoding target block is larger than the first vertical threshold size, the first inverse transform base selection unit 2066 obtains the first base selection signal 2066S. Then, the first inverse transform base selection unit 2066 selects the first inverse transform base in the vertical direction based on the first base selection signal 2066S.
[0379] (4) When the vertical size of the decoding target block is equal to or smaller than the first vertical threshold size, the first inverse transform base selection unit 2066 selects a fixed transform base in the vertical direction as the first inverse transform base in the vertical direction.
[0380] The first inverse transform unit 2067 performs a first inverse transform on the inverse quantized coefficients of the decoding target block by using the first inverse transform base selected by the first inverse transform base selection unit 2066, thereby restoring the residual of the decoding target block. Specifically, the first inverse transform unit 2067 performs a first inverse transform in the horizontal direction by using the first inverse transform base in the horizontal direction, and performs a first inverse transform in the vertical direction by using the first inverse transform base in the vertical direction.
[0381] [Processing of the Inverse Quantization Unit and Inverse Transformation Unit of the Decoding Device]
[0382] Next, together with the processing of the inverse quantization unit 204, refer to Fig.21 The processing of the inverse transform unit 206 configured as above will be described. Fig.21 This is a flowchart showing the processing of the inverse quantization unit 204 and the inverse transformation unit 206 of the decoding device 200 according to the fifth aspect of the first embodiment.
[0383] The inverse quantization unit 204 generates inverse quantization coefficients by inverse quantizing the quantized coefficients of the decoding target block decoded by the entropy decoding unit 202 ( S501 ).
[0384] The second inverse transform execution determination unit 2061 determines whether to perform the second inverse transform on the inverse quantized coefficients (S502). Here, the second inverse transform execution determination unit 2061 determines whether to perform the second inverse transform based on whether the adaptive transform base selection mode is valid in the decoding target block.
[0385] Here, when the adaptive transform base selection mode is valid (Yes in S502), neither the selection of the second inverse transform base nor the second inverse transform is performed. That is, step S503 and step S504 are skipped.
[0386] On the other hand, when the adaptive transform base selection mode is not effective (No in S502), the second inverse transform base selection unit 2062 selects the second inverse transform base based on the second base selection signal 2062S (S503). Furthermore, the second inverse transform unit 2063 performs a second inverse transform on the inverse quantized coefficients using the selected second inverse transform base (S504).
[0387] Next, the transform mode determination unit 2064 determines whether the adaptive transform base selection mode is valid in the decoding target block (S505). For example, the transform mode determination unit 2064 determines whether the adaptive transform base selection mode is valid based on the adaptive transform base selection mode signal 2064S.
[0388] When the adaptive transform base selection mode is not valid (No in S505), the first inverse transform base selection unit 2066 selects one basic transform base as the first inverse transform base in the horizontal direction and the vertical direction (S512). On the other hand, when the adaptive transform base selection mode is valid (Yes in S505), the size determination unit 2065 determines whether the transform size in the horizontal direction exceeds a certain range (S506). That is, the size determination unit 2065 determines whether the horizontal size of the decoding target block is larger than the first horizontal threshold size.
[0389] When the transform size in the horizontal direction exceeds a certain range ("Yes" in S506), the first inverse transform base selection unit 2066 selects a transform base in the horizontal direction from a plurality of adaptive transform bases as the first inverse transform base in the horizontal direction (S507). On the other hand, when the transform size in the horizontal direction is within a certain range ("No" in S506), the first inverse transform base selection unit 2066 selects a fixed transform base as the first inverse transform base in the horizontal direction (S508).
[0390] The size determination unit 2065 determines whether the transform size in the vertical direction exceeds a certain range (S509). In other words, the size determination unit 2065 determines whether the vertical size of the decoding target block is larger than a first vertical threshold size.
[0391] When the transform size in the vertical direction exceeds a certain range ("Yes" in S509), the first inverse transform base selection unit 2066 selects a transform base from a plurality of adaptive transform bases as the first inverse transform base in the vertical direction (S510). When the transform size in the vertical direction is within a certain range ("No" in S509), the first inverse transform base selection unit 2066 selects a fixed transform base as the first inverse transform base in the vertical direction (S511).
[0392] The first inverse transform unit 2067 performs the first inverse transform on the inverse quantized coefficients or the second inverse transform coefficients using the first inverse transform base selected as described above, thereby restoring the residual of the decoding target block ( S513 ).
[0393] In addition, the order of selecting the inverse transformation bases in the horizontal direction and the vertical direction may be the order of the horizontal direction and the vertical direction, or the reverse order thereof. In addition, the inverse transformation base in the horizontal direction and the inverse transformation base in the vertical direction may be selected at the same time.
[0394] In addition, selecting an inverse transform basis in the decoding device 200 means decoding information representing a basis used in the inverse transform contained in the encoded bit stream, determining the inverse transform basis based on the decoded information, or determining a uniquely represented inverse transform basis based on information such as an intra-frame prediction mode, a decoding object block size, or a basis in the first inverse transform.
[0395] Alternatively, you can also use Fig. 12A or Fig. 12B A decoding method of the encoding method of the first embodiment shown.
[0396] [Effects, etc.]
[0397] As described above, according to the decoding device 200 according to this embodiment, it is possible to achieve the same effects as those of the encoding device 100 according to the first embodiment.
[0398] [Combination with other methods]
[0399] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0400] (Sixth embodiment of implementation method 1)
[0401] Next, the sixth embodiment of the first embodiment is described. In this embodiment, an example of decoding various signals related to the first transform and the second transform in the fifth embodiment is described. In addition, the decoding device related to this embodiment corresponds to the encoding device of the second embodiment described above. Hereinafter, this embodiment will be specifically described with reference to the accompanying drawings, centering on the points different from the fifth embodiment.
[0402] In addition, the internal structure of the inverse transform unit 206 of the decoding device 200 related to this embodiment is the same as that of the fifth embodiment, and therefore is omitted from the illustration.
[0403] [Processing by the Entropy Decoding Unit, Inverse Quantization Unit, and Inverse Transformation Unit of the Decoding Device]
[0404] Reference Fig.22A and Fig. 22B , the processing of the entropy decoding unit 202, the inverse quantization unit 204, and the inverse transformation unit 206 of the decoding device 200 of this embodiment will be described. Fig.22A and Fig. 22B In the embodiment, the same reference numerals are used for the processes that are common to the fifth embodiment and the description thereof is omitted.
[0405] First, the entropy decoding unit 202 decodes the adaptive transform base selection mode signal from the bit stream (S601). Then, the transform mode determination unit 2064 determines whether the adaptive transform base selection mode is valid in the decoding target block based on the adaptive transform base selection mode signal (S602).
[0406] If the adaptive transform base selection mode is valid ("Yes" in S602), when the transform size in the horizontal direction exceeds a certain range ("Yes" in S603), the entropy decoding unit 202 decodes the first base selection signal in the horizontal direction from the bit stream (S604). On the other hand, when the transform size in the horizontal direction is within a certain range ("No" in S603), the entropy decoding unit 202 does not decode the first base selection signal in the horizontal direction. Furthermore, when the transform size in the vertical direction exceeds a certain range ("Yes" in S605), the entropy decoding unit 202 decodes the first base selection signal in the vertical direction from the bit stream (S606). On the other hand, when the transform size in the vertical direction is within a certain range ("No" in S605), the entropy decoding unit 202 does not decode the first base selection signal in the vertical direction.
[0407] When the adaptive transform base selection mode is not valid (No in S602 ), decoding of the first base selection signal is skipped ( S604 , S606 ).
[0408] Next, the entropy decoding unit 202 decodes the quantized coefficients ( S607 ).
[0409] Here, when the adaptive transform base selection mode is not valid (No in S608), the entropy decoding unit 202 decodes the second base selection signal from the bit stream (S609). On the other hand, when the adaptive transform base selection mode is valid (Yes in S608), the decoding of the second base selection signal is skipped (S609).
[0410] In addition, the order of each decoding may be pre-set to match the encoding method, and various signals may be decoded in a manner different from the order of the above decoding. In addition, when the second inverse transform is not implemented (skipped), the entropy decoding unit 202 may decode from the bit stream a signal indicating that the second inverse transform is not implemented, or may decode from the bit stream a signal for selecting a second inverse transform basis equivalent to no transform.
[0411] Alternatively, you can also use Fig.13A , Fig. 13B as well as Fig.14 A decoding method for the encoding method of the second embodiment shown.
[0412] [Effects, etc.]
[0413] As described above, according to the decoding device 200 according to this embodiment, it is possible to achieve the same effects as those of the encoding device 100 according to the second embodiment.
[0414] [Combination with other methods]
[0415] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0416] (Seventh embodiment of implementation method 1)
[0417] Next, the seventh embodiment of the first embodiment is described. This embodiment is different from the fifth embodiment described above in that, when the adaptive transform base selection mode is not effective, a different basic transform base is used as the first inverse transform base according to the size of the encoding target block. In addition, the decoding device related to this embodiment corresponds to the encoding device of the third embodiment described above. Hereinafter, this embodiment will be specifically described with reference to the accompanying drawings, centering on the points different from the fifth and sixth embodiments.
[0418] In addition, the internal structure of the inverse transform unit 206 of the decoding device 200 related to this embodiment is the same as that of the fifth embodiment, and therefore is omitted from the illustration.
[0419] [Processing of the Inverse Quantization Unit and Inverse Transformation Unit of the Decoding Device]
[0420] Reference Fig.23 The processing of the transform unit 106 and the quantization unit 108 of the encoding device 100 according to this embodiment will be described. Fig.23 1 is a flowchart showing the processing of the inverse quantization unit 204 and the inverse transformation unit 206 of the decoding device 200 according to the seventh aspect of the first embodiment. Fig.23 In the embodiment, the same reference numerals are used for the processes that are common to the fifth embodiment and the description thereof is omitted.
[0421] When the adaptive transform base selection mode is not valid (No in S505), the size determination unit 2065 determines whether the transform size is within a certain range (S701). That is, the size determination unit 2065 determines whether the horizontal size and the vertical size of the decoding target block are less than or equal to the second threshold size. Specifically, the size determination unit 2065 determines, for example, whether the product of the horizontal size and the vertical size of the decoding target block is less than or equal to the threshold.
[0422] Here, when the transformation size is within a certain range ("Yes" in S701), the first inverse transformation base selection unit 2066 selects the second basic transformation base as the first inverse transformation base in the horizontal direction and the vertical direction (S702). On the other hand, when the transformation size exceeds the certain range ("No" in S701), the first inverse transformation base selection unit 2066 selects the first basic transformation base as the first inverse transformation base in the horizontal direction and the vertical direction (S703).
[0423] Alternatively, you can also use Fig.16A decoding method for the encoding method according to the third embodiment is shown.
[0424] [Effects, etc.]
[0425] As described above, according to the decoding device 200 according to this embodiment, it is possible to achieve the same effects as those of the encoding device 100 according to the third embodiment.
[0426] [Combination with other methods]
[0427] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0428] (Eighth embodiment of implementation method 1)
[0429] Next, the eighth embodiment of the first embodiment is described. In this embodiment, an example of decoding various signals related to the first transform and the second transform in the seventh embodiment is described. In addition, the decoding device related to this embodiment corresponds to the encoding device of the fourth embodiment described above. Hereinafter, this embodiment will be specifically described with reference to the accompanying drawings, centering on the points different from the fifth to seventh embodiments.
[0430] In addition, the internal structure of the inverse transform unit 206 of the decoding device 200 related to this embodiment is the same as that of the fifth embodiment, and therefore is omitted from the illustration.
[0431] [Processing by the Entropy Decoding Unit, Inverse Quantization Unit, and Inverse Transformation Unit of the Decoding Device]
[0432] Reference Fig.24A and Fig. 24B The following describes the processing of the entropy decoding unit 202, the inverse quantization unit 204, and the inverse transformation unit 206 of the decoding device 200 according to this embodiment. Fig.24A and Fig. 24B In the present invention, the same reference numerals are used for the processing that is common to any one of the fifth to seventh methods, and the description thereof is omitted.
[0433] The entropy decoding unit 202 determines whether to skip decoding of the adaptive transform base selection mode signal (S801). For example, when any one of the following conditions (A) and (B) is satisfied, the entropy decoding unit 202 determines to skip decoding of the adaptive transform base selection mode signal, otherwise, the entropy decoding unit 202 determines not to skip decoding of the adaptive transform base selection mode signal.
[0434] (A) The adaptive transform basis selection mode is not effective.
[0435] (B) The adaptive transform base selection mode is valid and satisfies all of the following conditions (B1) to (B4).
[0436] (B1) The transformation size is equal to or smaller than the second threshold size W1xH1 used in step S701.
[0437] (B2) The transformation size in the horizontal direction is equal to or smaller than the first horizontal threshold size W2 used in step S506.
[0438] (B3) The transformation size in the vertical direction is equal to or smaller than the first vertical threshold size H2 used in step S509.
[0439] (B4) The second basic transformation basis and the fixed transformation basis in the horizontal direction and the vertical direction are the same transformation basis.
[0440] As a specific example, when the second threshold size W1xH1 is 4×4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and the second basic transformation basis and the fixed transformation basis are both DST-VII transformation basis, if the transformation size is less than 4×4 pixels, the entropy decoding unit 202 determines to skip the decoding of the adaptive transformation basis selection mode signal.
[0441] On the contrary, when any one of the above conditions (A) and (B) is not satisfied, the entropy decoding unit 202 determines not to skip decoding of the adaptive transform base selection mode signal.
[0442] Here, when it is determined that the decoding of the adaptive transform base selection mode signal is skipped ("Yes" in S801), the entropy decoding unit 202 skips steps S601 to S606 and decodes the quantized coefficients (S607). On the other hand, when it is determined that the decoding of the adaptive transform base selection mode signal is not skipped ("No" in S801), the entropy decoding unit 202 performs steps S601 to S606 in the same manner as in the sixth embodiment and then decodes the quantized coefficients (S207).
[0443] Alternatively, you can also use Fig.17A , Fig. 17B as well as Fig.18 A decoding method for the encoding method of the fourth embodiment shown.
[0444] [Effects, etc.]
[0445] As described above, according to the decoding device 200 related to this embodiment, it is possible to achieve the same effects as those of the encoding device 100 related to the fourth embodiment.
[0446] [Combination with other methods]
[0447] This embodiment may be implemented in combination with at least a portion of other embodiments of the present invention. In addition, a portion of the processing described in the flowchart of this embodiment, a portion of the device structure, a portion of the syntax, etc. may be implemented in combination with other embodiments.
[0448] (Variations of each aspect of Embodiment 1)
[0449] In addition, a signal indicating whether part or all of the processing described in any one of the first to eighth methods is valid may be encoded and decoded. Such a signal may be encoded in units of CU (Coding Unit) or CTU (Coding Tree Unit), or may be encoded in units of SPS (Sequence Parameter Set), PPS (Picture Parameter Set) or slices equivalent to the H.265 / HEVC standard.
[0450] Based on the picture type (I, P, B), slice type (I, P, B), transform size (4×4 pixels, 8x8 pixels or other), number of non-zero coefficients, quantization parameter, Temporal_id (layer of hierarchical coding) or any combination thereof, the selection of the first transform base and the first transform can be skipped, and the selection of the second transform base and the second transform can also be skipped.
[0451] When the encoding device according to the first to fourth aspects performs the above-mentioned operations, the decoding device according to the fifth to eighth aspects also performs corresponding operations. For example, when the encoding device encodes information indicating whether to enable the process of skipping the first transform or the second transform, the decoding device decodes the information and determines whether the first transform or the second transform is enabled and whether the information indicating the first transform or the second transform is encoded.
[0452] (Implementation Method 2)
[0453] In each of the above embodiments, each functional block can usually be implemented by an MPU and a memory. In addition, the processing of each functional block is usually implemented by a program execution unit such as a processor reading and executing the software (program) recorded in a recording medium such as a ROM. The software can be distributed by downloading, etc., or recorded in a recording medium such as a semiconductor memory for distribution. In addition, of course, each functional block can also be implemented by hardware (dedicated circuit).
[0454] In addition, the processing described in each embodiment can be implemented by centralized processing using a single device (system), or can be implemented by decentralized processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing can be performed or decentralized processing can be performed.
[0455] The aspects of the present invention are not limited to the above-described embodiments, and various modifications can be made, which are also included in the scope of the aspects of the present invention.
[0456] Furthermore, an application example of the moving picture encoding method (image encoding method) or the moving picture decoding method (image decoding method) shown in each of the above embodiments and a system using the same are described here. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding and decoding device having both. Other structures in the system can be appropriately changed according to the situation.
[0457] [Example of use]
[0458] Fig.25 1 is a diagram showing the overall configuration of a content providing system ex100 for realizing content distribution services. A communication service providing area is divided into cells of desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.
[0459] In the content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may combine some of the above elements. Instead of the base stations ex106 to ex110, which are fixed wireless stations, the devices may be directly or indirectly connected to each other via a telephone network or short-range wireless. In addition, the streaming media server ex103 is connected to the devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101. In addition, the streaming media server ex103 is connected to a terminal in a hot spot in an airplane ex117 via a satellite ex116.
[0460] Alternatively, wireless access points or hotspots may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or directly connected to the aircraft ex117 without going through the satellite ex116.
[0461] The camera ex113 is a device such as a digital camera capable of taking still images and moving images. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that supports mobile communication systems generally referred to as 2G, 3G, 3.9G, 4G, and 5G in the future.
[0462] The household appliance ex118 is a refrigerator or a device included in a household fuel cell cogeneration system.
[0463] In the content supply system ex100, a terminal having a camera function is connected to the streaming server ex103 via the base station ex106, etc., so that on-site distribution can be performed. In on-site distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) performs the encoding process described in the above embodiments on still images or moving image contents captured by a user using the terminal, multiplexes the image data obtained by the encoding and the sound data obtained by encoding the sound corresponding to the image, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.
[0464] On the other hand, the streaming server ex103 streams the content data sent by the client that requested it. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117 that can decode the coded data. Each device that receives the distributed data decodes the received data and reproduces it. That is, each device functions as an image decoding device according to one aspect of the present invention.
[0465] [Distributed processing]
[0466] In addition, the streaming media server ex103 can also be a plurality of servers or a plurality of computers, which distributes data by distributing processing or recording. For example, the streaming media server ex103 can also be implemented by CDN (Contents Delivery Network), which implements content distribution by connecting many edge servers scattered in the world to the network between the edge servers. In CDN, physically closer edge servers are dynamically allocated according to the client. And, by caching and distributing content to the edge server, delay can be reduced. In addition, in the case of a certain error or the communication state changes due to an increase in the amount of communication, it is possible to distribute the processing with multiple edge servers, or switch the distribution subject to other edge servers, or bypass the part of the network where the failure occurs and continue to distribute, so that high-speed and stable distribution can be achieved.
[0467] In addition, the encoding process of the captured data can be performed by each terminal, on the server side, or shared by each other, without being limited to the distributed processing of the distribution itself. As an example, two processing cycles are usually performed in the encoding process. In the first cycle, the complexity or code amount of the image of the frame or scene unit is detected. In addition, in the second cycle, processing is performed to maintain the image quality and improve the encoding efficiency. For example, by performing the first encoding process by the terminal and the second encoding process by the server side that receives the content, the quality and efficiency of the content can be improved while reducing the processing load in each terminal. In this case, if there is a request for receiving and decoding almost in real time, the data completed by the first encoding by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.
[0468] As another example, the camera ex113 or the like extracts feature quantities from an image, compresses data on the feature quantities as metadata, and transmits the data to the server. The server determines the importance of an object based on the feature quantities, switches the quantization accuracy, and performs compression corresponding to the meaning of the image. Feature quantity data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression in the server. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding) and the server may perform encoding with a large processing load such as CABAC (context adaptive binary arithmetic coding).
[0469] As another example, in a stadium, shopping mall, or factory, there are multiple video data obtained by shooting roughly the same scene by multiple terminals. In this case, multiple terminals that have shot, and other terminals and servers that have not shot as needed, are used to distribute the encoding processing, such as GOP (Group of Picture) units, picture units, or tile units obtained by dividing pictures, so as to perform distributed processing. This can reduce delays and better achieve real-time performance.
[0470] In addition, since the plurality of image data are substantially the same scene, the server can also manage and / or instruct the image data taken by each terminal to refer to each other. Alternatively, the server can receive the encoded data from each terminal and change the reference relationship between the plurality of data, or modify or replace the image itself and re-encode it. In this way, a stream with improved quality and efficiency of each data can be generated.
[0471] In addition, the server may perform transcoding to change the encoding method of the video data and distribute the video data. For example, the server may convert the encoding method of the MPEG type to the VP type, or convert H.264 to H.265.
[0472] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, the following uses "server" or "terminal" as the subject of the process, but part or all of the process performed by the server can also be performed by the terminal, and part or all of the process performed by the terminal can also be performed by the server. In addition, the same is true for the decoding process.
[0473] [3D, multi-angle]
[0474] In recent years, there has been an increasing number of cases where images or videos taken by terminals such as multiple cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, such as different scenes or the same scene taken from different angles, are combined and used. The images taken by each terminal are combined based on the relative positional relationship between the terminals obtained separately or the areas where the feature points contained in the images are consistent.
[0475] The server can not only encode two-dimensional moving images, but also encode still images automatically or at a user-specified time based on scene analysis of moving images and send them to the receiving terminal. When the server is able to obtain the relative position relationship between the shooting terminals, it can generate not only two-dimensional moving images but also three-dimensional shapes of the scene based on images of the same scene shot from different angles. In addition, the server can also encode three-dimensional data generated by point clouds, etc., and can also select or reconstruct images from images shot by multiple terminals based on the results of identifying or tracking people or targets using three-dimensional data to generate images sent to the receiving terminal.
[0476] In this way, users can arbitrarily select each image corresponding to each shooting terminal to enjoy the scene, and can also enjoy the content of the image cut out of any viewpoint from the three-dimensional data reconstructed using multiple images or images. Furthermore, like the image, the sound can also be collected from multiple different angles, and the server can match the image and multiplex and send the sound from a specific angle or space with the image.
[0477] In addition, in recent years, content that establishes correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server produces viewpoint images for the right eye and the left eye respectively, and can be encoded to allow reference between viewpoint images through Multi-View Coding (MVC) or the like, or encoded as different streams without reference to each other. When decoding different streams, they can be reproduced synchronously with each other according to the user's viewpoint to reproduce a virtual three-dimensional space.
[0478] In the case of AR images, the server may overlap the virtual object information in the virtual space with the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device obtains or maintains the virtual object information and three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and produces overlapping data by smoothly connecting them. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to the entrustment of the virtual object information, and the server may produce overlapping data matching the received movement of the viewpoint based on the three-dimensional data maintained in the server, encode the overlapping data and distribute it to the decoding device. In addition, the overlapping data has an alpha value indicating the transmittance other than RGB, and the server sets the alpha value of the part other than the target produced according to the three-dimensional data to 0, etc., and encodes the part in a transparent state. Alternatively, the server may also set the RGB value of a specified value as the background like a chroma key, and generate data that sets the part other than the target as the background color.
[0479] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, or it can be performed on the server side, or it can be shared and performed. As an example, a terminal may first send a receiving request to the server, and other terminals may receive the content corresponding to the request and perform decoding processing, and send the decoded signal to a device with a display. By distributing the processing regardless of the performance of the communicative terminal itself and selecting appropriate content, data with better image quality can be reproduced. In addition, as another example, large-size image data can also be received by a TV, etc., and a portion of the image such as tiles after being divided can be decoded and displayed by the personal terminal of the viewer. In this way, while making the overall image shared, it is possible to confirm one's own area of responsibility or the area that wants to be confirmed in more detail at hand.
[0480] In addition, it is expected that in the future, when multiple short-range, medium-range or long-range wireless communications can be used both indoors and outdoors, the distribution system standards such as MPEG-DASH can be used to seamlessly receive content while switching appropriate data for the communication being connected. As a result, users can not only use their own terminals, but also freely select decoding devices or display devices such as displays set up indoors and outdoors to switch in real time. In addition, it is possible to switch the decoding terminal and the display terminal for decoding based on their own location information. As a result, it is also possible to move while displaying map information on a part of the wall or ground of a building next to a displayable device embedded in the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as caching the encoded data in a server that can be accessed from the receiving terminal in a short time, or copying the encoded data in the edge server of the content distribution service.
[0481] [Scalable Coding]
[0482] To switch content, use Fig.26 1. The scalable stream compressed and coded using the moving picture coding method described in the above embodiments is described as shown. For the server, a plurality of streams having the same content but different qualities may be provided as a single stream, or a structure in which the content is switched by utilizing the characteristics of a temporally / spatially scalable stream coded in layers as shown in the figure. That is, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as the state of the communication band, so that the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, if the subsequent video viewed on the smartphone ex115 while on the move is to be viewed on a device such as an Internet TV after returning home, the device can simply decode the same stream into different layers, thereby reducing the burden on the server side.
[0483] Furthermore, in addition to the hierarchical structure in which the picture is encoded for each layer and the enhancement layer exists above the base layer as described above, the enhancement layer may include meta-information such as statistical information based on the image, and the decoding side may generate high-definition content by super-resolving the picture of the base layer based on the meta-information. Super-resolution may be either an increase in the SN ratio or an increase in resolution at the same resolution. The meta-information includes information for determining linear or nonlinear filter coefficients used in super-resolution processing, or information for determining parameter values in filter processing, machine learning, or least squares calculation used in super-resolution processing.
[0484] Alternatively, the image may be divided into tiles according to the meaning of the object in the image, and the decoding side may decode only a part of the area by selecting the tile to be decoded. In addition, by storing the attributes of the object (people, cars, balls, etc.) and the position in the image (coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and determine the tile that includes the object. For example, Fig. 27 As shown, the meta information is stored using a data storage structure different from pixel data, such as SEI messages in HEVC. The meta information indicates, for example, the position, size, or color of the main object.
[0485] In addition, the meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding side can obtain the time when a specific person appears in the image, and by matching the information of the picture unit, it can determine the picture where the target exists and the position of the target in the picture.
[0486] [Web page optimization]
[0487] Fig.28 This is a diagram showing an example of a display screen of a web page on the computer ex111 or the like. Fig.29 1 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. Fig.28 and Fig.29 As shown, there are cases where a web page includes a plurality of link images that are links to image contents, and the way in which they are visible differs depending on the device being browsed. When a plurality of link images are visible on the screen, before a user explicitly selects a link image, or before a link image approaches the center of the screen, or before the entire link image enters the screen, the display device (decoding device) displays a still image or I picture of each content as a link image, or displays an image such as a GIF animation using a plurality of still images or I pictures, or receives only a base layer, decodes the image, and displays it.
[0488] When a linked image is selected by the user, the display device decodes the base layer with the highest priority. In addition, if there is information indicating that the content is scalable in the HTML constituting the web page, the display device can also decode to the enhancement layer. In addition, in order to ensure real-time performance, before selection or when the communication band is very tight, the display device can reduce the delay between the decoding time and the display time of the head picture (the delay from the start of decoding of the content to the start of display) by decoding and displaying only the pictures that are referenced in the front (I pictures, P pictures, B pictures that are only referenced in the front). In addition, the display device can also forcibly ignore the reference relationship of the pictures and set all B pictures and P pictures as forward references and roughly decode them, and perform normal decoding as the number of pictures received increases over time.
[0489] [Automatic driving]
[0490] Furthermore, when still images or video data such as two-dimensional or three-dimensional map information are transmitted and received for the purpose of automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta-information in addition to image data belonging to one or more layers, and decode them by establishing a correspondence between them. Furthermore, the meta-information may belong to a layer or may be multiplexed with image data alone.
[0491] In this case, since the car, drone or airplane including the receiving terminal is moving, the receiving terminal can switch base stations ex106 to ex110 to perform seamless reception and decoding by transmitting the location information of the receiving terminal when receiving a request. In addition, the receiving terminal can dynamically switch the degree to which meta-information is received or the degree to which map information is updated according to the user's selection, the user's condition or the state of the communication band.
[0492] As described above, in the content providing system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0493] [Distribution of Personal Content]
[0494] Furthermore, in the content supply system ex100, not only high-quality, long-duration contents provided by video distributors but also low-quality, short-duration contents provided by individuals can be unicasted or multicasted. It is expected that such personal contents will increase in the future. In order to make personal contents better, the server may perform encoding after editing. This can be achieved, for example, by the following structure.
[0495] After taking pictures in real time or accumulating them, the server performs recognition processing such as shooting errors, scene search, meaning analysis and target detection based on the original image or encoded data. In addition, based on the recognition results, the server manually or automatically corrects focus deviation or hand shaking, deletes scenes of low importance such as scenes with lower brightness than other pictures or scenes that are not in focus, or emphasizes the edges of the target, or changes the color tone. The server encodes the edited data based on the editing results. In addition, it is known that if the shooting time is too long, the viewing rate will decrease. The server can also automatically limit not only the scenes of low importance as mentioned above, but also the scenes with less movement based on the image processing results according to the shooting time, so as to become content within a specific time range. Alternatively, the server can also generate a summary based on the results of the scene's meaning analysis and encode it.
[0496] In addition, personal content may be written with content that infringes copyright, author's personality rights or portrait rights in its original state, or the scope of sharing may exceed the desired scope, which is inconvenient for individuals. Therefore, for example, the server may forcibly change the faces of people in the peripheral part of the screen or the home to an out-of-focus image for encoding. In addition, the server may also identify whether the face of a person different from the pre-registered person is captured in the encoded object image, and if so, apply mosaic or other processing to the face part. Alternatively, as a pre-processing or post-processing of the encoding, from the perspective of copyright, etc., the user may specify the person or background area that he wants to process the image, and the server may replace the specified area with another image or blur the focus. If it is a person, the image of the face part can be replaced while tracking the person in the moving image.
[0497] In addition, the viewing and listening of personal content with a small amount of data has a strong real-time requirement, so although it also depends on the bandwidth, the decoding device first receives, decodes and reproduces the basic layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the reproduction is reproduced more than twice, such as in the case of looping, the enhancement layer is also included to reproduce the high-definition image. In this way, if the stream is scalable encoded, it is possible to provide an experience in which the stream gradually becomes smoother and the image becomes better although the moving image is relatively rough when it is not selected or at the beginning of viewing. In addition to scalable encoding, the same experience can be provided when the rougher stream reproduced for the first time and the second stream encoded with reference to the moving image of the first time are composed of one stream.
[0498] [Other usage examples]
[0499] In addition, these encoding or decoding processes are usually processed in the LSIex500 included in each terminal. The LSIex500 can be a single chip or a structure composed of multiple chips. In addition, the software for encoding or decoding of moving images can be loaded into a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by the computer ex111, etc., and the encoding process and decoding process can be performed using the software. Furthermore, when the smart phone ex115 has a camera, the moving image data obtained by the camera can be sent. In this case, the moving image data is the data that has been encoded by the LSIex500 included in the smart phone ex115.
[0500] In addition, LSIex500 may also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. If the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software, and then obtains and reproduces the content.
[0501] Furthermore, the content supply system ex100 is not limited to the content supply system ex100 via the Internet ex101, and at least one of the video encoding device (video encoding device) or video decoding device (video decoding device) in each of the above-mentioned embodiments can be incorporated into a digital broadcasting system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by using broadcasting radio waves such as satellites, the content supply system ex100 is different from the structure that is easy for unicast in that it is suitable for multicast, but the same application can be made to the encoding process and the decoding process.
[0502] [Hardware Structure]
[0503] Fig.30 is a diagram showing a smartphone ex115. Fig.311 is a diagram showing a configuration example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of shooting video and still images, and a display unit ex458 for displaying decoded data such as the video shot by the camera unit ex465 and the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing coded data or decoded data such as shot video or still images, recorded sound, received video or still images, and mails, and a slot unit ex464 as an interface unit with a SIM ex468 for identifying a user and authenticating access to various data such as a network. In addition, an external memory may be used instead of the memory unit ex467.
[0504] In addition, the main control unit ex460 that performs integrated control of the display unit ex458 and the operation unit ex466 is connected to the power circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / separation unit ex453, the sound signal processing unit ex454, the slot unit ex464 and the memory unit ex467 via a bus ex470.
[0505] When the power key is turned on by a user's operation, the power circuit unit ex461 supplies power from the battery pack to each unit, thereby activating the smartphone ex115 to be able to operate.
[0506] The smartphone ex115 performs processes such as calls and data communications under the control of a main control unit ex460 including a CPU, ROM, and RAM. During a call, the voice signal collected by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, which is then subjected to spectrum diffusion processing by the modulation / demodulation unit ex452. The transmitting / receiving unit ex451 performs digital-to-analog conversion processing and frequency conversion processing, and then transmits the signal via the antenna ex450. In addition, the received data is amplified, subjected to frequency conversion processing and analog-to-digital conversion processing, subjected to spectrum inverse diffusion processing by the modulation / demodulation unit ex452, and converted into an analog voice signal by the voice signal processing unit ex454, which is then output from the voice output unit ex457. During data communications, text, still image, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main unit, and similar transmission and reception processing is performed. In the data communication mode, when transmitting video, still images, or video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of taking a video or still image by the camera unit ex465, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a predetermined manner, and performs modulation and conversion processing on the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.
[0507] When receiving an image attached to an e-mail or a chat tool, or an image linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of image data and a bit stream of audio data, and supplies the encoded image data to the image signal processing unit ex455 via the synchronous bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal by a moving image decoding method corresponding to the moving image encoding method shown in the above-mentioned embodiments, and displays the image or still image included in the linked moving image file on the display unit ex458 via the display control unit ex459. In addition, the audio signal processing unit ex454 decodes the audio signal and outputs the audio from the audio output unit ex457. In addition, since real-time streaming media is becoming more and more popular, the reproduction of audio may be socially inappropriate depending on the user's situation. Therefore, as an initial value, it is preferable to have a structure in which only the image data is reproduced without reproducing the audio signal. The sound may be reproduced synchronously only when the user performs an operation such as clicking on the video data.
[0508] In the description here, the smartphone ex115 is used as an example. However, as a terminal, three types of installation forms are conceivable, namely, a transmitting terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcasting system, it is assumed that multiplexed data in which audio data and the like are multiplexed in video data is received and transmitted. However, in addition to audio data, character data related to the video may be multiplexed in the multiplexed data, and video data itself may be received or transmitted instead of the multiplexed data.
[0509] In addition, although the main control unit ex460 including a CPU is assumed to control the encoding or decoding process, the terminal is often equipped with a GPU. Therefore, a structure can also be made to use the performance of the GPU to process a larger area together through a memory shared by the CPU and GPU, or a memory that manages addresses in a way that can be used together. In this way, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset) and transformation / quantization processing are performed together in units such as pictures by the GPU instead of the CPU.
[0510] Industrial Applicability
[0511] The present invention can be utilized in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, or a digital video camera.
[0512] Description of symbols
[0513] 100 Encoding device
[0514] 102 Division
[0515] 104 Subtraction Department
[0516] 106 Transformation Department
[0517] 108 Quantitative Department
[0518] 110 Entropy Coding Unit
[0519] 112, 204 Inverse Quantization Unit
[0520] 114, 206 Inverse transformation unit
[0521] 116, 208 Addition Department
[0522] 118, 210 memory blocks
[0523] 120, 212 loop filter unit
[0524] 122, 214 frame memory
[0525] 124, 216 Intra-frame prediction unit
[0526] 126, 218 Inter-frame prediction unit
[0527] 128, 220 Prediction and Control Department
[0528] 200 Decoding device
[0529] 202 Entropy Decoding Department
[0530] 1061, 2064 conversion mode determination unit
[0531] 1062, 2065 Size determination unit
[0532] 1063 First conversion base selection unit
[0533] 1064 First Transformation Section
[0534] 1065 Second Transformation Implementation Judgment Unit
[0535] 1066 Second conversion base selection unit
[0536] 1067 The Second Transformation
[0537] 1141, 2062 Second inverse transformation base selection unit
[0538] 1142, 2063 Second inverse transformation unit
[0539] 1143, 2066 1st inverse transformation base selection unit
[0540] 1144, 2067 The first inverse transformation unit
[0541] 2061 Second inverse transformation execution determination unit
[0542] 2062S 2nd base selection signal
[0543] 2064S Adaptive transform base selection mode signal
[0544] 2065S Size Signal
[0545] 2066S 1st base selection signal
Claims
1. A coding device comprising a circuit and a memory, wherein: The above circuit uses the above memory to perform the following processing: Determine whether an adaptive transform base selection mode for selecting a transform base according to the size of a coding target block is effective, When the above adaptive transform basis selection mode is valid, When the horizontal size of the encoding target block is larger than a threshold size, a first transform base is selected from a plurality of transform base candidates as a transform base in the horizontal direction, When the horizontal size of the encoding target block is smaller than the threshold size, a second transform basis is selected as a transform basis in the horizontal direction, and the second transform basis is a fixed transform basis. By performing a first horizontal transform on the residual of the encoding target block using the selected horizontal transform basis, a first transform coefficient is generated. A bit stream including information indicating whether the adaptive transform base selection mode is effective is generated.
2. A decoding device comprising a circuit and a memory, wherein: The above circuit uses the above memory to perform the following processing: determining whether an adaptive transform base selection mode for selecting a transform base according to the size of a decoding target block is effective, When the above adaptive transform basis selection mode is valid, When the horizontal size of the decoding target block is larger than the threshold size, a first inverse transform basis is selected from a plurality of inverse transform basis candidates as an inverse transform basis in the horizontal direction, When the horizontal size of the decoding target block is smaller than the threshold size, a second inverse transformation basis is selected as an inverse transformation basis in the horizontal direction, wherein the second inverse transformation basis is a fixed inverse transformation basis. A prediction residual is generated by performing a first inverse transform in the horizontal direction on the coefficients of the decoding target block using the selected inverse transform base in the horizontal direction.
3. A decoding method, wherein: determining whether an adaptive transform base selection mode for selecting a transform base according to the size of a decoding target block is effective, When the above adaptive transform basis selection mode is valid, When the horizontal size of the decoding target block is larger than the threshold size, a first inverse transform basis is selected from a plurality of inverse transform basis candidates as an inverse transform basis in the horizontal direction, When the horizontal size of the decoding target block is smaller than the threshold size, a second inverse transformation basis is selected as an inverse transformation basis in the horizontal direction, wherein the second inverse transformation basis is a fixed inverse transformation basis. A prediction residual is generated by performing a first inverse transform in the horizontal direction on the coefficients of the decoding target block using the selected inverse transform base in the horizontal direction.
4. A coding method, wherein: Determine whether an adaptive transform base selection mode for selecting a transform base according to the size of a coding target block is effective, When the above adaptive transform basis selection mode is valid, When the horizontal size of the encoding target block is larger than a threshold size, a first transform base is selected from a plurality of transform base candidates as a transform base in the horizontal direction, When the horizontal size of the encoding target block is smaller than the threshold size, a second transform basis is selected as a transform basis in the horizontal direction, and the second transform basis is a fixed transform basis. By performing a first horizontal transform on the residual of the encoding target block using the selected horizontal transform basis, a first transform coefficient is generated. A bit stream including information indicating whether the adaptive transform base selection mode is effective is generated.
Citation Information
Patent Citations
Method of generating quantized block
CN103096068A
Low complexity transform coding using adaptive DCT / DST for intra-prediction
CN103098473A