Encoding method and decoding method
By selecting the transform base according to the size of the encoded object block in video encoding, the problem of difficulty in improving compression efficiency and processing load in the prior art is solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202211053433.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-12-28
- Filing Date
- 2018-12-19
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2038-12-19
AI Technical Summary
In the existing video encoding technology, compression efficiency and processing load are difficult to improve simultaneously.
During the encoding and decoding process, an appropriate transformation base is selected according to the size of the encoded target block, and the first and second transformations are performed to generate an effective transformation coefficient.
This further improves the compression efficiency of video encoding and reduces processing load.
Smart Images

Figure CN115278238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an encoding method and a decoding method. Background Art
[0002] A video coding standard specification called HEVC (High-Efficiency Video Coding) is standardized by JCT-VC (Joint Collaborative Team on Video Coding).
[0003] Prior Art Documents
[0004] Non-Patent Documents
[0005] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC (High Efficiency Video Coding))
[0006] Non-Patent Document 2: Jianle Chen et al., Algorithm Description of Joint Exploration Test Model 5 (JEM5), Joint Video Exploration Team (JVET) of ITU-T SG16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 5th Meeting: Geneva, CH, Document: JVET-E1001, January 2017 Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] In such encoding and decoding technologies, further improvement in compression efficiency and reduction in processing load are required.
[0009] Therefore, the present invention provides an encoding device, a decoding device, an encoding method, or a decoding method capable of achieving further improvement in compression efficiency and reduction in processing load.
[0010] Means for Solving the Problems
[0011] Encoding method for a technical solution of the present invention, wherein it is determined whether the mode of selecting a transform basis according to the size of an encoding target block is valid. When the above mode is valid, when the vertical size of the above encoding target block is greater than a threshold size, a first transform basis is selected from a plurality of transform basis candidates as the transform basis in the vertical direction. When the vertical size of the above encoding target block is less than the above threshold size, a second transform basis is selected as the transform basis in the vertical direction. The above second transform basis is a fixed transform basis. By performing a first transform on the residual of the above encoding target block using the selected transform basis in the vertical direction, first transform coefficients are generated.
[0012] Decoding method for a technical solution of the present invention, wherein it is determined whether the mode of selecting an inverse transform basis according to the size of a decoding target block is valid. When the above mode is valid, when the vertical size of the above decoding target block is greater than a threshold size, a first inverse transform basis is selected from a plurality of inverse transform basis candidates as the inverse transform basis in the vertical direction. When the vertical size of the above decoding target block is less than the above threshold size, a second inverse transform basis is selected as the inverse transform basis in the vertical direction. The above second inverse transform basis is a fixed inverse transform basis. By performing a first inverse transform on the coefficients of the above decoding target block using the selected inverse transform basis in the vertical direction, a prediction residual is generated.
[0013] In addition, these general or specific technical solutions can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0014] Advantages of the Invention
[0015] The present invention can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve a further improvement in compression efficiency and a reduction in processing load. Brief Description of the Drawings
[0016] Figure 1 It is a block diagram showing the functional structure of an encoding device according to Embodiment 1.
[0017] Figure 2 It is a diagram showing an example of block division in Embodiment 1.
[0018] Figure 3 It is a table showing transform basis functions corresponding to respective transform types.
[0019] Figure 4A It is a diagram showing an example of the shape of a filter used in ALF.
[0020] Figure 4B It is a diagram showing another example of the shape of a filter used in ALF.
[0021] Figure 4C It is a diagram showing another example of the shape of the filter used in the ALF.
[0022] Figure 5A It is a diagram showing 67 intra prediction modes of intra prediction.
[0023] Figure 5B It is a flowchart for explaining the outline of the predicted image correction process based on the OBMC process.
[0024] Figure 5C It is a conceptual diagram for explaining the outline of the predicted image correction process based on the OBMC process.
[0025] Figure 5D It is a diagram showing an example of FRUC.
[0026] Figure 6 It is a diagram for explaining the pattern matching (bidirectional matching) between two blocks along the motion trajectory.
[0027] Figure 7 It is a diagram for explaining the pattern matching (template matching) between a template in the current picture and a block in the reference picture.
[0028] Figure 8 It is a diagram for explaining a model assuming uniform linear motion.
[0029] Figure 9A It is a diagram for explaining the derivation of the motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.
[0030] Figure 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.
[0031] Figure 9C It is a conceptual diagram for explaining the outline of the DMVR process.
[0032] Figure 9D It is a diagram for explaining the outline of the predicted image generation method that employs the luminance correction process based on the LIC process.
[0033] Figure 10 It is a block diagram showing the functional structure of the decoding device related to Embodiment 1.
[0034] Figure 11A It is a block diagram showing the internal structure of the transformation unit of the encoding device related to the first mode of Embodiment 1.
[0035] Figure 11B It is a block diagram showing the internal structure of the inverse transformation unit of the encoding device related to the first mode of Embodiment 1.
[0036] Figure 12A It is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the first mode of Embodiment 1.
[0037] Figure 12B It is a flowchart showing a modified example of the processing of the transformation unit and the quantization unit of the encoding device according to the first mode of Embodiment 1.
[0038] Figure 13A It is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the second mode of Embodiment 1.
[0039] Figure 13B It is a flowchart showing the processing of the entropy encoding unit of the encoding device according to the second mode of Embodiment 1.
[0040] Figure 14 It is a diagram showing a specific example of the syntax in the second mode of Embodiment 1.
[0041] Figure 15 It is a table showing a specific example of the presence or absence of encoding of the transform basis and the signal used in the second mode of Embodiment 1.
[0042] Figure 16 It is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the third mode of Embodiment 1.
[0043] Figure 17A It is a flowchart showing the processing of the transformation unit and the quantization unit of the encoding device according to the fourth mode of Embodiment 1.
[0044] Figure 17B It is a flowchart showing the processing of the entropy encoding unit of the encoding device according to the fourth mode of Embodiment 1.
[0045] Figure 18 It is a diagram showing a specific example of the syntax in the fourth mode of Embodiment 1.
[0046] Figure 19 It is a table showing a specific example of the presence or absence of encoding of the transform basis and the signal used in the fourth mode of Embodiment 1.
[0047] Figure 20 It is a block diagram showing the internal structure of the inverse transformation unit of the decoding device according to the fifth mode of Embodiment 1.
[0048] Figure 21 It is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the fifth mode of Embodiment 1.
[0049] Figure 22AThis is a flowchart showing the processing of the entropy decoding unit of the decoding device according to the sixth aspect of the first embodiment.
[0050] Figure 22B This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the sixth aspect of the first embodiment.
[0051] Figure 23 This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the seventh aspect of the first embodiment.
[0052] Figure 24A This is a flowchart showing the processing of the entropy decoding unit of the decoding device according to the eighth aspect of the first embodiment.
[0053] Figure 24B This is a flowchart showing the processing of the inverse quantization unit and the inverse transformation unit of the decoding device according to the eighth aspect of the first embodiment.
[0054] Figure 25 This is the overall structure diagram of the content supply system that implements content distribution services.
[0055] Figure 26 This is a diagram showing an example of a coding structure for scalable coding.
[0056] Figure 27 This is a diagram showing an example of a coding structure for scalable coding.
[0057] Figure 28 This is a diagram showing an example of a display screen of a web page.
[0058] Figure 29 This is a diagram showing an example of a display screen of a web page.
[0059] Figure 30 This is a diagram showing an example of a smart phone.
[0060] Figure 31 This is a block diagram showing an example of the structure of a smart phone. Specific implementation method
[0061] The following is a detailed description of the implementation method with reference to the accompanying drawings.
[0062] In addition, the embodiments described below are all inclusive or specific examples. The numerical values, shapes, materials, components, configurations and connection forms of components, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the claims. In addition, among the components of the following embodiments, components that are not described in the independent claims representing the highest concept are described as arbitrary components.
[0063] (Embodiment 1)
[0064] First, as an example of an encoding device and a decoding device that can apply the processes and / or structures described in each aspect of the present invention described below, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device that can apply the processes and / or structures described in each aspect of the present invention, and the processes and / or structures described in each aspect of the present invention can also be implemented in encoding devices and decoding devices different from Embodiment 1.
[0065] When applying the processes and / or structures described in each aspect of the present invention to Embodiment 1, for example, any of the following may be performed.
[0066] (1) For the encoding device or decoding device of Embodiment 1, replace the constituent elements corresponding to the constituent elements described in each aspect of the present invention among the plurality of constituent elements constituting the encoding device or decoding device with the constituent elements described in each aspect of the present invention;
[0067] (2) For the encoding device or decoding device of Embodiment 1, after arbitrarily changing the functions or processes implemented by a part of the plurality of constituent elements constituting the encoding device or decoding device, such as adding, replacing, or deleting, replace the constituent elements corresponding to the constituent elements described in each aspect of the present invention with the constituent elements described in each aspect of the present invention;
[0068] (3) For the method implemented by the encoding device or decoding device of Embodiment 1, after adding a process and / or arbitrarily changing a part of the plurality of processes included in the method, such as replacing or deleting, replace the process corresponding to the process described in each aspect of the present invention with the process described in each aspect of the present invention;
[0069] (4) Combine a part of the plurality of constituent elements constituting the encoding device or decoding device of Embodiment 1 with the constituent elements described in each aspect of the present invention, the constituent elements having a part of the functions of the constituent elements described in each aspect of the present invention, or the constituent elements implementing a part of the processes implemented by the constituent elements described in each aspect of the present invention and implement;
[0070] (5) Combine an element that has a part of the functions of a part of the multiple elements constituting the encoding device or decoding device of Embodiment 1, or an element that implements a part of the processing implemented by a part of the multiple elements constituting the encoding device or decoding device of Embodiment 1, with the elements described in each aspect of the present invention, an element that has a part of the functions of the elements described in each aspect of the present invention, or an element that implements a part of the processing implemented by the elements described in each aspect of the present invention;
[0071] (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processing corresponding to the processing described in each aspect of the present invention among the multiple processes included in the method with the processing described in each aspect of the present invention;
[0072] (7) Combine a part of the processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processing described in each aspect of the present invention and implement them.
[0073] In addition, the implementation manners of the processing and / or structure described in each aspect of the present invention are not limited to the above examples. For example, it can also be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, and the processing and / or structure described in each aspect can also be implemented independently. In addition, the processing and / or structure described in different aspects can also be combined and implemented.
[0074] [Outline of Encoding Device]
[0075] First, the outline of the encoding device of Embodiment 1 will be described. Figure 1 FIG. is a block diagram showing the functional structure of the encoding device 100 of Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in units of blocks.
[0076] As Figure 1 shown, the encoding device 100 is a device that encodes images in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0077] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as a splitting unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a loop filtering unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. In addition, the encoding device 100 may also be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filtering unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0078] Hereinafter, each component included in the encoding device 100 will be described.
[0079] [Splitting Unit]
[0080] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128). Such blocks of a fixed size are sometimes referred to as coding tree units (CTUs). And the splitting unit 102 splits each block of a fixed size into blocks of a variable size (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block splitting. Such blocks of a variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture may also be used as the processing units for CUs, PUs, and TUs.
[0081] Figure 2 is a diagram showing an example of block splitting in Embodiment 1. In Figure 2 the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.
[0082] Here, the block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into 4 square 64×64 blocks (quadtree block splitting).
[0083] The 64×64 block in the upper left is further vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is further vertically divided into two rectangular 16×64 blocks (binary tree block division). As a result, the 64×64 block in the upper left is divided into two 16×64 blocks 11, 12 and a 32×64 block 13.
[0084] The 64×64 block in the upper right is horizontally divided into two rectangular 64×32 blocks 14, 15 (binary tree block division).
[0085] The 64×64 block in the lower left is divided into four square 32×32 blocks (quadtree block division). The upper left block and the lower right block among the four 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the 64×64 block in the lower left is divided into 16×32 blocks 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.
[0086] The 64×64 block 23 in the lower right is not divided.
[0087] As described above, in Figure 2 , the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block division. Such a division is sometimes referred to as QTBT (quad - tree plus binary tree) division.
[0088] In addition, in Figure 2 , one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block can also be divided into three blocks (ternary tree division). A division including such a ternary tree division is sometimes referred to as MBT (multi type tree) division.
[0089] [Subtraction section]
[0090] The subtraction section 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division section 102. That is, the subtraction section 104 calculates the prediction error (also referred to as the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction section 104 outputs the calculated prediction error to the transformation section 106.
[0091] The original signal is the input signal of the encoding device 100 and is a signal representing the images of the respective pictures constituting the moving image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0092] [Transformation unit]
[0093] The transformation unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the vectorization unit 108. Specifically, the transformation unit 106, for example, performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0094] In addition, the transformation unit 106 may adaptively select a transformation type from among multiple transformation types and use a transform basis function corresponding to the selected transformation type to transform the prediction error into transform coefficients. Such a transformation is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0095] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transform basis functions corresponding to the respective transformation types. In Figure 3 N represents the number of input pixels. The selection of the transformation type from among these multiple transformation types may depend, for example, on the type of prediction (intra prediction and inter prediction) or on the intra prediction mode.
[0096] Information indicating whether to apply such EMT or AMT (for example, referred to as an AMT flag) and information indicating the selected transformation type are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and may also be other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0097] In addition, the transformation unit 106 can also perform inverse transformation on the transformation coefficients (transformation results). Such inverse transformation is called AST (adaptive secondary transform) or NSST (non-separable secondary transform) in some cases. For example, the transformation unit 106 performs inverse transformation on each sub-block (e.g., 4×4 sub-block) included in the block of transformation coefficients corresponding to the intra-prediction error. Information indicating whether to apply NSST and information related to the transformation matrix used in NSST are signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0098] Here, a separable transformation refers to a method of performing multiple transformations by separating them in each direction according to the number of dimensions of the input, and a non-separable transformation refers to a method of treating two or more dimensions as one dimension and performing a transformation together when the input is multi-dimensional.
[0099] For example, as an example of a non-separable transformation, when the input is a 4×4 block, it can be regarded as a permutation with 16 elements, and a transformation process is performed on this permutation with a 16×16 transformation matrix.
[0100] In addition, similarly, a method of performing Givens rotation on this permutation multiple times (Hypercube Givens Transform) after regarding a 4×4 input block as a permutation with 16 elements is also an example of a non-separable transformation.
[0101] [Quantization Unit]
[0102] The quantization unit 108 quantizes the transformation coefficients output from the transformation unit 106. Specifically, the quantization unit 108 scans the transformation coefficients of the current block in a specified scanning order, and quantizes the transformation coefficients based on the quantization parameter (QP) corresponding to the scanned transformation coefficients. Then, the quantization unit 108 outputs the quantized transformation coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112.
[0103] The specified order is the order for quantization / inverse quantization of the transformation coefficients. For example, the specified scanning order is defined by ascending frequency (order from low frequency to high frequency) or descending frequency (order from high frequency to low frequency).
[0104] A quantization parameter refers to a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0105] [Entropy Encoding Unit]
[0106] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) by performing variable-length encoding on the quantized coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110 binarizes the quantized coefficients, for example, and performs arithmetic encoding on the binary signal.
[0107] [Inverse Quantization Unit]
[0108] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. And the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114.
[0109] [Inverse Transform Unit]
[0110] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. And the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0111] In addition, since information is lost through quantization in the restored prediction error, it is not consistent with the prediction error calculated by the subtraction unit 104. That is, the restored prediction error contains a quantization error.
[0112] [Addition Unit]
[0113] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. And the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a local decoded block.
[0114] [Block Memory]
[0115] The block memory 118 is a storage unit for storing blocks within the picture to be encoded (hereinafter referred to as the current picture) that are referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the addition unit 116.
[0116] [Loop Filter Unit]
[0117] The loop filter unit 120 performs loop filtering on the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to the filtering used within the coding loop (in-loop filtering), and includes, for example, deblocking filtering (DF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).
[0118] In ALF, a least squares error filter used to remove coding distortion is adopted. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters is adopted based on the direction and activity of the locality-based gradient.
[0119] Specifically, first, sub-blocks (such as 2×2 sub-blocks) are classified into multiple classes (such as 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (such as 0 to 2 or 0 to 4) and the activity value A of the gradient (such as 0 to 4), the classification value C (such as C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes (such as 15 or 25 classes).
[0120] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (such as horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantifying the added result.
[0121] Based on the result of such classification, the filter to be used for the sub-block is determined from among multiple filters.
[0122] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. Figures 4A - 4C It is a diagram showing multiple examples of the shape of the filter used in ALF. Figure 4A It represents a 5×5 diamond-shaped filter, Figure 4B It represents a 7×7 diamond-shaped filter, Figure 4C It represents a 9×9 diamond-shaped filter. The information indicating the shape of the filter is signaled at the picture level. In addition, the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, and can also be other levels (such as sequence level, slice level, tile level, CTU level, or CU level).
[0123] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined at the CU level whether to adopt ALF, and for chrominance difference, it is determined at the picture level whether to adopt ALF. The information indicating the on / off of ALF is signaled at the picture level or CU level. In addition, the signaling of the information indicating the on / off of ALF does not need to be limited to the picture level or CU level, and can also be other levels (such as sequence level, slice level, tile level, or CTU level).
[0124] Coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Additionally, the signaling of the coefficient sets does not need to be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0125] [Frame memory]
[0126] The frame memory 122 is a storage unit for storing reference pictures used in inter-frame prediction, and is also sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0127] [Intra-frame prediction unit]
[0128] The intra-frame prediction unit 124 performs intra-frame prediction (also referred to as intra-picture prediction) of the current block by referring to the block within the current picture stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.
[0129] For example, the intra-frame prediction unit 124 performs intra-frame prediction using one of a plurality of predefined intra-frame prediction modes. The plurality of intra-frame prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0130] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0131] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. Additionally, the plurality of directional prediction modes may also include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Figure 5A It is a diagram showing 67 intra-frame prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. The solid arrows represent 33 directions defined by the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.
[0132] In addition, in the intra prediction of the chrominance blocks, the luminance blocks may also be referred to. That is, the chrominance components of the current block may also be predicted based on the luminance component of the current block. Such intra prediction is called CCLM (cross-component linear model) prediction in some cases. The intra prediction mode of the chrominance blocks referring to the luminance blocks (for example, called the CCLM mode) may also be added as one of the intra prediction modes of the chrominance blocks.
[0133] The intra prediction unit 124 may also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is called PDPC (position dependent intraprediction combination) in some cases. Information indicating whether PDPC is adopted (for example, called the PDPC flag) is signaled at the CU level, for example. In addition, the signaling of this information is not limited to the CU level and may also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0134] [Inter prediction unit]
[0135] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or sub-blocks (for example, 4×4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. And, the inter prediction unit 126 performs motion compensation using the motion information (for example, motion vector) obtained by the motion estimation, thereby generating an inter prediction signal for the current block or sub-block. And, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0136] The motion information used in the motion compensation is signaled. A predicted motion vector may also be used in the signaling of the motion vector. That is, the difference between the motion vector and the predicted motion vector may also be signaled.
[0137] Alternatively, it may be that not only the motion information of the current block obtained by motion estimation is used, but also the motion information of adjacent blocks is used to generate an inter-frame prediction signal. Specifically, it may also be that a prediction signal based on the motion information obtained by motion estimation and a prediction signal based on the motion information of adjacent blocks are weighted and added together, thereby generating an inter-frame prediction signal in units of sub-blocks within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0138] In such an OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) is signaled at the sequence level. In addition, information indicating whether the OBMC mode is adopted (e.g., referred to as the OBMC flag) is signaled at the CU level. Additionally, the levels at which these pieces of information are signaled do not need to be limited to the sequence level and the CU level, and may also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0139] A more specific description of the OBMC mode will be given. Figure 5B and Figure 5C are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC processing.
[0140] First, using the motion vector (MV) assigned to the coding target block, a predicted image (Pred) obtained by normal motion compensation is acquired.
[0141] Next, the motion vector (MV_L) of the already-coded left adjacent block is adopted for the coding target block to obtain a predicted image (Pred_L), and the first correction of the predicted image is performed by weighted superposition of the above-mentioned predicted image and Pred_L.
[0142] Similarly, the motion vector (MV_U) of the already-coded upper adjacent block is adopted for the coding target block to obtain a predicted image (Pred_U), and the second correction of the predicted image is performed by weighted superposition of the predicted image after the above-mentioned first correction and Pred_U, and this is used as the final predicted image.
[0143] In addition, the method of two-stage correction using the left adjacent block and the upper adjacent block is described here, but it may also be configured to perform more than two-stage corrections using the right adjacent block and the lower adjacent block.
[0144] In addition, the region for superposition may not be the entire pixel region of the block, but only a partial region near the block boundary.
[0145] In addition, the prediction image correction process based on one reference image has been described here, but the same applies to the case of correcting the prediction image based on multiple reference images. After obtaining the corrected prediction images based on the respective reference images, the obtained prediction images are further superimposed to obtain the final prediction image.
[0146] In addition, the processing target block described above may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.
[0147] As a method for determining whether to use the OBMC process, for example, there is a method of using a signal indicating whether to use the OBMC process, that is, obmc_flag. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as obmc_flag and encoding is performed using the OBMC process. If it does not belong to a region with complex motion, the value 0 is set as obmc_flag and encoding is performed without using the OBMC process. On the other hand, in a decoding device, the obmc_flag described in the stream is decoded, and whether to use the OBMC process is switched according to the value for decoding.
[0148] In addition, the motion information may not be signaled and may be derived on the decoding device side. For example, the merge mode specified by the H.265 / HEVC standard may be used. In addition, for example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, motion estimation is performed without using the pixel values of the current block.
[0149] Here, the mode of performing motion estimation on the decoding device side is described. The mode of performing motion estimation on the decoding device side may be a mode called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.
[0150] In Figure 5D shows an example of FRUC processing. First, referring to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector are generated (which may be shared with the merge list). Then, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, the evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0151] And, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, a motion vector for the current block can also be derived by performing pattern matching in a peripheral region of the position within the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV can also be searched by the same method. In the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and this is used as the final MV for the current block. Additionally, a structure that does not perform this process can also be formed.
[0152] The exact same process can also be performed when processing in units of sub-blocks.
[0153] In addition, regarding the evaluation value, it is calculated by obtaining a difference value of the reconstructed image through pattern matching between the region within the reference picture corresponding to the motion vector and a specified region. Additionally, it can also be that, in addition to the difference value, other information is used to calculate the evaluation value.
[0154] As pattern matching, the first pattern matching or the second pattern matching is used. In some cases, the first pattern matching and the second pattern matching are respectively referred to as bilateral matching and template matching.
[0155] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate, a region within another reference picture along the motion trajectory of the current block is used.
[0156] Figure 6 It is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along the motion trajectory. As Figure 6 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two motion vectors (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the above candidate MV by the display time interval is obtained, and the evaluation value is calculated using the obtained difference value. A candidate MV with the best evaluation value among multiple candidate MVs can be selected as the final MV.
[0157] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, a mirror-symmetric bidirectional motion vector is derived.
[0158] In the second pattern matching, pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture (e.g., an upper and / or left adjacent block)) and a block within the reference picture. Thus, in the second pattern matching, as the specified region for calculating the evaluation value for the above candidates, a block adjacent to the current block within the current picture is used.
[0159] Figure 7 It is a diagram for explaining an example of pattern matching (template matching) between a template within the current picture and a block within the reference picture. As Figure 7 shown, in the second pattern matching, by searching for the block in the reference picture (Ref0) that best matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic), the motion vector of the current block is derived. Specifically, for the current block, the difference between the reconstructed images of the left adjacent and / or upper adjacent encoded regions and the reconstructed image at the equivalent position within the encoded reference picture (Ref0) specified by the candidate MV is derived, and the obtained difference value is used to calculate the evaluation value. Among the multiple candidate MVs, the candidate MV with the best evaluation value is selected as the best candidate MV.
[0160] Information indicating whether to adopt the FRUC mode (e.g., called the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the method of pattern matching (the first pattern matching or the second pattern matching) (e.g., called the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0161] Here, a mode of deriving a motion vector based on a model assuming uniform linear motion is described. This mode has a case called BIO (bi-directional optical flow).
[0162] Figure 8 It is a diagram for explaining a model assuming uniform linear motion. InFigure 8 Among them, (v x , v y ) represents the velocity vector, and τ 0 , τ 1 respectively represent the time distances between the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) represents the motion vector corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) represents the motion vector corresponding to the reference picture Ref 1 .
[0163] At this time, under the assumption of a uniform linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1 , MVy 1 ) are respectively expressed as (vxτ 0 , vyτ 0 ) and (-vxτ 1 , -vyτ 1 ), and the following optical flow equation (1) holds.
[0164] [Equation 1]
[0165]
[0166] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a merge list, etc. are corrected in pixel units.
[0167] In addition, the motion vector can also be derived on the decoding device side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.
[0168] Here, a mode of deriving a motion vector in units of sub-blocks based on motion vectors of multiple adjacent blocks will be described. In this mode, there is a case called the affine motion compensation prediction mode.
[0169] Figure 9A It is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on motion vectors of multiple adjacent blocks. In Figure 9A , the current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v 0 of the upper-left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v 1 of the upper-right control point of the current block is derived. And, using the two motion vectors v 0 and v 1 , the motion vectors (v x , v y ) of each sub-block within the current block are derived by the following equation (2).
[0170] [Equation 2]
[0171]
[0172] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a preset weight coefficient.
[0173] In such an affine motion compensation prediction mode, several modes with different methods for deriving the motion vectors of the upper-left and upper-right control points may also be included. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. In addition, the signaling of the information indicating this affine motion compensation prediction mode does not need to be limited to the CU level, and it may also be other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0174] [Prediction control unit]
[0175] The prediction control unit 128 selects one of the intra-prediction signal and the inter-prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.
[0176] Here, an example of deriving the motion vector of the coded picture by the merge mode will be described. Figure 9B It is a diagram for explaining the outline of the motion vector derivation process based on the merge mode.
[0177] First, a candidate prediction MV list registered with the prediction MV is generated. As candidates for the prediction MV, there are the MVs of multiple coded blocks spatially located around the coding target block, i.e., spatially adjacent prediction MVs, the MVs of blocks near the projection of the position of the coding target block in the coded reference picture, i.e., temporally adjacent prediction MVs, the MV generated by combining the MV values of the spatially adjacent prediction MV and the temporally adjacent prediction MV, i.e., combined prediction MV, and the MV with a value of zero, i.e., zero prediction MV, etc.
[0178] Next, by selecting one prediction MV from among the multiple prediction MVs registered in the prediction MV list, it is determined as the MV of the coding target block.
[0179] Furthermore, in the variable-length coding unit, merge_idx, which is a signal indicating which prediction MV is selected, is described in the stream and encoded.
[0180] In addition, in Figure 9B The prediction MVs registered in the prediction MV list described are an example, and it can also be a structure with a number different from that in the figure, or a structure that does not include some types of the prediction MVs in the figure, or a structure with prediction MVs other than the types in the figure added.
[0181] In addition, the MV of the coding target block derived through the merge mode can also be used for the subsequent DMVR processing to determine the final MV.
[0182] Here, an example of determining the MV using the DMVR processing is described.
[0183] Figure 9C is a conceptual diagram for explaining the outline of the DMVR processing.
[0184] First, the optimal MVP set for the processing target block is used as the candidate MV, and according to the above candidate MV, reference pixels are obtained from the first reference picture of the processed picture in the L0 direction and the second reference picture of the processed picture in the L1 direction respectively, and a template is generated by taking the average of each reference pixel.
[0185] Next, using the above template, the peripheral regions of the candidate MVs in the first reference picture and the second reference picture are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, regarding the cost value, it is calculated using the difference values between the pixel values of the template and the pixel values of the search region and the MV value, etc.
[0186] In addition, in the coding device and the decoding device, the outline of the processing described here is basically common.
[0187] In addition, even if it is not the processing itself described here, as long as it is a process that can search the periphery of the candidate MV and derive the final MV, other processes can also be used.
[0188] Here, a mode of generating a prediction image using the LIC process will be described.
[0189] Figure 9D It is a diagram for explaining an outline of a method for generating a prediction image using a luminance correction process based on the LIC process.
[0190] First, an MV is derived that is used to obtain a reference image corresponding to an encoding target block from a reference picture that is an encoded picture.
[0191] Next, for the encoding target block, using the luminance pixel values of the left-adjacent and upper-adjacent encoded peripheral reference regions and the luminance pixel values at the same position in the reference picture specified by the MV, information indicating how the luminance values change in the reference picture and the encoding target picture is extracted, and a luminance correction parameter is calculated.
[0192] By performing a luminance correction process on the reference image in the reference picture specified by the MV using the above luminance correction parameter, a prediction image for the encoding target block is generated.
[0193] In addition, Figure 9D The shape of the above peripheral reference region in [[ ]] is an example, and shapes other than this can also be used.
[0194] Furthermore, a process of generating a prediction image based on one reference picture has been described here, but the same applies to the case of generating a prediction image based on multiple reference pictures. After performing a luminance correction process on the reference images obtained from each reference picture in the same manner, a prediction image is generated.
[0195] As a method for determining whether to adopt the LIC process, for example, there is a method of using lic_flag which is a signal indicating whether to adopt the LIC process. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region where a luminance change has occurred. In the case where it belongs to a region where a luminance change has occurred, a value 1 is set as lic_flag and encoding is performed using the LIC process. In the case where it does not belong to a region where a luminance change has occurred, a value 0 is set as lic_flag and encoding is performed without using the LIC process. On the other hand, in a decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to adopt the LIC process according to its value.
[0196] As another method for determining whether to use LIC processing, there is also a method of determining whether LIC processing is used in surrounding blocks. As a specific example, when the encoding target block is in merge mode, it is determined whether the surrounding coded blocks selected when deriving the MV in merge mode processing have been encoded using LIC processing, and based on the result, it is switched whether to use LIC processing for encoding. In addition, in the case of this example, the processing in decoding is exactly the same.
[0197] [Overview of the decoding device]
[0198] Next, an overview of a decoding device that can decode the coded signal (coded bit stream) output from the coding device 100 will be described. Figure 10 is a block diagram showing the functional structure of a decoding device 200 according to Embodiment 1. The decoding device 200 is a moving picture / image decoding device that decodes a moving picture / image in units of blocks.
[0199] If Figure 10 As shown in , the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.
[0200] The decoding device 200 is implemented, for example, by a general-purpose processor and a memory. In this case, when the processor executes a software program stored in the memory, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a loop filter unit 212, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. In addition, the decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.
[0201] The following describes the various components included in the decoding device 200.
[0202] [Entropy Decoding Department]
[0203] The entropy decoding unit 202 performs entropy decoding on the coded bit stream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bit stream into a binary signal, for example. Next, the entropy decoding unit 202 debinarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks.
[0204] [Inverse Quantization Unit]
[0205] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0206] [Inverse Transform Unit]
[0207] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients that are the input from the inverse quantization unit 204.
[0208] For example, when the information read from the encoded bitstream indicates the use of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type read from the bitstream.
[0209] In addition, for example, when the information read from the encoded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re - transform to the transform coefficients.
[0210] [Addition Unit]
[0211] The addition unit 208 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 206, to the prediction sample, which is the input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0212] [Block Memory]
[0213] The block memory 210 is a storage unit for storing blocks within the picture to be decoded (hereinafter referred to as the current picture) that are referenced in intra - frame prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.
[0214] [Loop Filter Unit]
[0215] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208, and outputs the filtered reconstructed block to the frame memory 214, the display device, and so on.
[0216] When the information indicating the on / off of ALF read from the encoded bitstream indicates that ALF is on, one filter is selected from multiple filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.
[0217] [Frame Memory]
[0218] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.
[0219] [Intra Prediction Unit]
[0220] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction on the basis of the intra prediction mode read from the coded bitstream and referring to the blocks within the current picture stored in the block memory 210. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance differences) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0221] In addition, when the intra prediction mode that refers to the luminance block is selected for the intra prediction of the chrominance block, the intra prediction unit 216 may also predict the chrominance component of the current block on the basis of the luminance component of the current block.
[0222] Furthermore, when the information read from the coded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction on the basis of the gradients of the reference pixels in the horizontal / vertical directions.
[0223] [Inter-Frame Prediction Unit]
[0224] The inter-frame prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter-frame prediction unit 218 performs motion compensation using the motion information (e.g., motion vector) read from the coded bitstream, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.
[0225] In addition, when the information read from the coded bitstream indicates the adoption of the OBMC mode, the inter-frame prediction unit 218 generates an inter-frame prediction signal using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks.
[0226] Furthermore, when the information read from the coded bitstream indicates the adoption of the FRUC mode, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the coded stream, thereby deriving the motion information. And the inter-frame prediction unit 218 performs motion compensation using the derived motion information.
[0227] In addition, when the inter-frame prediction unit 218 adopts the BIO mode, it derives a motion vector based on a model assuming uniform linear motion. In addition, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is adopted, the inter-frame prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of multiple adjacent blocks.
[0228] [Prediction control unit]
[0229] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, uses the selected signal as the prediction signal, and outputs it to the addition unit 208.
[0230] (First mode of Embodiment 1)
[0231] Next, the first mode of Embodiment 1 will be specifically described with reference to the drawings.
[0232] [Internal structure of the transformation unit of the encoding device]
[0233] First, refer to Figure 11A This describes the internal structure of the transformation unit 106 of the encoding device 100 according to this mode. Figure 11A It is a block diagram showing the internal structure of the transformation unit 106 of the encoding device 100 according to the first mode of Embodiment 1.
[0234] As Figure 11A shown, the transformation unit 106 according to this mode includes a transformation mode determination unit 1061, a size determination unit 1062, a first transformation basis selection unit 1063, a first transformation unit 1064, a second transformation execution determination unit 1065, a second transformation basis selection unit 1066, and a second transformation unit 1067.
[0235] The transformation mode determination unit 1061 determines whether the adaptive transformation basis selection mode is valid for the block to be encoded. The adaptive transformation basis selection mode is a mode of adaptively selecting a transformation basis from one or more candidates of the first transformation basis. The determination of whether the adaptive transformation basis selection mode is valid is made based on, for example, the identification information of the first transformation basis or the adaptive transformation basis selection mode.
[0236] The size determination unit 1062 determines whether the horizontal size of the block to be encoded exceeds the first horizontal threshold size. In addition, the size determination unit 1062 determines whether the vertical size of the block to be encoded exceeds the first vertical threshold size. The first horizontal threshold size may be the same as or different from the first vertical threshold size. The first horizontal threshold size and the first vertical threshold size may also be predefined in a standard specification, for example. In addition, for example, the first horizontal threshold size and the first vertical threshold size may be sizes determined based on the image or encoded in the bitstream.
[0237] The first transform basis selection unit 1063 selects the first transform basis. In the present invention, selecting a basis means, in addition to selecting at least one basis from among a plurality of basis candidates, also including determining or setting at least one basis without a plurality of basis candidates.
[0238] When the adaptive transform basis selection mode is not valid, the first transform basis selection unit 1063 selects one basic transform basis as the first transform basis in the horizontal direction and the vertical direction. Further, when the adaptive transform basis selection mode is valid, based on the horizontal size and the vertical size of the coding target block, the first transform basis selection unit 1063 selects the first transform basis in the horizontal direction and the vertical direction as follows in (1) to (4).
[0239] (1) When the horizontal size of the coding target block is greater than the first horizontal threshold size, the first transform basis selection unit 1063 adaptively selects the first transform basis in the horizontal direction from among one or more transform basis candidates.
[0240] (2) When the horizontal size of the coding target block is less than or equal to the first horizontal threshold size, the first transform basis selection unit 1063 selects a fixed transform basis in the horizontal direction as the first transform basis in the horizontal direction.
[0241] (3) When the vertical size of the coding target block is greater than the first vertical threshold size, the first transform basis selection unit 1063 adaptively selects the first transform basis in the vertical direction from among one or more transform basis candidates.
[0242] (4) When the vertical size of the coding target block is less than or equal to the first vertical threshold size, the first transform basis selection unit 1063 selects a fixed transform basis in the vertical direction as the first transform basis in the vertical direction.
[0243] The fixed transform basis in the horizontal direction may be the same as or different from the fixed transform basis in the vertical direction. As the fixed transform basis in the horizontal direction and the vertical direction, for example, a transform basis of type 7 discrete sine transform (DST-VII) can be used.
[0244] The first transform unit 1064 performs a first transform on the residual of the coding target block using the first transform basis selected by the first transform basis selection unit 1063, thereby generating first transform coefficients. Specifically, the first transform unit 1064 performs a first transform in the horizontal direction using the first transform basis in the horizontal direction, and performs a first transform in the vertical direction using the first transform basis in the vertical direction.
[0245] The second transformation execution determination unit 1065 determines whether to perform a second transformation on the first transformation coefficients based on whether the adaptive transformation basis selection mode is valid in the coding target block. Specifically, the second transformation execution determination unit 1065 performs the second transformation when the adaptive transformation basis selection mode is not valid, and determines not to perform the second transformation when the adaptive transformation basis selection mode is valid.
[0246] The second transformation basis selection unit 1066 selects a second transformation basis when it is determined to perform the second transformation. That is, when the adaptive transformation basis selection mode is not valid, the second transformation basis selection unit 1066 selects the second transformation basis. On the contrary, when the adaptive transformation basis selection mode is valid, the second transformation basis selection unit 1066 does not select the second transformation basis. That is, when the adaptive transformation basis selection mode is valid, the second transformation basis selection unit 1066 skips the selection of the second transformation basis.
[0247] When it is determined to perform the second transformation, the second transformation unit 1067 uses the second transformation basis selected by the second transformation basis selection unit 1066 to transform the first transformation coefficients. That is, when the adaptive transformation basis selection mode is not valid, the second transformation unit 1067 generates second transformation coefficients by performing a second transformation on the first transformation coefficients using the second transformation basis. On the contrary, when the adaptive transformation basis selection mode is valid, the second transformation unit 1067 does not perform a second transformation on the first transformation coefficients. That is, when the adaptive transformation basis selection mode is valid, the second transformation unit 1067 skips the second transformation.
[0248] [Internal Structure of Inverse Transformation Unit of Encoding Device]
[0249] Next, refer to Figure 11B The internal structure of the inverse transformation unit 114 of the encoding device 100 according to this embodiment will be described. Figure 11B It is a block diagram showing the internal structure of the inverse transformation unit 114 of the encoding device 100 according to the first mode of Embodiment 1.
[0250] As Figure 11B shown, the inverse transformation unit 114 according to this embodiment includes a second inverse transformation basis selection unit 1141, a second inverse transformation unit 1142, a first inverse transformation basis selection unit 1143, and a first inverse transformation unit 1144.
[0251] When the adaptive transformation basis selection mode is not valid in the coding target block, the second inverse transformation basis selection unit 1141 selects the inverse transformation basis of the second transformation basis selected by the second transformation basis selection unit 1066 as the second inverse transformation basis.
[0252] When the adaptive transform basis selection mode is not valid for the coded object block, the second inverse transform unit 1142 performs a second inverse transform on the inverse quantized coefficients by using the second inverse transform basis selected by the second inverse transform basis selection unit 1141, and generates second inverse transform coefficients. The inverse quantized coefficients mean the coefficients inverse quantized by the inverse quantization unit 112.
[0253] The first inverse transform basis selection unit 1143 selects the inverse transform basis of the first transform basis selected by the first transform basis selection unit 1063 as the first inverse transform basis.
[0254] When the adaptive transform basis selection mode is not valid for the coded object block, the first inverse transform unit 1144 performs a first inverse transform on the second inverse transform coefficients by using the first inverse transform basis, and reconstructs the residual of the coded object block. On the other hand, when the adaptive transform basis selection mode is valid for the coded object block, the residual of the coded object block is reconstructed by performing a first inverse transform on the inverse quantized coefficients by using the first inverse transform basis.
[0255] [Processing of the transform unit and quantization unit of the coding device]
[0256] Next, with reference to the processing of the quantization unit 108 Figure 12A The processing of the transform unit 106 configured as above will be described. Figure 12A It is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the coding device 100 according to the first mode of Embodiment 1.
[0257] The transform mode determination unit 1061 determines whether the adaptive transform basis selection mode is valid for the coded object block (S101).
[0258] When the adaptive transform basis selection mode is not valid (No in S101), the first transform basis selection unit 1063 selects one basic transform basis as the first transform basis in the horizontal direction and the vertical direction (S102).
[0259] When the adaptive transform basis selection mode is valid (Yes in S101), the size determination unit 1062 determines whether the transform size in the horizontal direction exceeds a certain range (S103). That is, the size determination unit 1062 determines whether the horizontal size of the coded object block is greater than the first horizontal threshold size.
[0260] When the transform size in the horizontal direction exceeds a certain range (Yes in S103), the first transform basis selection unit 1063 selects the transform basis in the horizontal direction from among a plurality of adaptive transform bases as the first transform basis in the horizontal direction (S104).
[0261] When the transformed size in the horizontal direction is within a certain range (No in S103), the first transform basis selection unit 1063 selects a fixed transform basis as the first transform basis in the horizontal direction (S105).
[0262] Next, the size determination unit 1062 determines whether the transformed size in the vertical direction exceeds a certain range (S106). That is, the size determination unit 1062 determines whether the vertical size of the coding target block is greater than the first vertical threshold size.
[0263] When the transformed size in the vertical direction exceeds a certain range (Yes in S106), the first transform basis selection unit 1063 selects an adaptive transform basis from a plurality of adaptive transform bases as the first transform basis in the vertical direction (S107).
[0264] When the transformed size in the vertical direction is within a certain range (No in S106), the first transform basis selection unit 1063 selects a fixed transform basis as the first transform basis in the vertical direction (S108).
[0265] In addition, the selection order of the transform bases in the horizontal direction and the vertical direction can be the order of the horizontal direction and the vertical direction, or the reverse order. In addition, the transform basis in the horizontal direction and the transform basis in the vertical direction can also be selected simultaneously.
[0266] The first transform unit 1064 uses the first transform basis selected in step S102, S107 or step S108 to perform the first transform on the prediction residual to generate the first transform coefficient (S109).
[0267] Next, the second transform execution determination unit 1065 determines whether to perform the second transform on the first transform coefficient (S110). Here, the second transform execution determination unit 1065 determines whether to perform the second transform based on whether the adaptive transform basis selection mode is valid in the coding target block.
[0268] When the adaptive transform basis selection mode is valid (Yes in S110), neither the selection of the second transform basis nor the second transform is performed, and the quantization unit 108 generates quantization coefficients by performing quantization of the first transform coefficient (S113). That is, steps S111 and S112 Figure 12A are skipped.
[0269] In the case where the adaptive transform basis selection mode is not valid (No in S110), the second transform basis selection unit 1066 selects a second transform basis from one or more candidates for the second transform basis (S111). Then, the second transform unit 1067 generates second transform coefficients by performing a second transform on the first transform coefficients using the selected second transform basis (S112). Then, the quantization unit 108 generates quantized coefficients by performing quantization of the second transform coefficients (S113).
[0270] As the above-described basic transform basis, a prescribed transform basis can be used. In this case, it is possible to determine whether the adaptive transform basis selection mode is valid based on whether the first transform bases in the horizontal direction and the vertical direction are the prescribed transform basis. In addition, the prescribed transform basis can be one transform basis or two or more transform bases.
[0271] In addition, in the case where the second transform is not performed (skipped), the second transform may not be performed, or a transform equivalent to not performing the transform may be performed as the second transform. In the former case, information indicating that the second transform is not performed may also be encoded in the bitstream. In addition, in the latter case, information indicating a transform equivalent to not performing the transform may also be encoded in the bitstream. Hereinafter, the same can be said for the process of skipping each transform.
[0272] In addition, Figure 12A The steps shown and the order of the steps, etc. are examples and are not limited thereto. For example, as Figure 12B shown, it is also possible to combine Figure 12A the determination of the adaptive transform basis selection mode (S101) and the determination of the implementation of the second transform (S110). Figure 12B is a flowchart showing a modification example of the processing of the transform unit 106 and the quantization unit 108 of the encoding apparatus 100 according to the first mode of Embodiment 1. Figure 12B The flowchart of Figure 12A is a flowchart substantially equivalent to the flowchart of
[0273] In Figure 12B the determination of the implementation of the second transform (S110) is deleted, and the first transform (S109) is divided into two (S109A, S109B). In this case, the transform unit 106 of the encoding apparatus 100 may not include the second transform implementation determination unit 1065.
[0274] Regarding the selection of the second inverse transform basis and the second inverse transform, and the selection of the first inverse transform basis and the first inverse transform in the inverse transform unit 114, it is only necessary to perform them according to the transform of the Figure 12A transform unit 106, so the description and illustration are omitted.
[0275] In addition, the first transformation may be a frequency transformation that can adaptively select a transformation basis, such as the EMT described in Non-Patent Document 2, a frequency transformation that switches the transformation basis under certain conditions, or other general transformations. For example, instead of selecting the first transformation basis, a fixed transformation basis may be set. Additionally, a first transformation basis equivalent to not performing the first transformation may be used. Furthermore, in the first transformation, identification information indicating which of the adaptive transformation basis selection mode and the transformation basis fixed mode using a fixed basic transformation basis (e.g., the transformation basis of type 2 discrete cosine transform (DCT-II)) is effective can be used to select either of the two modes. In this case, based on the identification information, it is also possible to determine which of the adaptive transformation basis selection mode and the transformation basis fixed mode is effective in the coding target block. For example, in the EMT described in Non-Patent Document 2, since there is identification information (emt_cu_flag) indicating whether the adaptive transformation basis selection mode is effective in units such as CU (Coding Unit), this identification information can be used to determine whether the adaptive transformation basis selection mode is effective in the coding target block.
[0276] In addition, the second transformation may be a secondary transformation process such as the NSST described in Non-Patent Document 2, a transformation that switches the transformation basis under certain conditions, or other general transformations. For example, instead of selecting the second transformation basis, a fixed transformation basis may be set. Additionally, a second transformation basis equivalent to not performing the second transformation may be used. Additionally, NSST may be a frequency space transformation after DCT or DST. For example, a KLT (Karhunen Loveve Transform) or a basis equivalent to KLT representing the transformation coefficients of DCT or DST obtained offline, or a HyGT (Hypercube-Givens Transform) represented by a combination of rotation transformations may be used.
[0277] Furthermore, this processing can be applied to either the luminance signal or the chrominance signal. As long as the input signal is in RGB format, it can also be applied to each of the R, G, and B signals. Moreover, in the luminance signal and the chrominance signal, the selectable bases in the first transformation or the second transformation may be different. For example, since the frequency band of the luminance signal is wider than that of the chrominance signal, in order to perform an optimal transformation, in the first transformation or the second transformation of the luminance signal, more types of bases may be used as selection candidates than for the chrominance. Additionally, this processing can be applied to either intra-frame processing or inter-frame processing.
[0278] [Effects, etc.]
[0279] In the first transformation (primary transformation) and the second transformation (secondary transformation) described in Non-Patent Document 2, the best transformation basis or transformation coefficients (filters) are selected to achieve the overall best coding efficiency. Therefore, in order to search for the best combination of candidates for the transformation basis and transformation coefficients (filters) used in the first transformation and the second transformation, the first transformation and the second transformation need to be tried multiple times. That is, in the transformation method described in Non-Patent Document 2, it is necessary to calculate the evaluation value for all combinations of candidates for the transformation basis of the first transformation and candidates for the transformation basis of the second transformation, and select the combination with the smallest evaluation value. Therefore, the present inventor has found that in the transformation method described in Non-Patent Document 2, the processing amount becomes huge.
[0280] Therefore, the encoding device 100 according to this method does not always perform both the first transformation and the second transformation, but skips the second transformation based on whether the adaptive transformation basis selection mode is effective. Thus, the encoding device 100 can reduce the number of combinations of candidates for the transformation basis of the first transformation and candidates for the transformation basis of the second transformation, and can reduce the processing amount.
[0281] In addition, according to the encoding device 100 according to this method, it is possible to limit the candidates for the first transformation basis based on the conditions of the transformation sizes in the horizontal and vertical directions. Thus, it is possible to reduce the processing amount for searching for the best first transformation basis by trial. In addition, based on conditions such as the basis selected in the first transformation basis, it is possible to reduce the processing for searching for the best second transformation basis by trial. In addition, it is possible to reduce the processing amount for the trial of the combination of the first transformation and the second transformation.
[0282] As an example, as the basic transformation basis, the transformation basis of DCT-II can be used. DCT-II is highly likely to be adopted when the residual shape is flat or randomly adopted. For example, if DCT-II is used as the first transformation basis, there is a possibility that the effect of the second transformation is improved because of the tendency of the concentration toward the low frequency to increase. On the other hand, in transformation bases other than DCT-II, high-frequency components are likely to remain, and there is a possibility that the effect of the second transformation is reduced.
[0283] In addition, as an example, as the fixed transformation basis selected when the transformation size is within a certain range, the DST-VII transformation basis can be used. In particular, in the case of intra-frame processing, DST-VII has a tendency to be selected with a very high probability when the residual shape is inclined and the size is small.
[0284] In addition, as the basic transformation basis, it is not limited to one specified transformation basis, and multiple specified transformation bases can also be used.
[0285] In addition, whether to perform the selection of the second transformation basis and the second transformation can be switched according to the transformation size. In addition, the candidates for the second transformation basis can also be switched according to the transformation size.
[0286] Alternatively, it can also be configured to switch whether to perform the second transformation only based on whether the adaptive transformation basis selection mode is valid, without switching the first transformation basis corresponding to the transformation size. That is, in Figure 12A , steps S103, S105, S106, and step S108 can also be deleted. Here, it can also be determined whether the adaptive transformation basis selection mode is valid based on the identification information indicating the use of the mode or the type of the first transformation basis.
[0287] Similarly, it can also be configured not to switch whether to perform the second transformation based on whether the adaptive transformation basis selection mode is valid, but only to switch the first transformation basis corresponding to the transformation size. That is, in Figure 12A , step S110 can also be deleted.
[0288] In addition, regardless of whether the adaptive transformation basis selection mode is valid, the selection of the second transformation basis and the second transformation can also not be skipped. Also, regardless of the method of selecting the first transformation basis, when the adaptive transformation basis selection mode is not valid, the selection of the second transformation basis and the second transformation can be performed, and when the adaptive transformation basis selection mode is valid, the selection of the second transformation basis and the second transformation can be skipped.
[0289] In addition, as the threshold values of the specific transformation sizes in the horizontal or vertical directions for selecting a candidate as the first transformation basis or a fixed transformation basis from multiple adaptive transformation bases (i.e., the first horizontal threshold size and the first vertical threshold size), 4, 8, 16, 32, or 64 pixels, etc. can also be used.
[0290] [Combination with other methods]
[0291] This method can also be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. can also be combined with other methods for implementation.
[0292] (The second method of Embodiment 1)
[0293] Next, the second method of Embodiment 1 will be described. In this method, an example of encoding various signals related to the first transformation and the second transformation in the first method will be described. Hereinafter, with the differences from the first method as the center, this method will be specifically described with reference to the drawings.
[0294] In addition, since the internal structures of the transformation unit 106 and the inverse transformation unit 114 of the encoding device 100 related to this method are the same as those in the first method, the illustration thereof is omitted.
[0295] [Processing of the transformation unit, quantization unit, and entropy encoding unit of the encoding device]
[0296] Refer to Figure 13A and Figure 13B , and the processing of the transformation unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding device 100 according to this embodiment will be described. Figure 13A is a flowchart showing the processing of the transformation unit 106 and quantization unit 108 of the encoding device 100 according to the second embodiment of Embodiment 1. Figure 13B is a flowchart showing the processing of the entropy encoding unit 110 of the encoding device 100 according to the second embodiment of Embodiment 1. In Figure 13A and Figure 13B , the same reference numerals are used for the processing shared with the first embodiment, and the description thereof is omitted.
[0297] After quantization is performed (S113), the entropy encoding unit 110 encodes the adaptive transform basis selection mode signal (S201). The adaptive transform basis selection mode signal is an example of the identification information of the adaptive transform basis selection mode.
[0298] Then, if the adaptive transform basis selection mode is valid ("Yes" in S202), when the transform size in the horizontal direction exceeds a certain range ("Yes" in S203), the entropy encoding unit 110 encodes the first basis selection signal in the horizontal direction (S204). On the other hand, when the transform size in the horizontal direction is within a certain range ("No" in S203), the entropy encoding unit 110 does not encode the first basis selection signal in the horizontal direction. Further, when the transform size in the vertical direction exceeds a certain range ("Yes" in S205), the entropy encoding unit 110 encodes the first basis selection signal in the vertical direction (S206). On the other hand, when the transform size in the vertical direction is within a certain range ("No" in S205), the entropy encoding unit 110 does not encode the first basis selection signal in the vertical direction.
[0299] When the adaptive transform basis selection mode is not valid ("No" in S202), the encoding of the first basis selection signal (S204, S206) is skipped.
[0300] Next, the entropy encoding unit 110 encodes the quantization coefficients (S207).
[0301] Here, when the adaptive transform basis selection mode is not valid ("No" in S208), the entropy encoding unit 110 encodes the second basis selection signal (S209). On the other hand, when the adaptive transform basis selection mode is valid ("Yes" in S208), the encoding of the second basis selection signal (S209) is skipped.
[0302] In addition, the order of each coding can be preset, and various signals can be coded in a manner different from the order of the above-mentioned coding.
[0303] In the case of not performing (skipping) the second transformation, a signal indicating non - performance of the second transformation can be coded, and a signal selecting the second basis equivalent to no transformation can also be coded.
[0304] [Syntax]
[0305] Here, the syntax in this method is described. Figure 14 It represents a specific example of the syntax in the second method of Embodiment 1.
[0306] In Figure 14 , for example, in the case where an adaptive transform basis selection mode signal (emt_cu_flag) is set (line 4), if the transform size in the horizontal direction (horizontal_tu_size) is greater than the first horizontal threshold size (horizontal_tu_size_th) (line 5), then the first basis selection signal in the horizontal direction (emt_horizontal_tridx) is coded (line 6). In addition, if the transform size in the vertical direction (vertical_tu_size) is greater than the first vertical threshold size (vertical_tu_size_th) (line 11), then the first basis selection signal in the vertical direction (emt_vertical_tridx) is coded (line 12). Under other conditions (lines 8 and 14), the coding of the first basis selection signal is skipped (lines 9 and 15).
[0307] In addition, in the case where the adaptive transform basis selection mode signal (emt_cu_flag) is not set (line 19), the second basis selection signal (secondary_tridx) is coded (line 20). On the contrary, in the case where the adaptive transform basis selection mode signal (emt_cu_flag) is set (line 22), the coding of the second basis selection signal (secondary_tridx) is skipped (line 23).
[0308] [Specific Examples of Transform Bases and Coding Signals]
[0309] Next, specific examples of transform bases and coding signals are described. Figure 15 It represents a specific example of the presence or absence of coding of the transform basis and signal used in the second method of Embodiment 1.
[0310] In Figure 15In the case where the adaptive transform basis selection mode is not valid, regardless of the size of the coding block, the transform basis of DCT-II is used as the first transform basis in the horizontal and vertical directions. That is, the transform basis of DCT-II is used as the basic transform basis. In addition, while the second transform (ON) is being performed, the second basis selection signal (secondary_tridx) indicating the second transform basis used in this second transform is encoded into the bitstream.
[0311] On the other hand, in the case where the adaptive transform basis selection mode is valid, depending on the horizontal size H and the vertical size V of the coding target block, a combination of DST-VII transform bases and other transform bases (index0 to index3) is used as candidates for the first transform basis in the horizontal and vertical directions. Additionally, regardless of the size of the coding target block, the second transform is not performed (OFF). Furthermore, although the second basis selection signal (secondary_tridx) is not encoded either, the adaptive transform basis selection mode signal (emt_cu_flag) is encoded into the bitstream. In addition, in the case where the horizontal size H of the coding target block is greater than 4 pixels, the first basis selection signal in the horizontal direction (emt_horizontal_tridx) is encoded into the bitstream. In addition, in the case where the vertical size V of the coding target block is greater than 4 pixels, the first basis selection signal in the vertical direction (emt_vertical_tridx) is encoded in the bitstream.
[0312] For example, in the case where the horizontal size H is 4 pixels or less and the vertical size V is 4 pixels or less, only the DST-VII transform basis is used as a candidate for the first transform basis in the horizontal and vertical directions. At this time, the first basis selection signals in the horizontal and vertical directions (emt_horizontal_tridx and emt_vertical_tridx) are not encoded.
[0313] In addition, for example, in the case where the horizontal size H is 4 pixels or less and the vertical size V is greater than 4 pixels, only the DST-VII transform basis is used as a candidate for the first transform basis in the horizontal direction, and the DST-VII transform basis and other transform bases are used as candidates for the first transform basis in the vertical direction. At this time, although the first basis selection signal in the horizontal direction (emt_horizontal_tridx) is not encoded, the first basis selection signal in the vertical direction (emt_vertical_tridx) is encoded.
[0314] In addition, for example, when the horizontal size H is greater than 4 pixels and the vertical size V is 4 pixels or less, the DST-VII transform basis and other transform bases are used as candidates for the first transform basis in the horizontal direction, and only the DST-VII transform basis is used as a candidate for the first transform basis in the vertical direction. At this time, the first basis selection signal (emt_horizontal_tridx) in the horizontal direction is encoded, but the first basis selection signal (emt_vertical_tridx) in the vertical direction is not encoded.
[0315] In addition, for example, when the horizontal size H is greater than 4 pixels and the vertical size V is greater than 4 pixels, the DST-VII transform basis and other transform bases are used as candidates for the first transform basis in both the horizontal and vertical directions. At this time, the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) in the horizontal and vertical directions are encoded.
[0316] [Effects, etc.]
[0317] As described above, according to the encoding device 100 of the present embodiment, it is possible to encode the information representing the first transform basis (the first basis selection signal) only when the adaptive transform basis selection mode is valid and the transform size exceeds a certain range, and there is a possibility of reducing the amount of code required for signaling the first transform basis. In addition, it is possible to encode the information representing the second transform basis (the second basis selection signal) only when the adaptive transform basis selection mode is not valid, and there is a possibility of reducing the amount of code required for signaling the second transform basis. In addition, by encoding the information for determining whether to skip the second transform (such as the adaptive transform basis selection mode signal) before the information representing the second transform basis, it is possible to determine at the time of decoding whether the information representing the second transform basis is encoded.
[0318] Alternatively, the encoding of the second basis selection signal may be always performed regardless of the adaptive transform basis selection mode. In addition, regardless of the transform size, if it is the adaptive transform basis selection mode, the encoding of the first basis selection signal may be always performed. In addition, it is possible to independently determine the presence or absence of encoding of the first basis selection signal using the size in the horizontal direction and the size in the vertical direction, or it is possible to make a combined determination.
[0319] [Combination with other embodiments]
[0320] This embodiment may also be implemented in combination with at least a part of other embodiments of the present invention. In addition, a part of the processing described in the flowchart of this embodiment, a part of the structure of the device, a part of the grammar, etc. may be combined with other embodiments for implementation.
[0321] (Third Mode of Embodiment 1)
[0322] Next, the third mode of Embodiment 1 will be described. In this mode, the difference from the above-described first mode is that when the adaptive transform basis selection mode is not valid, different basic transform bases are used as the first transform basis according to the size of the coding target block. Hereinafter, this mode will be specifically described with reference to the drawings, centering on the points different from the first mode and the second mode.
[0323] In addition, the internal structures of the transform unit 106 and the inverse transform unit 114 of the coding device 100 in this mode are the same as those in the first mode, so the illustration is omitted.
[0324] [Processing of the Transform Unit and Quantization Unit of the Coding Device]
[0325] Refer to Figure 16 The processing of the transform unit 106 and the quantization unit 108 of the coding device 100 in this mode will be described. Figure 16 is a flowchart showing the processing of the transform unit 106 and the quantization unit 108 of the coding device 100 according to the third mode of Embodiment 1. In Figure 16 the same reference numerals are given to the processing shared with the first mode and the description thereof is omitted.
[0326] When the adaptive transform basis selection mode is not valid (No in S101), the size determination unit 1062 determines whether the transform size is within a certain range (S301). That is, the size determination unit 1062 determines whether the size of the coding target block is equal to or less than the second threshold size. For example, the size determination unit 1062 determines whether the product of the horizontal size and the vertical size of the coding target block is equal to or less than the threshold, thereby determining whether the size of the coding target block is equal to or less than the second threshold size.
[0327] Here, when the transform size is within a certain range (Yes in S301), the first transform basis selection unit 1063 selects the second basic transform basis as the first transform basis in the horizontal direction and the vertical direction (S302). On the other hand, when the transform size exceeds a certain range (No in S301), the first transform basis selection unit 1063 selects the first basic transform basis as the first transform basis in the horizontal direction and the vertical direction (S303).
[0328] As an example, as the first basic transform basis, the transform basis of DCT-II can be used, and as the second basic transform basis, the DST-VII transform basis can be used.
[0329] In addition, the basic transform basis can also be selected from among candidates of a plurality of basic transform bases.
[0330] In addition, regardless of whether the adaptive transform basis selection mode is valid, the selection of the second transform basis and the second transform may not be skipped. Further, regardless of the method for selecting the first transform basis, when the adaptive transform basis selection mode is not valid, the selection of the second transform basis and the second transform may be performed, and when the adaptive transform basis selection mode is valid, the selection of the second transform basis and the second transform may be skipped.
[0331] In addition, when the adaptive transform basis selection mode is not valid, as the second threshold size for selecting one of the first basic transform basis and the second basic transform basis, for example, pixel sizes such as 4x4, 4x8, 8x4, 8x8 can be used. Further, as the transform size for comparison with the threshold, the product of the horizontal size and the vertical size of the coding target block can be used as in this method, or the horizontal size and the vertical size can be used separately.
[0332] In addition, when the adaptive transform basis selection mode is not valid, if the product of the horizontal size and the vertical size is within a certain range, the second basic transform basis may be selected as the first transform basis in the horizontal direction and the vertical direction, and the selection of the second transform basis and the second transform may be skipped.
[0333] [Effects, etc.]
[0334] As described above, according to the coding device 100 related to this method, when the adaptive transform basis selection mode is not valid, the first transform basis can be switched between the first basic transform basis and the second basic transform basis according to the transform size. Therefore, the first transform can be performed using the first transform basis corresponding to the transform size, and a reduction in the amount of code can be achieved.
[0335] [Combination with other methods]
[0336] This method may be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. may be combined with other methods for implementation.
[0337] (Fourth method of Embodiment 1)
[0338] Next, the fourth method of Embodiment 1 will be described. In this method, an example of the coding of various signals for the first transform and the second transform related to the third method will be described. Hereinafter, this method will be specifically described with reference to the drawings, centering on the points different from the first to third methods.
[0339] In addition, the internal structures of the transform unit 106 and the inverse transform unit 114 of the coding device 100 related to this method are the same as those of the first method, and thus the illustration thereof is omitted.
[0340] [Processing of the transformation unit, quantization unit, and entropy encoding unit of the encoding device]
[0341] Refer to Figure 17A and Figure 17B , and the processing of the transformation unit 106, quantization unit 108, and entropy encoding unit 110 of the encoding device 100 according to this method will be described. Figure 17A is a flowchart showing the processing of the transformation unit 106 and quantization unit 108 of the encoding device 100 according to the fourth method of Embodiment 1. Figure 17B is a flowchart showing the processing of the entropy encoding unit 110 of the encoding device 100 according to the fourth method of Embodiment 1. In Figure 17A and Figure 17B , for the processing shared with any of the first to third methods, the same reference numerals are used and the description is omitted.
[0342] After quantization is performed (S113), the entropy encoding unit 110 determines whether to skip the encoding of the adaptive transform basis selection mode signal (S401). For example, when any one of the following conditions (A) and (B) is satisfied, the entropy encoding unit 110 determines to skip the encoding of the adaptive transform basis selection mode signal; otherwise, the entropy encoding unit 110 determines not to skip the encoding of the adaptive transform basis selection mode signal.
[0343] (A) The adaptive transform basis selection mode is not valid.
[0344] (B) The adaptive transform basis selection mode is valid and all of the following conditions (B1) to (B4) are satisfied.
[0345] (B1) The transform size is equal to or smaller than the second threshold size W1xH1 used in step S301.
[0346] (B2) The transform size in the horizontal direction is equal to or smaller than the first horizontal threshold size W2 used in step S103.
[0347] (B3) The transform size in the vertical direction is equal to or smaller than the first vertical threshold size H2 used in step S106.
[0348] (B4) The second basic transform basis and the fixed transform bases in the horizontal and vertical directions are the same transform basis.
[0349] As a specific example, when the second threshold size W1xH1 is 4×4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and both the second basic transform basis and the fixed transform basis are the DST-VII transform basis, if the transform size is 4×4 pixels or smaller, the entropy encoding unit 110 determines to skip the encoding of the adaptive transform basis selection mode signal.
[0350] On the contrary, when neither of the above conditions (A) and (B) is satisfied, the entropy encoding unit 110 determines not to skip the encoding of the adaptive transform basis selection mode signal.
[0351] Here, when it is determined to skip the encoding of the adaptive transform basis selection mode signal (Yes in S401), the entropy encoding unit 110 skips steps S201 to S206 and encodes the quantization coefficients (S207). On the other hand, when it is determined not to skip the encoding of the adaptive transform basis selection mode signal (No in S401), the entropy encoding unit 110, in the same manner as the second method, performs steps S201 to S206 and then encodes the quantization coefficients (S207).
[0352] In addition, the encoding order of each encoding can be preset, and various signals can be encoded in a manner different from the above encoding order.
[0353] [Syntax]
[0354] Here, the syntax in this method will be described. Figure 18 Shows a specific example of the syntax in the fourth method related to Embodiment 1.
[0355] In Figure 18 , for example, when skipping the encoding of the adaptive transform basis selection mode signal (line 20), the encoding of the adaptive transform basis selection mode signal (emt_cu_flag) and the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) is skipped (line 21). Here, when the transform size in the horizontal direction (horizontal_tu_size) is less than or equal to the first horizontal threshold size (horizontal_tu_size_th) and the transform size in the vertical direction (vertical_tu_size) is less than or equal to the first vertical threshold size (vertical_tu_size_th), the encoding of the adaptive transform basis selection mode signal is skipped. When not skipping the encoding of the adaptive transform basis selection mode signal (lines 3 - 4), the adaptive transform basis selection mode signal (emt_cu_flag) is encoded (line 5), and, in the same manner as the second method, the first basis selection signals (emt_horizontal_tridx and emt_vertical_tridx) are encoded as needed (lines 7 - 16).
[0356] In addition, when skipping the encoding of the adaptive transform basis selection mode signal, the selection of the second transform basis and the second transform can also be skipped.
[0357] [Specific Examples of Transform Bases and Encoded Signals]
[0358] Next, specific examples of the transform basis and the encoded signal will be described. Figure 19 Specific examples showing the presence or absence of encoding of the transform basis and the signal used in the fourth mode of Embodiment 1 are presented. In Figure 19 , when both the horizontal size and the vertical size of the encoding target block are 4 pixels or less, the presence or absence of the transform basis and encoding is different from Figure 15 . Centering on the differences from Figure 15 , Figure 19 will be described.
[0359] In Figure 19 , when the adaptive transform basis selection mode is not effective, if both the horizontal size H and the vertical size V of the encoding target block are 4 pixels or less, the DST-VII transform basis instead of the DCT-II transform basis is used as the first transform basis in the horizontal and vertical directions.
[0360] Furthermore, when the adaptive transform basis selection mode is effective, if both the horizontal size H and the vertical size V of the encoding target block are 4 pixels or less, the adaptive transform basis selection mode signal (emt_cu_flag) is not encoded.
[0361] [Effects, etc.]
[0362] As described above, according to the encoding apparatus 100 of the present mode, when the conditions for skipping the encoding of the adaptive transform basis selection mode signal are satisfied, it is possible to omit all encoding of the adaptive transform basis selection mode signal and the first basis selection signal, and there is a possibility of reducing the amount of code.
[0363] [Combination with Other Modes]
[0364] This mode may also be implemented in combination with at least a part of other modes in the present invention. In addition, a part of the processing described in the flowchart of this mode, a part of the structure of the apparatus, a part of the syntax, etc. may be combined with other modes for implementation.
[0365] (Fifth Mode of Embodiment 1)
[0366] Next, the fifth mode of Embodiment 1 will be described. In this mode, a decoding apparatus will be described. In addition, the decoding apparatus of this mode corresponds to the encoding apparatus of the first mode described above. That is, the decoding apparatus of this mode can decode the bitstream encoded by the encoding apparatus of the first mode described above. Hereinafter, this mode will be specifically described with reference to the drawings.
[0367] [Internal Configuration of the Transform Unit and Inverse Transform Unit of the Decoding Apparatus]
[0368] First, the internal structure of the inverse transformation unit 206 of the decoding device 200 according to this method will be described. Figure 20 It is a block diagram showing the internal structure of the inverse transformation unit 206 of the decoding device 200 according to the fifth method of Embodiment 1.
[0369] As Figure 20 shown, the inverse transformation unit 206 according to this method includes a second inverse transformation execution determination unit 2061, a second inverse transformation basis selection unit 2062, a second inverse transformation unit 2063, a transformation mode determination unit 2064, a size determination unit 2065, a first inverse transformation basis selection unit 2066, and a first inverse transformation unit 2067.
[0370] The second inverse transformation execution determination unit 2061 determines whether to perform the second inverse transformation on the inverse quantization coefficients of the decoding target block based on whether the adaptive transformation basis selection mode is valid in the decoding target block. Specifically, the second inverse transformation execution determination unit 2061 performs the second inverse transformation when the adaptive transformation basis selection mode is not valid, and determines not to perform the second inverse transformation when the adaptive transformation basis selection mode is valid.
[0371] When it is determined to perform the second inverse transformation, the second inverse transformation basis selection unit 2062 selects the second inverse transformation basis. Specifically, when the adaptive transformation basis selection mode is not valid, the second inverse transformation basis selection unit 2062 obtains the second basis selection signal 2062S representing the second inverse transformation basis decoded from the bit stream by the entropy decoding unit 202. Then, the second inverse transformation basis selection unit 2062 selects the second inverse transformation basis based on the second basis selection signal 2062S. On the contrary, when the adaptive transformation basis selection mode is valid, the second inverse transformation basis selection unit 2062 does not select the second inverse transformation basis. That is, when the adaptive transformation basis selection mode is valid, the second inverse transformation basis selection unit 2062 skips the selection of the second inverse transformation basis.
[0372] When it is determined to perform the second inverse transformation, the second inverse transformation unit 2063 performs the second inverse transformation on the inverse quantization coefficients of the decoding target block using the second inverse transformation basis selected by the second inverse transformation basis selection unit 2062. That is, when the adaptive transformation basis selection mode is not valid, the second inverse transformation unit 2063 generates the second inverse transformation coefficients by performing the second inverse transformation on the inverse quantization coefficients using the second inverse transformation basis. On the contrary, when the adaptive transformation basis selection mode is valid, the second inverse transformation unit 2063 does not perform the second inverse transformation on the inverse quantization coefficients. That is, when the adaptive transformation basis selection mode is valid, the second inverse transformation unit 2063 skips the second inverse transformation.
[0373] The transformation mode determination unit 2064 determines whether the adaptive transformation basis selection mode is valid in the decoding target block. The determination of whether the adaptive transformation basis selection mode is valid is performed based on the first basis selection signal 2066S or the adaptive transformation basis selection mode signal 2064S decoded from the bitstream by the entropy decoding unit 202. That is, the determination is performed based on the identification information of the first inverse transformation basis or the adaptive transformation basis selection mode.
[0374] The size determination unit 2065 determines whether the horizontal size of the decoding target block exceeds the first horizontal threshold size. In addition, the size determination unit 1062 determines whether the vertical size of the decoding target block exceeds the first vertical threshold size. The determination of the horizontal size and the vertical size is performed based on the size signal 2065S decoded from the bitstream by the entropy decoding unit 202.
[0375] The first inverse transformation basis selection unit 2066 selects the first inverse transformation basis. Specifically, when the adaptive transformation basis selection mode is not valid, the first inverse transformation basis selection unit 2066 selects one basic transformation basis as the first inverse transformation basis in the horizontal direction and the vertical direction. In addition, when the adaptive transformation basis selection mode is valid, the first inverse transformation basis selection unit 2066 selects the first inverse transformation basis in the horizontal direction and the vertical direction as follows in (1) to (4).
[0376] (1) When the horizontal size of the decoding target block is larger than the first horizontal threshold size, the first inverse transformation basis selection unit 2066 obtains the first basis selection signal 2066S representing the first inverse transformation basis, which is decoded from the bitstream by the entropy decoding unit 202. Then, the first inverse transformation basis selection unit 2066 selects the first inverse transformation basis in the horizontal direction based on the first basis selection signal 2066S.
[0377] (2) When the horizontal size of the decoding target block is less than or equal to the first horizontal threshold size, the first inverse transformation basis selection unit 2066 selects a fixed transformation basis in the horizontal direction as the first inverse transformation basis in the horizontal direction.
[0378] (3) When the vertical size of the decoding target block is larger than the first vertical threshold size, the first inverse transformation basis selection unit 2066 obtains the first basis selection signal 2066S. Then, the first inverse transformation basis selection unit 2066 selects the first inverse transformation basis in the vertical direction based on the first basis selection signal 2066S.
[0379] (4) When the vertical size of the decoding target block is less than or equal to the first vertical threshold size, the first inverse transformation basis selection unit 2066 selects a fixed transformation basis in the vertical direction as the first inverse transformation basis in the vertical direction.
[0380] The first inverse transform unit 2067 performs a first inverse transform on the inverse quantized coefficients of the block to be decoded by using the first inverse transform basis selected by the first inverse transform basis selection unit 2066, thereby restoring the residual of the block to be decoded. Specifically, the first inverse transform unit 2067 performs a horizontal first inverse transform by using the first inverse transform basis in the horizontal direction and performs a vertical first inverse transform by using the first inverse transform basis in the vertical direction.
[0381] [Processing of the inverse quantization unit and the inverse transform unit of the decoding device]
[0382] Next, refer to it together with the processing of the inverse quantization unit 204 Figure 21 Explain the processing of the inverse transform unit 206 configured as described above. Figure 21 It is a flowchart showing the processing of the inverse quantization unit 204 and the inverse transform unit 206 of the decoding device 200 according to the fifth mode of Embodiment 1.
[0383] The inverse quantization unit 204 generates inverse quantized coefficients by performing inverse quantization on the quantized coefficients of the block to be decoded decoded by the entropy decoding unit 202 (S501).
[0384] The second inverse transform execution determination unit 2061 determines whether to perform a second inverse transform on the inverse quantized coefficients (S502). Here, the second inverse transform execution determination unit 2061 determines whether to perform the second inverse transform based on whether the adaptive transform basis selection mode is valid in the block to be decoded.
[0385] Here, when the adaptive transform basis selection mode is valid ("Yes" in S502), neither the selection of the second inverse transform basis nor the second inverse transform is performed. That is, steps S503 and S504 are skipped.
[0386] On the other hand, when the adaptive transform basis selection mode is not valid ("No" in S502), the second inverse transform basis selection unit 2062 selects the second inverse transform basis based on the second basis selection signal 2062S (S503). Further, the second inverse transform unit 2063 performs a second inverse transform on the inverse quantized coefficients by using the selected second inverse transform basis (S504).
[0387] Next, the transform mode determination unit 2064 determines whether the adaptive transform basis selection mode is valid in the block to be decoded (S505). For example, the transform mode determination unit 2064 determines whether the adaptive transform basis selection mode is valid based on the adaptive transform basis selection mode signal 2064S.
[0388] When the adaptive transform basis selection mode is not valid (No in S505), the first inverse transform basis selection unit 2066 selects one basic transform basis as the first inverse transform basis in the horizontal and vertical directions (S512). On the other hand, when the adaptive transform basis selection mode is valid (Yes in S505), the size determination unit 2065 determines whether the transform size in the horizontal direction exceeds a certain range (S506). That is, the size determination unit 2065 determines whether the horizontal size of the decoding target block is larger than the first horizontal threshold size.
[0389] When the transform size in the horizontal direction exceeds a certain range (Yes in S506), the first inverse transform basis selection unit 2066 selects the transform basis in the horizontal direction from among a plurality of adaptive transform bases as the first inverse transform basis in the horizontal direction (S507). On the other hand, when the transform size in the horizontal direction is within a certain range (No in S506), the first inverse transform basis selection unit 2066 selects a fixed transform basis as the first inverse transform basis in the horizontal direction (S508).
[0390] The size determination unit 2065 determines whether the transform size in the vertical direction exceeds a certain range (S509). That is, the size determination unit 2065 determines whether the vertical size of the decoding target block is larger than the first vertical threshold size.
[0391] When the transform size in the vertical direction exceeds a certain range (Yes in S509), the first inverse transform basis selection unit 2066 selects a transform basis from among a plurality of adaptive transform bases as the first inverse transform basis in the vertical direction (S510). When the transform size in the vertical direction is within a certain range (No in S509), the first inverse transform basis selection unit 2066 selects a fixed transform basis as the first inverse transform basis in the vertical direction (S511).
[0392] The first inverse transform unit 2067 performs a first inverse transform on the inverse quantization coefficients or the second inverse transform coefficients using the first inverse transform basis selected as above, and restores the residual of the decoding target block (S513).
[0393] In addition, the selection order of the inverse transform bases in the horizontal and vertical directions may be the order of the horizontal and vertical directions, or the reverse order thereof. In addition, the inverse transform bases in the horizontal and vertical directions may be selected simultaneously.
[0394] Furthermore, selecting an inverse transform basis in the decoding device 200 means decoding information indicating the basis used in the inverse transform included in the encoded bit stream, determining the inverse transform basis based on the decoded information, or determining the uniquely represented inverse transform basis based on information such as the intra prediction mode, the size of the decoding target block, or the basis in the first inverse transform.
[0395] In addition, a decoding method matching the Figure 12A or Figure 12B encoding method of the first method shown can also be used.
[0396] [Effects, etc.]
[0397] As described above, according to the decoding device 200 related to this method, the same effects as those of the encoding device 100 related to the first method can be achieved.
[0398] [Combination with other methods]
[0399] This method can also be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. can also be combined with other methods for implementation.
[0400] (The sixth method of Embodiment 1)
[0401] Next, the sixth method of Embodiment 1 will be described. In this method, an example of decoding various signals related to the first transformation and the second transformation in the fifth method will be described. In addition, the decoding device related to this method corresponds to the encoding device of the above-mentioned second method. Hereinafter, with the differences from the fifth method as the center, this method will be specifically described with reference to the drawings.
[0402] In addition, the internal structure of the inverse transformation unit 206 of the decoding device 200 related to this method is the same as that of the fifth method, so the illustration is omitted.
[0403] [Processing of the entropy decoding unit, inverse quantization unit, and inverse transformation unit of the decoding device]
[0404] Refer to Figure 22A and Figure 22B , and the processing of the entropy decoding unit 202, inverse quantization unit 204, and inverse transformation unit 206 of the decoding device 200 related to this method will be described. In Figure 22A and Figure 22B , the same reference numerals are used for the processing shared with the fifth method and the description is omitted.
[0405] First, the entropy decoding unit 202 decodes the adaptive transform basis selection mode signal from the bit stream (S601). Then, based on the adaptive transform basis selection mode signal, the transform mode determination unit 2064 determines whether the adaptive transform basis selection mode is valid in the block to be decoded (S602).
[0406] If the adaptive transform basis selection mode is valid ("Yes" in S602), then when the transform size in the horizontal direction exceeds a certain range ("Yes" in S603), the entropy decoding unit 202 decodes the first basis selection signal in the horizontal direction from the bitstream (S604). On the other hand, when the transform size in the horizontal direction is within a certain range ("No" in S603), the entropy decoding unit 202 does not decode the first basis selection signal in the horizontal direction. Further, when the transform size in the vertical direction exceeds a certain range ("Yes" in S605), the entropy decoding unit 202 decodes the first basis selection signal in the vertical direction from the bitstream (S606). On the other hand, when the transform size in the vertical direction is within a certain range ("No" in S605), the entropy decoding unit 202 does not decode the first basis selection signal in the vertical direction.
[0407] When the adaptive transform basis selection mode is not valid ("No" in S602), the decoding of the first basis selection signal is skipped (S604, S606).
[0408] Next, the entropy decoding unit 202 decodes the quantization coefficients (S607).
[0409] Here, when the adaptive transform basis selection mode is not valid ("No" in S608), the entropy decoding unit 202 decodes the second basis selection signal from within the bitstream (S609). On the other hand, when the adaptive transform basis selection mode is valid ("Yes" in S608), the decoding of the second basis selection signal is skipped (S609).
[0410] In addition, it is also possible to preset the respective decoding orders in accordance with the encoding method, and decode various signals in a manner different from the above decoding order. Further, when the second inverse transform is not performed (skipped), the entropy decoding unit 202 can decode the signal indicating the non - performance of the second inverse transform from the bitstream, or can decode the signal for selecting the second inverse transform basis equivalent to no transform from the bitstream.
[0411] In addition, it is also possible to adopt Figure 13A , Figure 13B and Figure 14 the decoding method matching the encoding method of the second method shown.
[0412] [Effects, etc.]
[0413] As described above, according to the decoding device 200 related to this mode, the same effects as those of the encoding device 100 related to the second mode can be achieved.
[0414] [Combinations with Other Modes]
[0415] This method can also be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. can be combined with other methods for implementation.
[0416] (The seventh mode of Embodiment 1)
[0417] Next, the seventh mode of Embodiment 1 will be described. In this mode, the difference from the above-described fifth mode is that when the adaptive transform basis selection mode is not valid, different basic transform bases are used as the first inverse transform basis according to the size of the coding target block. In addition, the decoding device related to this mode corresponds to the coding device of the above-described third mode. Hereinafter, this mode will be specifically described with reference to the drawings centering on the points different from the fifth mode and the sixth mode.
[0418] In addition, since the internal structure of the inverse transform unit 206 of the decoding device 200 related to this mode is the same as that of the fifth mode, the illustration thereof is omitted.
[0419] [Processing of the inverse quantization unit and the inverse transform unit of the decoding device]
[0420] Refer to Figure 23 The processing of the transform unit 106 and the quantization unit 108 of the coding device 100 related to this mode will be described. Figure 23 is a flowchart showing the processing of the inverse quantization unit 204 and the inverse transform unit 206 of the decoding device 200 related to the seventh mode of Embodiment 1. In Figure 23 the same reference numerals are given to the processing shared with the fifth mode and the description thereof is omitted.
[0421] When the adaptive transform basis selection mode is not valid (\"No\" in S505), the size determination unit 2065 determines whether the transform size is within a certain range (S701). That is, the size determination unit 2065 determines whether the horizontal size and the vertical size of the decoding target block are equal to or less than the second threshold size. Specifically, the size determination unit 2065 determines, for example, whether the product of the horizontal size and the vertical size of the decoding target block is equal to or less than the threshold.
[0422] Here, when the transform size is within a certain range (\"Yes\" in S701), the first inverse transform basis selection unit 2066 selects the second basic transform basis as the first inverse transform basis in the horizontal direction and the vertical direction (S702). On the other hand, when the transform size exceeds a certain range (\"No\" in S701), the first inverse transform basis selection unit 2066 selects the first basic transform basis as the first inverse transform basis in the horizontal direction and the vertical direction (S703).
[0423] In addition, it is also possible to adopt matching with Figure 16Decoding method for the encoding method of the third mode shown.
[0424] [Effects, etc.]
[0425] As described above, according to the decoding device 200 related to this mode, the same effects as those of the encoding device 100 related to the third mode can be achieved.
[0426] [Combination with other modes]
[0427] This mode can also be implemented in combination with at least a part of other modes in the present invention. In addition, a part of the processing described in the flowchart of this mode, a part of the structure of the device, a part of the grammar, etc. can also be combined with other modes for implementation.
[0428] (Eighth mode of Embodiment 1)
[0429] Next, the eighth mode of Embodiment 1 will be described. In this mode, an example of decoding various signals related to the first transformation and the second transformation in the seventh mode will be described. In addition, the decoding device related to this mode corresponds to the encoding device of the fourth mode described above. Hereinafter, this mode will be specifically described with reference to the drawings centering on the differences from the fifth to seventh modes.
[0430] In addition, since the internal structure of the inverse transformation unit 206 of the decoding device 200 related to this mode is the same as that of the fifth mode, the illustration thereof is omitted.
[0431] [Processing of the entropy decoding unit, inverse quantization unit, and inverse transformation unit of the decoding device]
[0432] Refer to Figure 24A and Figure 24B to describe the processing of the entropy decoding unit 202, inverse quantization unit 204, and inverse transformation unit 206 of the decoding device 200 related to this mode. In Figure 24A and Figure 24B , for the processing shared with any of the fifth to seventh modes, the same reference numerals are given and the description thereof is omitted.
[0433] The entropy decoding unit 202 determines whether to skip the decoding of the adaptive transform basis selection mode signal (S801). For example, when any one of the following conditions (A) and (B) is satisfied, the entropy decoding unit 202 determines to skip the decoding of the adaptive transform basis selection mode signal; otherwise, the entropy decoding unit 202 determines not to skip the decoding of the adaptive transform basis selection mode signal.
[0434] (A) The adaptive transform basis selection mode is not valid.
[0435] (B) The adaptive transform basis selection mode is valid and all of the following conditions (B1) to (B4) are satisfied.
[0436] (B1) The transformed size is equal to or smaller than the second threshold size W1xH1 used in step S701.
[0437] (B2) The transformed size in the horizontal direction is equal to or smaller than the first horizontal threshold size W2 used in step S506.
[0438] (B3) The transformed size in the vertical direction is equal to or smaller than the first vertical threshold size H2 used in step S509.
[0439] (B4) The second basic transformation basis and the fixed transformation bases in the horizontal and vertical directions are the same transformation basis.
[0440] As a specific example, when the second threshold size W1xH1 is 4×4 pixels, the first horizontal threshold size W2 is 4 pixels, the first vertical threshold size H2 is 4 pixels, and both the second basic transformation basis and the fixed transformation basis are DST-VII transformation bases, if the transformed size is equal to or smaller than 4×4 pixels, the entropy decoding unit 202 determines to skip the decoding of the adaptive transformation basis selection mode signal.
[0441] Conversely, when any of the conditions in the above (A) and (B) is not satisfied, the entropy decoding unit 202 determines not to skip the decoding of the adaptive transformation basis selection mode signal.
[0442] Here, when it is determined to skip the decoding of the adaptive transformation basis selection mode signal (Yes in S801), the entropy decoding unit 202 skips steps S601 to S606 and decodes the quantization coefficients (S607). On the other hand, when it is determined not to skip the decoding of the adaptive transformation basis selection mode signal (No in S801), the entropy decoding unit 202, in the same manner as in the sixth mode, executes steps S601 to S606 and then decodes the quantization coefficients (S207).
[0443] In addition, it is also possible to adopt a decoding method that matches the Figure 17A , Figure 17B and Figure 18 shown in the fourth encoding method.
[0444] [Effects, etc.]
[0445] As described above, according to the decoding device 200 related to this mode, the same effects as those of the encoding device 100 related to the fourth mode can be achieved.
[0446] [Combination with other modes]
[0447] This method can also be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. can be combined with other methods for implementation.
[0448] (Modification examples of each method of Embodiment 1)
[0449] In addition, signals indicating whether a part or all of the processing described in any one of the first to eighth methods is effective can be encoded and decoded. Such signals can be encoded in units of CU (Coding Unit) or CTU (Coding Tree Unit), or can be encoded in units equivalent to SPS (Sequence Parameter Set), PPS (Picture Parameter Set) or slices of the H.265 / HEVC standard.
[0450] Based on the picture type (I, P, B), slice type (I, P, B), transform size (4×4 pixels, 8x8 pixels or others), number of non-zero coefficients, quantization parameter, Temporal_id (layer of hierarchical coding), or any combination thereof, the selection of the first transform basis and the first transform can be skipped, and the selection of the second transform basis and the second transform can also be skipped.
[0451] When the encoding device related to the first to fourth methods performs the above actions, the decoding device related to the fifth to eighth methods also performs corresponding actions. For example, when the encoding device encodes information indicating whether the process of skipping the first transform or the second transform is effective, the decoding device decodes the information and determines whether the first transform or the second transform is effective and whether the information indicating the first transform or the second transform is encoded.
[0452] (Embodiment 2)
[0453] In the above embodiments, each functional block can generally be implemented by an MPU, a memory, etc. In addition, the processing of each functional block is generally realized by a program execution unit such as a processor reading and executing software (program) recorded in a recording medium such as a ROM. This software can be distributed by downloading or the like, or can be recorded in a recording medium such as a semiconductor memory for distribution. In addition, of course, each functional block can also be implemented by hardware (special circuit).
[0454] In addition, the processing described in each embodiment can be implemented by centralized processing using a single device (system), or can also be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, either centralized processing or distributed processing can be performed.
[0455] The form of the present invention is not limited to the above embodiments, and various modifications can be made, and they are also included in the scope of the form of the present invention.
[0456] Furthermore, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in the above embodiments and a system using the same are described here. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Regarding other structures in the system, they can be appropriately changed according to circumstances.
[0457] [Usage Example]
[0458] Figure 25 It is a diagram showing the overall structure of a content supply system ex100 that realizes a content distribution service. The provision area of the communication service is divided into desired sizes, and base stations ex106, ex107, ex108, ex109, and ex110 as fixed wireless stations are respectively provided in each unit.
[0459] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. Some of the above elements in the content supply system ex100 can also be combined and connected. Each device can also be directly or indirectly connected to each other via a telephone network or short-range wireless without going through the base stations ex106 to ex110 as fixed wireless stations. In addition, a streaming media server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smart phone ex115 via the Internet ex101 and the like. In addition, the streaming media server ex103 is connected to terminals in a hotspot in an airplane ex117 via a satellite ex116.
[0460] In addition, a wireless access point or a hotspot etc. can be used in place of the base stations ex106 to ex110. Moreover, the streaming media server ex103 can be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or can be directly connected to the aircraft ex117 without going through the satellite ex116.
[0461] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a mobile phone, or a PHS (Personal Handyphone System) etc. corresponding to the modes of mobile communication systems generally referred to as 2G, 3G, 3.9G, 4G, and those to be referred to as 5G in the future.
[0462] The home appliance ex118 is a refrigerator or a device included in a household fuel cell cogeneration system etc.
[0463] In the content supply system ex100, a terminal having a photographing function is connected to the streaming media server ex103 via the base station ex106 etc., and thereby live distribution etc. can be performed. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smart phone ex115, and the terminal in the aircraft ex117 etc.) performs the encoding process described in the above respective embodiments on the still image or moving image content photographed by the user using the terminal, multiplexes the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and sends the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.
[0464] On the other hand, the streaming media server ex103 performs stream distribution on the content data sent by a requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or a terminal in the aircraft ex117 etc. that can decode the data after the above encoding process. Each device that receives the distributed data performs a decoding process on the received data and reproduces it. That is, each device functions as an image decoding device according to one aspect of the present invention.
[0465] [Distributed processing]
[0466] In addition, the streaming media server ex103 can also be multiple servers or multiple computers, which disperse the processing or recording of data and then distribute it. For example, the streaming media server ex103 can also be implemented by a CDN (Content Delivery Network), and content distribution is achieved through a network that connects many edge servers scattered around the world to each other. In a CDN, physically closer edge servers are dynamically allocated according to the client. Moreover, by caching and distributing content to this edge server, latency can be reduced. In addition, in the case of a certain error or when the communication state changes due to an increase in traffic, etc., the processing can be dispersed among multiple edge servers, or the distribution entity can be switched to another edge server, or a part of the network with a fault can be bypassed to continue the distribution, so high-speed and stable distribution can be achieved.
[0467] In addition, not limited to the decentralized processing of the distribution itself, the encoding process of the captured data can be performed by each terminal, on the server side, or they can share the work. As an example, usually two processing loops are performed in the encoding process. In the first loop, the complexity or code amount of the image in units of frames or scenes is detected. In addition, in the second loop, a process is performed to improve the encoding efficiency while maintaining the image quality. For example, by having the terminal perform the first encoding process and the server side that receives the content perform the second encoding process, the processing load in each terminal can be reduced while improving the quality and efficiency of the content. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.
[0468] As another example, a camera ex113, etc., extracts feature amounts from an image, compresses the data regarding the feature amounts as metadata, and sends it to the server. The server, for example, judges the importance of the target based on the feature amounts and switches the quantization accuracy, etc., to perform compression corresponding to the meaning of the image. Feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction during re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed by the server.
[0469] As other examples, in a stadium, a shopping mall, a factory, etc., there are cases where there are multiple pieces of image data obtained by multiple terminals photographing substantially the same scene. In this case, the multiple terminals that have performed the photographing, and other terminals and servers that have not performed the photographing as needed are used. For example, encoding processing is separately assigned and distributed in units of GOP (Group of Picture), picture units, or tile units obtained by dividing a picture, etc. Thereby, latency can be reduced and real-time performance can be better achieved.
[0470] In addition, since the multiple pieces of image data are of substantially the same scene, the server can also manage and / or instruct to cross-reference the image data photographed by each terminal. Alternatively, the server can receive the encoded data from each terminal and change the cross-reference relationship between the multiple data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.
[0471] In addition, the server can also transcode the image data by changing the encoding method and then distribute the image data. For example, the server can change the encoding method of the MPEG type to the VP type, or can change H.264 to H.265.
[0472] In this way, the encoding process can be performed by the terminal or one or more servers. Therefore, the following descriptions use "server" or "terminal", etc. as the main body of the processing, but a part or all of the processing performed by the server can also be performed by the terminal, and a part or all of the processing performed by the terminal can also be performed by the server. In addition, the same applies to the decoding process regarding these.
[0473] [3D, Multi-angle]
[0474] In recent years, the cases of merging and using images or videos of different scenes photographed by multiple terminals such as multiple cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or of the same scene photographed from different angles have increased. The images photographed by each terminal are merged based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the images are the same, etc.
[0475] The server not only encodes two-dimensional moving images, but can also automatically or at a user-specified time encode still images based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative positional relationship between the shooting terminals, it can not only generate three-dimensional shapes of the same scene from images taken from different angles based on two-dimensional moving images, but also encode three-dimensional data generated from point clouds separately. Additionally, based on the results of identifying or tracking people or objects using the three-dimensional data, the server can select or reconstruct images from the images captured by multiple terminals and generate images to be sent to the receiving terminal.
[0476] In this way, users can not only arbitrarily select each image corresponding to each shooting terminal to view the scene, but also view the content of the image taken from an arbitrary viewpoint cut from the three-dimensional data reconstructed using multiple images or videos. Furthermore, similar to images, sound can also be collected from multiple different angles, and the server multiplexes and sends the sound from a specific angle or space with the image according to the image.
[0477] In addition, in recent years, content that correlates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has been spreading. In the case of VR images, the server separately creates viewpoint images for the right eye and the left eye, and can perform encoding that allows reference between the images at each viewpoint through Multi-View Coding (MVC) or the like, or can encode them as different streams without mutual reference. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.
[0478] In the case of AR images, it can also be that the server overlays virtual object information in the virtual space on the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device obtains or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates the overlay data by smoothly connecting them. Or, it can also be that the decoding device sends the movement of the user's viewpoint to the server in addition to entrusting the virtual object information, and the server creates the overlay data according to the three-dimensional data held in the server, matches the received movement of the viewpoint, encodes the overlay data, and distributes it to the decoding device. Additionally, the overlay data has an α value representing transparency in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted through this part. Or, the server can also set the RGB value of a specified value as the background like chroma keying and generate data with the part other than the target set as the background color.
[0479] Similarly, the decoding process of the distributed data can be performed by each terminal acting as a client, on the server side, or can be shared between them. As an example, a certain terminal may first send a reception request to the server, and another terminal may receive the content corresponding to the request and perform the decoding process, and then send the decoded signal to a device with a display. By dispersing the processing regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with better image quality. In addition, as another example, a large-sized image data may be received by a TV or the like, and a personal terminal of the viewer may decode and display a part of the area such as tiles after the picture is segmented. Thus, while making the overall image shared, it is possible to confirm one's own responsible area or the area that one wants to confirm in more detail at hand.
[0480] In addition, it is envisioned that in the future, in a situation where multiple wireless communications, whether short-range, medium-range, or long-range, can be used both indoors and outdoors, using a distribution system standard such as MPEG-DASH, content can be received seamlessly while appropriately switching data for the connected communication. Thus, the user can not only use their own terminal but also freely select a decoding device or a display device such as a monitor installed indoors and outdoors to perform real-time switching. In addition, based on their own location information and the like, it is possible to switch the decoding terminal and the display terminal for decoding. Thus, it is also possible to display map information on a part of the wall or the ground of a building next to a displayable device while moving towards the destination. In addition, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time or replicated in an edge server of the content distribution service, it is possible to switch the bit rate of the received data.
[0481] [Scalable Coding]
[0482] Regarding the content switching, a scalable stream that is compressed and encoded using the motion image encoding method shown in Figure 26 each of the above embodiments will be described. For the server, it may have multiple streams with the same content but different qualities as separate streams, or may be structured to switch content using the characteristics of a temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, by the decoding side determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication band, the decoding side can freely switch between decoding low-resolution content and high-resolution content. For example, when wanting to view the subsequent video that was viewed on a smartphone ex115 while moving at home using a device such as an Internet TV, the device only needs to decode the same stream to different layers, so the burden on the server side can be reduced.
[0483] Furthermore, in addition to the structure in which images are encoded layer by layer as described above and a hierarchical structure with an enhancement layer above the base layer is implemented, the enhancement layer may include meta-information such as statistical information of the image. On the decoding side, the image of the base layer is super-resolved based on the meta-information to generate high-quality content. The super-resolution can be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least-squares operation used in the super-resolution process, etc.
[0484] Alternatively, the image may be segmented into tiles or the like according to the meaning of objects or the like in the image, and on the decoding side, only a part of the region is decoded by selecting the tiles to be decoded. In addition, by saving the attributes of the object (person, car, ball, etc.) and the position in the image (coordinate position in the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including the object. For example, as Figure 27 shown, a data storage structure different from the pixel data, such as the SEI message in HEVC, is used to store the meta-information. The meta-information represents, for example, the position, size, or color of the main object.
[0485] In addition, the meta-information may be stored in units composed of multiple images, such as a stream, sequence, or random access unit. As a result, the decoding side can obtain the time when a specific person appears in the video, and by matching with the information of the image unit, can determine the image in which the object exists and the position of the object in the image.
[0486] [Optimization of Web Page]
[0487] Figure 28 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. Figure 29 is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As Figure 28 and Figure 29 shown, there are cases where a web page includes multiple linked images that are links to image content, and the visible manner thereof varies depending on the viewing device. When multiple linked images can be seen on the screen, before the user explicitly selects a linked image, or before the linked image approaches near the center of the screen or the whole of the linked image enters the screen, the display device (decoding device) displays the still image or I picture that each content has as a linked image, or displays an image such as a gif animation using multiple still images or I pictures, or only receives the base layer and decodes and displays the image.
[0488] When the user selects a linked image, the display device decodes the base layer with the highest priority. Additionally, if there is information indicating scalable content in the HTML that constitutes the web page, the display device may also decode up to the enhancement layer. Furthermore, in cases where real-time performance needs to be ensured or the communication bandwidth is extremely tight before selection, the display device can reduce the delay between the decoding time and the display time of the leading picture (from the start of content decoding to the start of display) by decoding and displaying only the pictures with forward references (I pictures, P pictures, B pictures with only forward references). Additionally, the display device may also forcibly ignore the reference relationships of the pictures and roughly decode all B pictures and P pictures as forward references, and as more pictures are received over time, perform normal decoding.
[0489] [Autonomous driving]
[0490] Furthermore, in cases where still images or video data such as two-dimensional or three-dimensional map information are transmitted and received for the autonomous driving or driving assistance of a vehicle, the receiving terminal may also receive information such as weather or construction information as meta-information in addition to the image data belonging to one or more layers, and decode them by establishing correspondences. Additionally, the meta-information may belong to a layer or may only be multiplexed with the image data.
[0491] In this case, since vehicles, drones, or airplanes including the receiving terminal are moving, the receiving terminal can switch between base stations ex106 to ex110 to perform seamless reception and decoding by sending the location information of the receiving terminal at the time of the reception request. Additionally, the receiving terminal can dynamically switch the degree to which meta-information is received or the degree to which map information is updated according to the user's selection, the user's situation, or the status of the communication bandwidth.
[0492] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.
[0493] [Distribution of personal content]
[0494] Furthermore, in the content supply system ex100, not only high-quality and long-duration content provided by video distribution providers but also unicast or multicast distribution of low-quality and short-duration content provided by individuals can be performed. Additionally, it is conceivable that such personal content will increase in the future. In order to make personal content better, the server may also perform an encoding process after an editing process. This can be achieved, for example, through the following structure.
[0495] During shooting in real time or after accumulating the shots, the server performs recognition processing such as shooting error, scene search, meaning analysis, and object detection based on the original image or the encoded data. And, based on the recognition results, the server manually or automatically performs editing such as correcting focus deviation or camera shake, deleting scenes with low importance such as scenes with lower brightness than other pictures or out-of-focus scenes, emphasizing the edges of the object, or changing the color tone. Based on the editing results, the server encodes the edited data. In addition, it is known that the viewing rate decreases if the shooting time is too long. The server can also automatically limit scenes with low importance as described above, as well as scenes with little movement, etc. based on the image processing results according to the shooting time, so as to become content within a specific time range. Or, the server can also generate a summary and encode it based on the result of the meaning analysis of the scene.
[0496] In addition, in the case of personal content, there are situations where the original state contains content that infringes on copyright, the author's personality rights, or portrait rights, etc., and there are also inconvenient situations for individuals such as the sharing range exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or at home, etc. into out-of-focus images. In addition, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user designates the person or background area for which the image is to be processed, and the server performs processing such as replacing the designated area with another image or blurring the focus. In the case of a person, the image of the face part can be replaced while tracking the person in the moving image.
[0497] In addition, the real-time viewing of personal content with a small data volume has strong requirements. So although it also depends on the bandwidth, the decoding device first receives and decodes and reproduces the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period, and when the reproduction is looped 2 or more times, reproduce the high-quality image including the enhancement layer. In this way, if it is a scalable encoded stream, it is possible to provide an experience where the moving image is rough at the stage of not being selected or just starting to watch, but the stream gradually becomes smooth and the image quality improves. In addition to scalable encoding, the same experience can also be provided when the first reproduced rough stream and the second stream encoded with reference to the first moving image form one stream.
[0498] [Other usage examples]
[0499] In addition, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSIex500 can be either a single chip or a structure composed of multiple chips. Additionally, software for motion image encoding or decoding can be loaded into a certain recording medium (such as CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, in the case where a smartphone ex115 is equipped with a camera, it is also possible to transmit the motion image data obtained by this camera. The motion image data at this time is the data after being encoded by the LSIex500 possessed by the smartphone ex115.
[0500] In addition, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. In the case where the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software, and then performs content acquisition and reproduction.
[0501] In addition, it is not limited to the content supply system ex100 via the Internet ex101, and at least one of the motion image encoding device (image encoding device) or motion image decoding device (image decoding device) of the above-described embodiments can also be assembled in a digital broadcast system. Since the broadcast radio wave uses a satellite, etc. to carry multiplexed data that multiplexes video and audio for transmission and reception, there is a difference in being suitable for multicast compared to the structure of the content supply system ex100 that is easy for unicast, but the same application can be made for the encoding process and decoding process.
[0502] [Hardware Structure]
[0503] Figure 30 It is a diagram showing the smartphone ex115. In addition, Figure 31FIG. 0 is a diagram showing a structural example of a smart phone ex115. The smart phone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from a base station ex110, a camera unit ex465 capable of capturing images and still images, and a display unit ex458 for displaying the images captured by the camera unit ex465 and decoding data such as the images received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the captured images or still images, recorded sound, received images or still images, encoded or decoded data such as e-mails, or a slot unit ex464 as an interface unit with a SIM ex468 for identifying a user and performing authentication for accessing various data represented by a network. In addition, an external memory may be used instead of the memory unit ex467.
[0504] In addition, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc. is connected to a power supply circuit unit ex461, an operation input control unit ex462, an image signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, a sound signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via a bus ex470.
[0505] If the power key is turned on by a user operation, the power supply circuit unit ex461 starts the smart phone ex115 to an operable state by supplying power to each unit from a battery pack.
[0506] The smart phone ex115 performs processes such as calls and data communication under the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal processing unit ex454 converts the voice signal collected by the voice input unit ex456 into a digital voice signal, performs spread spectrum processing on it using the modulation / demodulation unit ex452, and after the digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, it is transmitted via the antenna ex450. In addition, the received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, inverse spread spectrum processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. During data communication, text, still image, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body, etc., and the sending and receiving processes are performed in the same way. In the data communication mode, when sending video, still image, or video and voice, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method shown in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. In addition, the voice signal processing unit ex454 encodes the voice signal collected by the voice input unit ex456 during the process of shooting video, still image, etc. by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a specified manner, and modulation processing and conversion processing are performed by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and it is transmitted via the antenna ex450.
[0507] When receiving an image attached to an email or a chat tool, or an image linked on a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bitstream of video data and a bitstream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in the above embodiments, and displays the image or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. In addition, the audio signal processing unit ex454 decodes the audio signal and outputs the sound from the audio output unit ex457. In addition, since real-time streaming media is becoming popular, depending on the user's situation, there may be occasions where the reproduction of sound is inappropriate in society. Therefore, as an initial value, a structure that does not reproduce the audio signal but only reproduces the video data is preferred. It is also possible to reproduce the sound synchronously only when the user performs an operation such as clicking on the video data.
[0508] In addition, here, the smart phone ex115 is taken as an example for explanation. However, as the terminal, in addition to the transceiver terminal having both an encoder and a decoder, three installation forms can be considered: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. Furthermore, in the digital broadcast system, it is assumed that the multiplexed data in which audio data and the like are multiplexed in the video data is received and transmitted for explanation. However, in the multiplexed data, character data related to the video and the like can be multiplexed in addition to the audio data, and it is also possible to receive or transmit the video data itself instead of the multiplexed data.
[0509] In addition, it is assumed that the main control unit ex460 including the CPU controls the encoding or decoding process for explanation. However, in many cases, the terminal has a GPU. Therefore, it is also possible to configure a structure in which the performance of the GPU is utilized to process a larger area together through a memory shared by the CPU and the GPU or a memory that manages addresses in a shared manner. As a result, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization are performed not by the CPU but by the GPU together in units of pictures or the like.
[0510] Industrial applicability
[0511] The present invention can be applied to, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, or a digital video camera, etc.
[0512] Reference Numeral Explanation
[0513] 100 Encoding device
[0514] 102 Splitting unit
[0515] 104 Subtraction unit
[0516] 106 Transformation unit
[0517] 108 Quantization unit
[0518] 110 Entropy encoding unit
[0519] 112, 204 Inverse quantization unit
[0520] 114, 206 Inverse transformation unit
[0521] 116, 208 Addition unit
[0522] 118, 210 Block memory
[0523] 120, 212 Loop filtering unit
[0524] 122, 214 Frame memory
[0525] 124, 216 Intra prediction unit
[0526] 126, 218 Inter prediction unit
[0527] 128, 220 Prediction control unit
[0528] 200 Decoding device
[0529] 202 Entropy decoding unit
[0530] 1061, 2064 Transformation mode determination unit
[0531] 1062, 2065 Size determination unit
[0532] 1063 First transformation basis selection unit
[0533] 1064 First transformation unit
[0534] 1065 Second transformation execution determination unit
[0535] 1066 Second transformation basis selection unit
[0536] 1067 Second transformation unit
[0537] 1141, 2062 Second Inverse Transformation Basis Selection Unit
[0538] 1142, 2063 Second Inverse Transformation Unit
[0539] 1143, 2066 First Inverse Transformation Basis Selection Unit
[0540] 1144, 2067 First Inverse Transformation Unit
[0541] 2061 Second Inverse Transformation Execution Determination Unit
[0542] 2062S Second Basis Selection Signal
[0543] 2064S Adaptive Transformation Basis Selection Mode Signal
[0544] 2065S Dimension Signal
[0545] 2066S First Basis Selection Signal
Claims
1. A coding method, wherein: Determine whether the adaptive transform basis selection mode is valid, When the above adaptive transform basis selection mode is valid, When the vertical size of the encoding target block is larger than the threshold size, a first transform base is selected from a plurality of transform base candidates as a transform base in the vertical direction. When the vertical size of the encoding target block is smaller than the threshold size, a second transform basis is selected as a transform basis in the vertical direction, and the second transform basis is a fixed transform basis. A first transform is performed on the residual of the encoding target block using the selected transform basis in the vertical direction, thereby generating a first transform coefficient.
2. A decoding method, wherein: Determine whether the adaptive transform basis selection mode is valid, When the above adaptive transform basis selection mode is valid, When the vertical size of the decoding target block is larger than the threshold size, a first inverse transform basis is selected from a plurality of inverse transform basis candidates as an inverse transform basis in the vertical direction, When the vertical size of the decoding target block is smaller than the threshold size, a second inverse transform basis is selected as an inverse transform basis in the vertical direction, wherein the second inverse transform basis is a fixed inverse transform basis. A prediction residual is generated by performing a first inverse transform on the coefficients of the decoding target block using the selected inverse transform basis in the vertical direction.
Citation Information
Patent Citations
Video encoding device, and video decoding device
CN102763416A
Enhanced multiple transforms for prediction residual
CN107211144A