Encoding device, decoding device, and bit stream generating device

The coding device efficiently quantizes rectangular blocks in moving images by converting square quantization matrices to rectangular forms, addressing inefficiencies in conventional methods and improving coding efficiency.

JP2026001099APending Publication Date: 2026-01-06PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025159843
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-03-30
Filing Date
2025-09-26
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Conventional video encoding methods, such as H.265/HEVC, face inefficiencies in quantizing rectangular blocks within moving images, leading to decreased coding efficiency.

Method used

A coding device that utilizes a circuit and memory to perform up-conversion and down-conversion processes on quantization matrices, converting square blocks to rectangular blocks, and encodes only the information related to the quantization matrices, thereby efficiently quantizing rectangular blocks of various shapes.

Benefits of technology

This approach allows for efficient quantization of rectangular blocks in moving images, enhancing coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001099000001_ABST
    Figure 2026001099000001_ABST
Patent Text Reader

Abstract

To efficiently quantize rectangular blocks of various shapes.SOLUTION: The encoding device 100 includes a circuit 160 and a memory 162, wherein the circuit 160 generates a second quantization matrix for rectangular blocks by up-converting a first quantization matrix for square blocks only in a first direction and down-converting the first quantization matrix only in the first direction, quantizes a plurality of transform coefficients of a target block using the second quantization matrix, and encodes the plurality of quantized transform coefficients of the target block, and wherein the circuit 160 copies matrix elements of the first quantization matrix in the first direction in the up-conversion process; In the down-conversion process, the circuit 160 thins out the matrix elements of the first quantization matrix in the first direction and encodes only the information about the first quantization matrix out of the first quantization matrix and the second quantization matrix into a bit stream, where the first direction is the vertical direction.SELECTED DRAWING: Figure 31
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an encoding device that encodes moving images. [Background technology]

[0002] BACKGROUND ART Conventionally, H.265, also known as HEVC (High Efficiency Video Coding), exists as a standard for encoding moving images (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] H.265(ISO / IEC 23008-2 HEVC) / HEVC(High Efficiency Video Coding) Summary of the Invention [Problem to be solved by the invention]

[0004] However, unless the various rectangular blocks contained in a video are efficiently quantized, the coding efficiency decreases.

[0005] Therefore, the present disclosure provides an encoding device and the like that can efficiently quantize rectangular blocks of various shapes included in a moving image. [Means for solving the problem]

[0006] A coding device according to an aspect of the present disclosure is a coding device that codes a moving image by performing quantization, the coding device including a circuit and a memory, wherein the circuit uses the memory to perform up-conversion processing only in a first direction on a first quantization matrix for square blocks and down-conversion processing only in the first direction to generate a second quantization matrix for rectangular blocks, quantize a plurality of transform coefficients of a block to be processed using the second quantization matrix, and encode the quantized plurality of transform coefficients of the block to be processed, wherein the first quantization matrix is ​​a square quantization matrix having a first number of rows and a first number of columns. the second quantization matrix is ​​a rectangular matrix having a second number of rows and a third number of columns different from the second number, the second number or the third number being the same as the first number, in the up-conversion process, the circuit copies matrix elements of the first quantization matrix in the first direction, in the down-conversion process, the circuit thins out matrix elements of the first quantization matrix in the first direction, and encodes only information relating to the first quantization matrix out of the first quantization matrix and the second quantization matrix into a bitstream, the first direction being a vertical direction.

[0007] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]

[0008] An encoding device and the like according to an aspect of the present disclosure can efficiently quantize rectangular blocks of various shapes included in a moving image. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of a coding device according to the first embodiment. [Figure 2]FIG. 2 is a diagram showing an example of block division according to the first embodiment. [Figure 3] FIG. 3 is a table showing the transformation basis functions corresponding to each transformation type. [Figure 4A] FIG. 4A is a diagram showing an example of the shape of a filter used in ALF. [Figure 4B] FIG. 4B is a diagram showing another example of the shape of the filter used in ALF. [Figure 4C] FIG. 4C is a diagram showing another example of the shape of the filter used in ALF. [Figure 5A] FIG. 5A is a diagram showing 67 intra prediction modes in intra prediction. [Figure 5B] FIG. 5B is a flowchart for explaining an outline of the predicted image correction process using the OBMC process. [Figure 5C] FIG. 5C is a conceptual diagram for explaining an outline of the predicted image correction process using the OBMC process. [Figure 5D] FIG. 5D is a diagram showing an example of FRUC. [Figure 6] FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 8] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. [Figure 9A] FIG. 9A is a diagram for explaining derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 9B] FIG. 9B is a diagram for explaining an outline of the motion vector derivation process in the merge mode. [Figure 9C] FIG. 9C is a conceptual diagram for explaining an outline of the DMVR process. [Figure 9D]FIG. 9D is a diagram for explaining an outline of a predicted image generation method using luminance correction processing by LIC processing. [Figure 10] FIG. 10 is a block diagram showing a functional configuration of a decoding device according to the first embodiment. [Figure 11] FIG. 11 is a diagram showing a first example of the flow of encoding processing using a quantization matrix in an encoding device. [Figure 12] FIG. 12 is a diagram showing an example of a decoding process flow using a quantization matrix in a decoding device corresponding to the encoding device described in FIG. [Figure 13] FIG. 13 is a diagram illustrating a first example of generating a quantization matrix for rectangular blocks from a quantization matrix for square blocks in step S102 of FIG. 11 and step S202 of FIG. [Figure 14] FIG. 14 is a diagram for explaining a method for generating the quantization matrix for the rectangular block described in FIG. 13 by down-converting from the quantization matrix for the corresponding square block. [Figure 15] FIG. 15 is a diagram for explaining a second example of generating a quantization matrix for rectangular blocks from a quantization matrix for square blocks in step S102 of FIG. 11 and step S202 of FIG. [Figure 16] FIG. 16 is a diagram for explaining a method for generating the quantization matrix for the rectangular block described in FIG. 15 by up-converting from the quantization matrix for the corresponding square block. [Figure 17] FIG. 17 is a diagram showing a second example of the encoding process flow using a quantization matrix in the encoding device. [Figure 18] FIG. 18 is a diagram showing an example of a decoding process flow using a quantization matrix in a decoding device corresponding to the encoding device described in FIG. [Figure 19] FIG. 19 is a diagram for explaining an example of a quantization matrix corresponding to the size of the valid transform coefficient area for each block size in step S701 in FIG. 17 and step S801 in FIG. [Figure 20]FIG. 20 is a diagram showing a modification of the second example of the encoding process flow using a quantization matrix in the encoding device. [Figure 21] FIG. 21 is a diagram showing an example of a decoding process flow using a quantization matrix in a decoding device corresponding to the encoding device described in FIG. [Figure 22] FIG. 22 is a diagram illustrating a first example of generating a quantization matrix for rectangular blocks from a quantization matrix for square blocks in step S1002 in FIG. 20 and step S1102 in FIG. [Figure 23] FIG. 23 is a diagram for explaining a method for generating the quantization matrix for the rectangular block described in FIG. 22 by down-converting from the quantization matrix for the corresponding square block. [Figure 24] FIG. 24 is a diagram illustrating a second example of generating a quantization matrix for rectangular blocks from a quantization matrix for square blocks in step S1002 in FIG. 20 and step S1102 in FIG. [Figure 25] FIG. 25 is a diagram for explaining a method for generating the quantization matrix for the rectangular block described in FIG. 24 by up-converting from the quantization matrix for the corresponding square block. [Figure 26] FIG. 26 is a diagram showing a third example of the encoding process flow using a quantization matrix in the encoding device. [Figure 27] FIG. 27 is a diagram showing an example of a decoding process flow using a quantization matrix in a decoding device corresponding to the encoding device described in FIG. [Figure 28] Figure 28 is a diagram illustrating an example of a method for generating a quantization matrix using a common method from the values ​​of quantization coefficients of a quantization matrix having only diagonal components for each block size in step S1601 of Figure 26 and step S1701 of Figure 27. [Figure 29]Figure 29 is a diagram illustrating another example of a method for generating a quantization matrix using a common method from the values ​​of quantization coefficients of a quantization matrix having only diagonal components in a processing target block of each block size in step S1601 of Figure 26 and step S1701 of Figure 27. [Figure 30] FIG. 30 is a block diagram showing an example of implementation of an encoding device. [Figure 31] FIG. 31 is a flowchart showing an example of the operation of the encoding device shown in FIG. [Figure 32] FIG. 32 is a block diagram showing an implementation example of a decoding device. [Figure 33] FIG. 33 is a flowchart showing an example of the operation of the decoding device shown in FIG. [Figure 34] FIG. 34 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 35] FIG. 35 is a diagram showing an example of a coding structure for scalable coding. [Figure 36] FIG. 36 is a diagram showing an example of a coding structure for scalable coding. [Figure 37] FIG. 37 is a diagram showing an example of a display screen of a web page. [Figure 38] FIG. 38 is a diagram showing an example of a display screen of a web page. [Figure 39] FIG. 39 is a diagram illustrating an example of a smartphone. [Figure 40] FIG. 40 is a block diagram showing an example of the configuration of a smartphone. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, the embodiments will be specifically described with reference to the drawings.

[0011] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concepts are described as optional components.

[0012] (Embodiment 1) First, an overview of the first embodiment will be described as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure can be applied. However, the first embodiment is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure can be applied, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from the first embodiment.

[0013] When applying the processing and / or configurations described in each aspect of the present disclosure to the first embodiment, for example, any of the following may be performed.

[0014] (1) For the encoding device or decoding device of the first embodiment, among the multiple components constituting the encoding device or decoding device, components corresponding to the components described in each aspect of the present disclosure are replaced with the components described in each aspect of the present disclosure. (2) Any modification, such as addition, replacement, or deletion, of the functions or processes performed by some of the components constituting the encoding device or decoding device of the first embodiment may be made to the encoding device or decoding device, and then components corresponding to the components described in each aspect of the present disclosure may be replaced with the components described in each aspect of the present disclosure. (3) The method implemented by the encoding device or decoding device of the first embodiment may be modified by adding a process and / or replacing or deleting some of the processes included in the method, and then replacing the process described in each aspect of the present disclosure with the process described in each aspect of the present disclosure. (4) Some of the components constituting the encoding device or decoding device of the first embodiment may be implemented in combination with components described in each aspect of the present disclosure, components having some of the functions of the components described in each aspect of the present disclosure, or components performing some of the processing performed by the components described in each aspect of the present disclosure. (5) A component having some of the functions of some of the components constituting the encoding device or decoding device of the first embodiment, or a component that performs some of the processing performed by some of the components constituting the encoding device or decoding device of the first embodiment, is implemented in combination with a component described in each aspect of the present disclosure, a component having some of the functions of the components described in each aspect of the present disclosure, or a component that performs some of the processing performed by the components described in each aspect of the present disclosure. (6) In the method implemented by the encoding device or decoding device of the first embodiment, among the multiple processes included in the method, processes corresponding to the processes described in each aspect of the present disclosure are replaced with the processes described in each aspect of the present disclosure. (7) Some of the processes included in the method implemented by the encoding device or decoding device of the first embodiment may be implemented in combination with the processes described in each aspect of the present disclosure.

[0015] It should be noted that the manner of implementing the processes and / or configurations described in each aspect of the present disclosure is not limited to the above examples. For example, they may be implemented in a device used for a purpose different from the video / image encoding device or video / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented independently. Furthermore, the processes and / or configurations described in different aspects may be implemented in combination.

[0016] [Outline of the encoding device] First, an overview of a coding device according to Embodiment 1 will be described. Fig. 1 is a block diagram showing a functional configuration of a coding device 100 according to Embodiment 1. The coding device 100 is a video / image coding device that codes a video / image on a block-by-block basis.

[0017] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0018] The encoding device 100 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Alternatively, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0019] Each component included in the encoding device 100 will be described below.

[0020] [Divided part] The division unit 102 divides each picture included in the input video into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides a picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks are sometimes called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in this embodiment, there is no need to distinguish between CUs, PUs, and TUs, and some or all of the blocks in a picture may serve as the processing units of CUs, PUs, and TUs.

[0021] Fig. 2 is a diagram showing an example of block division according to embodiment 1. In Fig. 2, solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0022] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0023] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0024] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14 and 15 (binary tree block division).

[0025] The lower-left 64x64 block is divided into four square 32x32 blocks (quadtree block decomposition). Of the four 32x32 blocks, the upper-left and lower-right blocks are further divided. The upper-left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block decomposition). The lower-right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block decomposition). As a result, the lower-left 64x64 block is divided into 16x32 block 16, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.

[0026] The bottom right 64x64 block 23 is not split.

[0027] 2, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. This type of division is sometimes called QTBT (quad-tree plus binary tree) division.

[0028] In Fig. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0029] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (prediction sample) from an original signal (original sample) for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also referred to as a residual) of a block to be coded (hereinafter referred to as a current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0030] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing an image may also be referred to as a sample.

[0031] [Conversion section] The transform unit 106 transforms the spatial domain prediction errors into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction errors.

[0032] The transform unit 106 may adaptively select a transform type from among a plurality of transform types and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0033] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Fig. 3 is a table showing transform basis functions corresponding to each transform type. In Fig. 3, N represents the number of input pixels. Selection of a transform type from among these multiple transform types may depend, for example, on the type of prediction (intra prediction or inter prediction) or the intra prediction mode.

[0034] Information indicating whether EMT or AMT is applied (e.g., referred to as an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0035] Furthermore, the transform unit 106 may retransform the transform coefficients (transform results). Such retransformation may be referred to as an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transform unit 106 performs retransformation for each sub-block (e.g., 4x4 sub-block) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether or not to apply NSST and information regarding the transform matrix used for NSST are signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, picture level, slice level, tile level, or CTU level).

[0036] Here, a separable transformation is a method in which the transformation is performed multiple times by separating the input into directions equal to the number of dimensions, and a non-separable transformation is a method in which, when the input is multidimensional, two or more dimensions are treated as one dimension and the transformation is performed all at once.

[0037] For example, one example of a non-separable transformation is when the input is a 4x4 block, it is treated as a single array with 16 elements, and the transformation process is performed on that array using a 16x16 transformation matrix.

[0038] Similarly, a non-separable transformation is one that treats a 4x4 input block as a single array with 16 elements and then performs multiple Givens rotations on that array (Hypercube Givens Transform).

[0039] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on quantization parameters (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0040] The predetermined order is an order for quantizing / dequantizing the transform coefficients. For example, the predetermined scanning order is defined as an ascending order (low frequency to high frequency) or a descending order (high frequency to low frequency).

[0041] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. In other words, as the value of the quantization parameter increases, the quantization error also increases.

[0042] [Entropy coding section] The entropy coding unit 110 generates a coded signal (coded bit stream) by variable-length coding the quantized coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110, for example, binarizes the quantized coefficients and arithmetically codes the binary signal.

[0043] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. The inverse quantization unit 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0044] [Inverse conversion section] The inverse transform unit 114 restores the prediction error by inverse transforming the transform coefficients that are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients that corresponds to the transform performed by the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0045] Note that the restored prediction error does not match the prediction error calculated by the subtraction unit 104 because information has been lost due to quantization. In other words, the restored prediction error includes a quantization error.

[0046] [Addition section] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0047] [Block Memory] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be coded (hereinafter referred to as a current picture). Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0048] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0049] ALF applies a least squares error filter to remove coding artifacts, for example, for each 2x2 sub-block in the current block, one filter selected from multiple filters based on local gradient direction and activity.

[0050] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The sub-blocks are classified based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (e.g., 15 or 25 classes).

[0051] The gradient direction value D is derived by, for example, comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived by, for example, adding gradients in multiple directions and quantizing the sum.

[0052] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0053] The filter shape used in ALF is, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing several examples of filter shapes used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the filter shape is signaled at the picture level. Note that signaling of the information indicating the filter shape does not need to be limited to the picture level, and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0054] Whether ALF is turned on or off is determined, for example, at the picture level or the CU level. For example, whether ALF is applied to luminance is determined at the CU level, and whether ALF is applied to chrominance is determined at the picture level. Information indicating whether ALF is turned on or off is signaled at the picture level or the CU level. Note that signaling of information indicating whether ALF is turned on or off does not need to be limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0055] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are signaled at the picture level. Note that the signaling of the coefficient sets does not need to be limited to the picture level, but may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0056] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0057] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also referred to as intra-picture prediction) of the current block with reference to blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0058] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0059] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / High-Efficiency Video Coding (HEVC) standard (Non-Patent Document 1).

[0060] The multiple directional prediction modes include, for example, the 33 prediction modes defined in the H.265 / HEVC standard. Note that the multiple directional prediction modes may also include 32 prediction modes in addition to the 33 directions (65 directional prediction modes in total). Fig. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and dashed arrows represent the additional 32 directions.

[0061] Note that a luminance block may be referenced in intra prediction of a chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. This type of intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that references such a luminance block (e.g., called a CCLM mode) may be added as one of the intra prediction modes for the chrominance block.

[0062] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on gradients of reference pixels in the horizontal / vertical directions. Intra prediction involving such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is signaled, for example, at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0063] [Inter prediction section] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also referred to as inter prediction) on the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation on the current block or sub-block within the reference picture. The inter prediction unit 126 then generates an inter prediction signal for the current block or sub-block by performing motion compensation using motion information (e.g., a motion vector) obtained by the motion estimation. The inter prediction unit 126 then outputs the generated inter prediction signal to the prediction control unit 128.

[0064] The motion information used for motion compensation is signaled. For the signaling of the motion vector, a motion vector predictor may be used, i.e., the difference between the motion vector and the motion vector predictor may be signaled.

[0065] Note that an inter-prediction signal may be generated using not only the motion information of the current block obtained by motion estimation, but also the motion information of adjacent blocks. Specifically, an inter-prediction signal may be generated for each sub-block in the current block by weighting and adding a prediction signal based on the motion information obtained by motion estimation and a prediction signal based on the motion information of adjacent blocks. Such inter-prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0066] In such an OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called an OBMC block size) is signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., called an OBMC flag) is signaled at the CU level. Note that the signaling level of this information is not limited to the sequence level and the CU level, and may be other levels (e.g., the picture level, slice level, tile level, CTU level, or sub-block level).

[0067] The OBMC mode will now be described in more detail. Figures 5B and 5C are a flowchart and a conceptual diagram for explaining an outline of the predictive image correction process using the OBMC process.

[0068] First, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to the block to be coded.

[0069] Next, the motion vector (MV_L) of the coded left adjacent block is applied to the block to be coded to obtain a predicted image (Pred_L), and the predicted image is weighted and superimposed with Pred_L to perform the first correction of the predicted image.

[0070] Similarly, the motion vector (MV_U) of the already coded upper adjacent block is applied to the block to be coded to obtain a predicted image (Pred_U), and the predicted image that has been corrected the first time is weighted and overlaid with Pred_U to perform a second correction of the predicted image, which is then used as the final predicted image.

[0071] Although a two-stage correction method using the left adjacent block and the upper adjacent block has been described here, it is also possible to configure a method in which correction is performed more than two times using the right adjacent block or the lower adjacent block.

[0072] The area to be superimposed does not have to be the pixel area of ​​the entire block, but may be only a part of the area near the block boundary.

[0073] Although the process of correcting a predicted image from one reference picture has been described here, the process is similar when correcting a predicted image from multiple reference pictures. After obtaining corrected predicted images from each reference picture, the obtained predicted images are further superimposed to form the final predicted image.

[0074] The target block to be processed may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.

[0075] As a method for determining whether to apply OBMC processing, for example, there is a method using obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, an encoding device determines whether a block to be encoded belongs to an area with complex motion, and if it belongs to an area with complex motion, sets the value of obmc_flag to 1 and performs encoding by applying OBMC processing, and if it does not belong to an area with complex motion, sets the value of obmc_flag to 0 and performs encoding without applying OBMC processing. On the other hand, a decoding device decodes obmc_flag described in a stream, and switches whether to apply OBMC processing depending on the value, and performs decoding.

[0076] Alternatively, the motion information may be derived on the decoding device side without being signaled. For example, a merge mode defined in the H.265 / HEVC standard may be used. Alternatively, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation is performed without using pixel values ​​of the current block.

[0077] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.

[0078] An example of the FRUC process is shown in Figure 5D. First, a list of multiple candidates (which may be the same as the merge list) each having a predicted motion vector is generated by referring to the motion vectors of coded blocks spatially or temporally adjacent to the current block. Next, a best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0079] Then, a motion vector for the current block is derived based on the motion vector of the selected candidate. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as the motion vector for the current block as is. Also, for example, the motion vector for the current block may be derived by performing pattern matching in a peripheral area of ​​a position in a reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed in a similar manner in a peripheral area of ​​the best candidate MV, and if an MV with a better evaluation value is found, the best candidate MV may be updated to the MV and used as the final MV for the current block. Note that a configuration may be adopted in which this process is not performed.

[0080] The same processing may be performed when processing is performed in sub-block units.

[0081] The evaluation value is calculated by finding the difference between the reconstructed image and a predetermined area by pattern matching between the area in the reference picture corresponding to the motion vector. The evaluation value may be calculated using other information in addition to the difference.

[0082] As the pattern matching, first pattern matching or second pattern matching is used. The first pattern matching and second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0083] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture that is along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the candidate.

[0084] FIG. 6 is a diagram illustrating an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most closely matched pair of two blocks along the motion trajectory of a current block (Cur block) in two different reference pictures (Ref0, Ref1). Specifically, for the current block, a difference is derived between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by a candidate MV and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV by the display time interval, and an evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs may be selected as the final MV.

[0085] Under the assumption of continuous motion trajectories, motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between a current picture (CurPic) and two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures temporally and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives bidirectional motion vectors that are mirror-symmetric.

[0086] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the candidate.

[0087] 7 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of a current block is derived by searching a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in the current picture (Cur Pic). Specifically, a difference is derived between a reconstructed image of both or either of the coded areas adjacent to the left and / or above the current block and a reconstructed image at the same position in the coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the multiple candidate MVs is selected as the best candidate MV.

[0088] Information indicating whether such a FRUC mode is applied (e.g., called an FRUC flag) is signaled at the CU level. Furthermore, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (e.g., called an FRUC mode flag) is signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0089] Here, we will explain a mode in which motion vectors are derived based on a model that assumes uniform linear motion. This mode is sometimes called BIO (bi-directional optical flow) mode.

[0090] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v y) denotes a velocity vector, and τ0 and τ1 denote the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0 and Ref1), respectively. (MVx0,MVy0) denotes a motion vector corresponding to reference picture Ref0, and (MVx1,MVy1) denotes a motion vector corresponding to reference picture Ref1.

[0091] At this time, the velocity vector (v x ,v y ), (MVx0,MVy0) and (MVx1,MVy1) are respectively (v x τ0,v y τ0) and (-v x τ1,-v y τ1), and the following optical flow equation (1) holds:

[0092]

number

[0093] Here, I (k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like are corrected pixel by pixel.

[0094] Note that the decoding device may derive motion vectors using a method other than that based on a model assuming constant-velocity linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0095] Here, a mode in which a motion vector is derived for each sub-block based on the motion vectors of a plurality of neighboring blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.

[0096] FIG. 9A is a diagram for explaining the derivation of motion vectors for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v x ,v y ) is derived.

[0097]

number

[0098] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting coefficient.

[0099] Such an affine motion compensation prediction mode may include several modes in which the methods of deriving the motion vectors of the upper-left and upper-right corner control points are different. Information indicating such an affine motion compensation prediction mode (e.g., called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0100] [Predictive control unit] The prediction control unit 128 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.

[0101] Here, an example of deriving a motion vector for a picture to be coded in merge mode will be described. Fig. 9B is a diagram for explaining an overview of the motion vector derivation process in merge mode.

[0102] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include spatially adjacent prediction MVs, which are MVs held by multiple coded blocks spatially located around the block to be coded, temporally adjacent prediction MVs, which are MVs held by blocks in the vicinity of the block to be coded projected onto the coded reference picture, joint prediction MVs, which are MVs generated by combining the MV values ​​of the spatially adjacent prediction MVs and the temporally adjacent prediction MVs, and zero prediction MVs, which are MVs with a value of zero.

[0103] Next, one prediction MV is selected from the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV for the block to be coded.

[0104] Furthermore, the variable length coding unit encodes the stream by describing merge_idx, which is a signal indicating which predicted MV has been selected.

[0105] Note that the predicted MVs registered in the predicted MV list described in Figure 9B are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include predicted MVs other than the types of predicted MVs shown in the figure.

[0106] The final MV may be determined by performing the DMVR process, which will be described later, using the MV of the block to be coded derived in the merge mode.

[0107] Here, an example of determining the MV using the DMVR process will be described.

[0108] FIG. 9C is a conceptual diagram for explaining an outline of the DMVR process.

[0109] First, the optimal MVP set for the block to be processed is set as a candidate MV, and reference pixels are obtained from the first reference picture, which is a processed picture in the L0 direction, and the second reference picture, which is a processed picture in the L1 direction, according to the candidate MV, and a template is generated by averaging each reference pixel.

[0110] Next, the template is used to search the surrounding areas of the candidate MVs in the first and second reference pictures, and the MV with the smallest cost is determined as the final MV. The cost value is calculated using the difference between each pixel value of the template and each pixel value of the search area, the MV value, etc.

[0111] The outline of the processing described here is basically the same for the encoding device and the decoding device.

[0112] Note that other processing may be used instead of the processing described here, as long as it is processing that can search the vicinity of the candidate MV and derive the final MV.

[0113] Here, a mode for generating a predicted image using LIC processing will be described.

[0114] FIG. 9D is a diagram for explaining an outline of a predicted image generation method using luminance correction processing by LIC processing.

[0115] First, an MV for obtaining a reference image corresponding to a block to be coded is derived from a reference picture that is a coded picture.

[0116] Next, for the block to be coded, the luminance pixel values ​​of the coded surrounding reference areas adjacent to the left and above and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the MV are used to extract information indicating how the luminance values ​​have changed between the reference picture and the picture to be coded, and a luminance correction parameter is calculated.

[0117] A predicted image for the block to be coded is generated by performing luminance correction processing on a reference image in a reference picture specified by the MV using the luminance correction parameters.

[0118] The shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0119] Although the process of generating a predicted image from one reference picture has been described here, the process is similar when generating a predicted image from multiple reference pictures, and a luminance correction process is performed in a similar manner on the reference images obtained from each reference picture before generating a predicted image.

[0120] As a method for determining whether to apply LIC processing, for example, there is a method using lic_flag, which is a signal indicating whether to apply LIC processing. As a specific example, an encoding device determines whether the encoding target block belongs to an area where a luminance change occurs, and if it belongs to an area where a luminance change occurs, sets the value of lic_flag to 1 and performs encoding by applying LIC processing, and if it does not belong to an area where a luminance change occurs, sets the value of lic_flag to 0 and performs encoding without applying LIC processing. On the other hand, a decoding device decodes lic_flag described in the stream, and switches whether to apply LIC processing depending on the value, and performs decoding.

[0121] As another method for determining whether to apply LIC processing, for example, there is also a method for determining whether LIC processing has been applied to surrounding blocks.As a specific example, when the block to be coded is in merge mode, it is determined whether the surrounding coded blocks selected when deriving MV in merge mode processing have been coded using LIC processing, and depending on the result, whether to apply LIC processing is switched and coded.In addition, in this example, the process in decoding is exactly the same.

[0122] [Overview of the decoding device] Next, an overview will be given of a decoding device capable of decoding the coded signal (coded bitstream) output from the above coding device 100. Fig. 10 is a block diagram showing the functional configuration of a decoding device 200 according to Embodiment 1. The decoding device 200 is a video / image decoding device that decodes video / images on a block-by-block basis.

[0123] As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0124] The decoding device 200 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0125] Each component included in the decoding device 200 will be described below.

[0126] [Entropy Decoding] The entropy decoding unit 202 entropy-decodes the coded bitstream. Specifically, the entropy decoding unit 202 arithmetically decodes the coded bitstream into a binary signal. The entropy decoding unit 202 then debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantized coefficients to the inverse quantization unit 204 on a block-by-block basis.

[0127] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of a block to be decoded (hereinafter referred to as a current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. The inverse quantization unit 204 then outputs the inverse quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0128] [Inverse conversion section] The inverse transform unit 206 restores the prediction error by inverse transforming the transform coefficients input from the inverse quantization unit 204 .

[0129] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0130] Also, for example, if the information decoded from the coded bitstream indicates that NSST is to be applied, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0131] [Addition section] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transformer 206 and the prediction sample input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0132] [Block Memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter referred to as a current picture). Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208.

[0133] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0134] If the information indicating ALF on / off read from the encoded bitstream indicates that ALF is on, one filter is selected from multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0135] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0136] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction based on the intra prediction mode interpreted from the encoded bitstream, by referring to blocks in the current picture stored in the block memory 210. Specifically, the intra prediction unit 216 generates the intra prediction signal by performing intra prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0137] Note that when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0138] Furthermore, when information interpreted from the coded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects pixel values ​​after intra prediction based on the gradients of reference pixels in the horizontal and vertical directions.

[0139] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the coded bitstream, and outputs the inter prediction signal to the prediction control unit 220.

[0140] In addition, if the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks.

[0141] Furthermore, if the information interpreted from the coded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the coded bitstream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0142] Furthermore, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming constant-velocity linear motion. Furthermore, when information interpreted from the coded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks.

[0143] [Predictive control unit] The prediction control unit 220 selects either the intra-prediction signal or the inter-prediction signal, and outputs the selected signal to the addition unit 208 as a prediction signal.

[0144] [First aspect] The encoding device 100, the decoding device 200, the encoding method, and the decoding method according to the first embodiment of the present disclosure will be described below.

[0145] [First example of encoding and decoding processes using quantization matrices] 11 is a diagram showing a first example of a coding process flow using a quantization matrix (QM) in the coding device 100. The coding device 100 described here performs coding processing for each square or rectangular block obtained by dividing a picture (hereinafter also referred to as a screen) included in a moving image.

[0146] First, in step S101, the quantizer 108 generates a QM for a square block. The QM for a square block is a quantization matrix for multiple transform coefficients of the square block. Hereinafter, the QM for a square block is also referred to as a first quantization matrix. The quantizer 108 may generate the QM for a square block from a value defined by the user and set in the encoding device 100, or may adaptively generate the QM using encoding information of a picture that has already been encoded. The entropy encoder 110 may also describe a signal related to the QM for a square block generated by the quantizer 108 in the bitstream. In this case, the QM for a square block may be coded in the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. The QM for a square block does not have to be described in the stream. In this case, the quantizer 108 may use a default QM value for a square block predefined by a standard. Furthermore, the entropy coding unit 110 may write only some of the quantization coefficients required to generate the QM for square blocks into the stream, rather than writing all of the matrix coefficients (i.e., quantization coefficients) of the QM for square blocks into the stream, thereby reducing the amount of information to be coded.

[0147] Next, in step S102, the quantization unit 108 generates a QM for a rectangular block using the QM for the square block generated in step S101. The QM for a rectangular block is a quantization matrix for multiple transform coefficients of the rectangular block. Hereinafter, the QM for a rectangular block is also referred to as a second quantization matrix. Note that the entropy coding unit 110 does not write a QM signal for the rectangular block into the stream.

[0148] The processes of steps S101 and S102 may be configured to be performed collectively at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processes every time in block-by-block processing. Furthermore, the QMs generated by the quantization unit 108 in steps S101 and S102 may be configured to generate multiple types of QMs for blocks of the same block size according to conditions such as for luminance blocks / chrominance blocks, for intra-prediction blocks / inter-prediction blocks, and other conditions.

[0149] Next, a block-by-block loop is started. First, in step S103, the intra prediction unit 124 or the inter prediction unit 126 performs a prediction process using intra prediction or inter prediction, etc., on a block-by-block basis. In step S104, the transform unit 106 performs a transform process using a discrete cosine transform (DCT), etc., on the generated prediction residual image. In step S105, the quantization unit 108 performs a quantization process on the generated transform coefficients using the QM for square blocks and the QM for rectangular blocks, which are the outputs of steps S101 and S102. Note that, in inter prediction, a mode that references a block in a picture other than the picture to which the current block belongs may be used, as well as a mode that references a block in the picture to which the current block belongs. In this case, the QM for inter prediction may be used in common for both modes, or the QM for intra prediction may be used for the mode that references a block in the picture to which the current block belongs. Furthermore, in step S106, the inverse quantization unit 112 performs inverse quantization processing on the quantized transform coefficients using the QM for square blocks and the QM for rectangular blocks output from steps S101 and S102. In step S107, the inverse transform unit 114 performs inverse transform processing on the inverse quantized transform coefficients to generate a residual (prediction error) image. Next, in step S108, the adder 116 adds the residual image and the prediction image to generate a reconstructed image. This series of processing flows is repeated to complete the block-by-block loop.

[0150] This allows encoding processing to be performed even in encoding formats with rectangular blocks of various shapes by describing only QMs corresponding to square blocks in the stream, without describing QMs corresponding to rectangular blocks of each shape in the stream. That is, according to the encoding device 100 of the first aspect of the present disclosure, the QM corresponding to rectangular blocks is not described in the stream, thereby reducing the amount of code in the header region. Furthermore, according to the encoding device 100 of the first aspect of the present disclosure, a QM corresponding to a rectangular block can be generated from a QM corresponding to a square block, thereby enabling the use of an appropriate QM for rectangular blocks without increasing the amount of code in the header region. Therefore, according to the encoding device 100 of the first aspect of the present disclosure, quantization can be efficiently performed on rectangular blocks of various shapes, potentially improving encoding efficiency. The QM for square blocks does not need to be described in the stream; instead, a default QM value for square blocks predefined by the standard may be used.

[0151] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0152] Fig. 12 is a diagram showing an example of a decoding process flow using a quantization matrix (QM) in a decoding device 200 corresponding to the encoding device 100 described in Fig. 11. Note that the decoding device 200 described here performs decoding processing for each square or rectangular block obtained by dividing a screen.

[0153] First, in step S201, the entropy decoding unit 202 decodes a signal related to a QM for square blocks from the stream and generates a QM for square blocks using the decoded signal related to the QM for square blocks. The QM for square blocks may be decoded from the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. The QM for square blocks may not be decoded from the stream. In this case, a default value predefined by a standard may be used as the QM for square blocks. The entropy decoding unit 202 may generate the QM by decoding only a portion of the quantization coefficients required to generate the QM from the stream, rather than decoding all of the quantization coefficients of the matrix of the QM for square blocks from the stream.

[0154] Next, in step S202, the entropy decoding unit 202 generates a QM for a rectangular block using the QM for a square block generated in step S201. Note that the entropy decoding unit 202 does not decode the QM signal for a rectangular block from the stream.

[0155] The processes of steps S201 and S202 may be configured to be performed collectively at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processes every time in block-by-block processing. Furthermore, the QMs generated by the entropy decoding unit 202 in steps S201 and S202 may be configured to generate multiple types of QMs for blocks of the same block size, such as for luminance blocks / chrominance blocks, intra-prediction blocks / inter-prediction blocks, and depending on other conditions.

[0156] Next, a block-by-block loop is started. First, in step S203, the intra prediction unit 216 or the inter prediction unit 218 performs a prediction process using intra prediction, inter prediction, or the like, on a block-by-block basis. In step S204, the inverse quantization unit 204 performs an inverse quantization process on the quantized transform coefficients (i.e., quantized coefficients) decoded from the stream, using the QM for square blocks and the QM for rectangular blocks output from steps S201 and S202. Note that, in inter prediction, a mode that references a block in a picture other than the picture to which the current block belongs may be used, as well as a mode that references a block in the picture to which the current block belongs. In this case, the QM for inter prediction may be used in common for both modes, or the QM for intra prediction may be used for the mode that references a block in the picture to which the current block belongs. Next, in step S205, the inverse transform unit 206 performs an inverse transform process on the inversely quantized transform coefficients to generate a residual (prediction error) image. Next, in step S206, the adder 208 generates a reconstructed image by adding the residual image and the predicted image. This series of processing steps is repeated to complete the block-by-block loop.

[0157] As a result, even in a decoding method using rectangular blocks of various shapes, decoding is possible as long as only the QM corresponding to the square blocks is described in the stream, even if the QM corresponding to each rectangular block shape is not described in the stream. In other words, according to the decoding device 200 of the first aspect of the present disclosure, the QM corresponding to the rectangular blocks is not described in the stream, thereby reducing the amount of code in the header region. Furthermore, according to the decoding device 200 of the first aspect of the present disclosure, a QM corresponding to the rectangular block can be generated from a QM corresponding to a square block, thereby enabling the use of an appropriate QM for the rectangular block without increasing the amount of code in the header region. Therefore, according to the decoding device 200 of the first aspect of the present disclosure, efficient quantization of rectangular blocks of various shapes is possible, thereby potentially improving coding efficiency. The QM for the square block does not need to be described in the stream; instead, a default QM for square blocks predefined by the standard may be used.

[0158] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0159] [First example of how to generate a QM for the rectangular block in the first example] Fig. 13 is a diagram illustrating a first example of generating a QM for a rectangular block from a QM for a square block in step S102 in Fig. 11 and step S202 in Fig. 12. Note that the process described here is common to the encoding device 100 and the decoding device 200.

[0160] 13 illustrates the size of a QM for a rectangular block generated from a QM for each square block having a size ranging from 2×2 to 256×256. The example illustrated in FIG. 13 is characterized in that the length of the long side of each rectangular block is the same as the length of one side of the corresponding square block. In other words, this example is characterized in that the size of the rectangular block serving as the target block is smaller than the size of the square block. In other words, the encoding device 100 and decoding device 200 according to the first aspect of the present disclosure generate a QM for a rectangular block by down-converting a QM for a square block having one side the same length as the long side of the target rectangular block.

[0161] Note that FIG. 13 shows the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each QM for square blocks, without distinguishing between luminance blocks and chrominance blocks. The correspondence between QMs for square blocks and QMs for rectangular blocks that is adapted to the actual format used may be derived as appropriate. For example, in the case of the 4:2:0 format, luminance blocks are twice the size of chrominance blocks. Therefore, when referring to luminance blocks in the process of generating a QM for a rectangular block from a QM for a square block, usable QMs for square blocks correspond to square blocks with sizes ranging from 4×4 to 256×256. In this case, only QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used as QMs for rectangular blocks generated from a QM for a square block. Furthermore, when referring to chrominance blocks in the process of generating a QM for a rectangular block from a QM for a square block, usable QMs for square blocks correspond to square blocks with sizes ranging from 2×2 to 128×128. In this case, only QMs corresponding to rectangular blocks whose short side length is 2 or more and whose long side length is 128 or less are used as QMs for rectangular blocks generated from QMs for square blocks.

[0162] Also, for example, in the 4:4:4 format, luma blocks are blocks of the same size as chroma blocks, so when referring to chroma blocks in the process of generating a QM for rectangular blocks, just like when referring to luma blocks, the usable QMs for square blocks correspond to square blocks of sizes ranging from 4x4 to 256x256.

[0163] In this way, it is advisable to appropriately derive the correspondence between the QM for square blocks and the QM for rectangular blocks depending on the format actually used.

[0164] Note that the block sizes shown in Fig. 13 are merely examples and are not limiting. For example, QMs of block sizes other than those shown in Fig. 13 may be usable, or only QMs for square blocks of some of the block sizes shown in Fig. 13 may be usable.

[0165] FIG. 14 is a diagram for explaining a method for generating the QM for the rectangular block described in FIG. 13 by down-converting from the corresponding QM for the square block.

[0166] In the example of FIG. 14, a QM for an 8×4 rectangular block is generated from a QM for an 8×8 square block.

[0167] In the down-conversion process, the multiple matrix elements of the QM for square blocks are divided into groups with the same number as the number of multiple matrix elements of the QM for rectangular blocks, and for each of the multiple groups, the multiple matrix elements included in that group are arranged consecutively in the horizontal or vertical direction of the square block, and for each of the multiple groups, the matrix element located at the lowest frequency among the multiple matrix elements included in that group may be determined to be the matrix element corresponding to that group in the QM for rectangular blocks.

[0168] For example, in FIG. 14, a plurality of matrix elements of a QM for an 8×8 square block are surrounded by a thick line every predetermined number. The predetermined number of matrix elements surrounded by the thick line constitute one group. In the down-conversion process illustrated in FIG. 14, the QM for the 8×8 square block is divided so that the number of these groups is the same as the number of matrix elements (also called quantization coefficients) of the QM for a rectangular block generated from the QM for the 8×8 square block. In the example of FIG. 14, two quantization coefficients adjacent in the vertical direction constitute one group. Next, in the QM for the 8×8 square block, the quantization coefficient located at the lowest frequency side (the upper side in the example of FIG. 14) of each group is selected and used as the value of the QM for the 8×4 rectangular block.

[0169] Note that the method of selecting one quantization coefficient from each group as the QM value for the rectangular block is not limited to the above example, and other methods may be used. For example, the quantization coefficient located at the highest frequency in the group may be used as the QM value for the rectangular block, or the quantization coefficient located at the middle frequency may be used as the QM value for the rectangular block. Alternatively, the average, minimum, maximum, or intermediate value of all or some of the quantization coefficients in the group may be used. Note that if a decimal point is generated as a result of calculating these values, it may be rounded up, down, or to the nearest integer.

[0170] The method of selecting one quantization coefficient from each group in the QM for square blocks may be switched depending on the frequency range that each group is located in in the QM for square blocks. For example, the quantization coefficient located at the lowest frequency in the group may be selected for a group located in the low frequency range, the quantization coefficient located at the highest frequency in the group may be selected for a group located in the high frequency range, and the quantization coefficient located at the middle frequency in the group may be selected.

[0171] Note that the lowest frequency component of the QM for the rectangular block to be generated (the quantization coefficient at the top left in the example of Figure 14) may be written in the stream and set directly from the stream, rather than being derived from the QM for the square block. In this case, the amount of information written in the stream increases, which increases the amount of code in the header area, but it becomes possible to directly control the quantization coefficient of the QM for the lowest frequency component, which has the greatest impact on image quality, which increases the possibility of improving image quality.

[0172] Note that, although an example has been described here in which a QM for a square block is down-converted in the vertical direction to generate a QM for a rectangular block, a method similar to that of the example in Figure 14 may also be used to down-convert a QM for a square block in the horizontal direction to generate a QM for a rectangular block.

[0173] [Second example of how to generate a QM for the rectangular block in the first example] Fig. 15 is a diagram illustrating a second example of generating a QM for a rectangular block from a QM for a square block in step S102 in Fig. 11 and step S202 in Fig. 12. Note that the process described here is common to the encoding device 100 and the decoding device 200.

[0174] 15 illustrates the size of a QM for a rectangular block generated from a QM for each square block having a size ranging from 2×2 to 256×256. The example illustrated in FIG. 15 is characterized in that the length of the short side of each rectangular block is the same as the length of one side of the corresponding square block. In other words, this example is characterized in that the size of the rectangular block serving as the current block is larger than the size of the square block. In other words, the encoding device 100 and decoding device 200 according to the first aspect of the present disclosure generate a QM for a rectangular block by up-converting a QM for a square block having a side the same length as the short side of the current rectangular block.

[0175] Note that FIG. 15 shows the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from the QMs for each square block, without distinguishing between luminance blocks and chrominance blocks. The correspondence between QMs for square blocks and QMs for rectangular blocks adapted to the format actually used may be derived as appropriate. For example, in the case of the 4:2:0 format, when a luminance block is referenced in the process of generating a QM for a rectangular block from a QM for a square block, only QMs corresponding to rectangular blocks having a short side length of 4 or more and a long side length of 256 or less are used for the QM for the rectangular block generated from the QM for the square block. Also, when a chrominance block is referenced in the process of generating a QM for a rectangular block from a QM for a square block, only QMs corresponding to rectangular blocks having a short side length of 2 or more and a long side length of 128 or less are used for the QM for the rectangular block generated from the QM for the square block. Note that the content of the 4:4:4 format is the same as that described in FIG. 13, so a description thereof will be omitted here.

[0176] In this way, it is advisable to appropriately derive the correspondence between the QM for square blocks and the QM for rectangular blocks depending on the format actually used.

[0177] Note that the block sizes shown in Fig. 15 are merely examples and are not limiting. For example, it may be possible to use QMs for square blocks of block sizes other than those shown in Fig. 15, or it may be possible to use QMs for square blocks of only some of the block sizes shown in Fig. 15.

[0178] FIG. 16 is a diagram for explaining a method for generating the QM for the rectangular block described in FIG. 15 by up-converting from the QM for the corresponding square block.

[0179] In the example of FIG. 16, a QM for an 8×4 rectangular block is generated from a QM for a 4×4 square block.

[0180] In the upconversion process, (i) the matrix elements of the QM for rectangular blocks may be divided into groups equal in number to the number of matrix elements of the QM for square blocks, and for each of the groups, the matrix elements contained in that group may be overlapped to determine the matrix elements corresponding to that group in the QM for rectangular blocks, or (ii) the matrix elements of the QM for rectangular blocks may be determined by performing linear interpolation between adjacent matrix elements among the matrix elements of the QM for rectangular blocks.

[0181] For example, in FIG. 16, a predetermined number of matrix elements of a QM for an 8×4 rectangular block are surrounded by thick lines. The predetermined number of matrix elements surrounded by thick lines constitute one group. In the upconversion process illustrated in FIG. 16, the QM for the 8×4 rectangular block is divided so that the number of these groups is the same as the number of matrix elements (also referred to as quantization coefficients) of the QM for the corresponding 4×4 square block. In the example of FIG. 16, two quantization coefficients adjacent in the horizontal direction constitute one group. Next, in the QM for the 8×4 rectangular block, the value of the quantization coefficient of the QM for the square block corresponding to that group is selected and spread within that group as the quantization coefficients constituting each group, thereby obtaining the value of the QM for the 8×4 rectangular block.

[0182] The method for deriving the quantization coefficients in each group in the QM for rectangular blocks is not limited to the above example, and other methods may be used. For example, the quantization coefficients in each group may be derived so that they are continuous values ​​by performing linear interpolation or the like with reference to the values ​​of the quantization coefficients in adjacent frequency domains. If a decimal point is generated as a result of the calculation of these values, it may be rounded to an integer by rounding up, rounding down, or rounding to the nearest integer.

[0183] The method for deriving the quantization coefficients in each group in the QM for rectangular blocks may be switched depending on the frequency range that each group is located in. For example, in a group located in a low frequency range, the quantization coefficients in the group may be derived so that the values ​​of each quantization coefficient in the group are as small as possible, and in a group located in a high frequency range, the quantization coefficients in the group may be derived so that the values ​​of each quantization coefficient in the group are as large as possible.

[0184] Note that the lowest frequency component of the QM for the rectangular block to be generated (the quantization coefficient at the top left in the example of Figure 16) may be written in the stream and set directly from the stream, rather than being derived from the QM for the square block. In this case, the amount of information written in the stream increases, which increases the amount of code in the header area, but it becomes possible to directly control the quantization coefficient of the QM for the lowest frequency component, which has the greatest impact on image quality, which increases the possibility of improving image quality.

[0185] Note that, although an example has been described here in which a QM for a square block is upconverted vertically to generate a QM for a rectangular block, a method similar to that shown in the example of Figure 16 may also be used to upconvert a QM for a square block horizontally to generate a QM for a rectangular block.

[0186] [Other variations of the first example of encoding and decoding processes using quantization matrices] As a method for generating a QM for rectangular blocks from a QM for square blocks, the first example of the method for generating a QM for rectangular blocks described with reference to Figures 13 and 14 and the second example of the method for generating a QM for rectangular blocks described with reference to Figures 15 and 16 may be switched depending on the size of the rectangular block to be generated. For example, one method is to compare the ratio of the length and width of the rectangular block (the down-conversion or up-conversion ratio) with a threshold, and use the first example if it is greater than the threshold, and use the second example if it is less than the threshold. Another method is to write a flag in the stream for each rectangular block size indicating whether to use the first or second example, and switch between them. This makes it possible to switch between down-conversion and up-conversion depending on the size of the rectangular block, thereby enabling the generation of a more appropriate QM for rectangular blocks.

[0187] Instead of switching between up-conversion and down-conversion for each rectangular block size, a combination of up-conversion and down-conversion may be used for a single rectangular block. For example, a QM for a 32x32 square block may be up-converted horizontally to generate a QM for a 32x64 rectangular block, and then this QM for the 32x64 rectangular block may be down-converted vertically to generate a QM for a 16x64 rectangular block.

[0188] Alternatively, a single rectangular block may be upconverted in two directions. For example, a QM for a 16x16 square block may be upconverted vertically to generate a QM for a 32x16 rectangular block, and then this QM for the 32x16 rectangular block may be upconverted horizontally to generate a QM for a 32x64 rectangular block.

[0189] Note that down-conversion processing may be performed in two directions for one rectangular block. For example, a QM for a 64x64 square block may be down-converted horizontally to generate a QM for a 64x32 rectangular block, and then this QM for the 64x32 rectangular block may be down-converted vertically to generate a QM for a 16x32 rectangular block.

[0190] [Effect of the first example of encoding and decoding processes using quantization matrices] According to the encoding device 100 and the decoding device 200 of the first aspect of the present disclosure, the configurations described with reference to FIGS. 11 and 12 enable encoding and decoding of rectangular blocks even in encoding formats having rectangular blocks of various shapes by describing only QMs corresponding to square blocks in the stream, without describing QMs corresponding to rectangular blocks of each shape in the stream. In other words, according to the encoding device 100 and the decoding device 200 of the first aspect of the present disclosure, the QM corresponding to rectangular blocks is not described in the stream, thereby reducing the amount of code in the header region. Furthermore, according to the encoding device 100 and the decoding device 200 of the first aspect of the present disclosure, a QM corresponding to a rectangular block can be generated from a QM corresponding to a square block, thereby enabling the use of an appropriate QM for a rectangular block without increasing the amount of code in the header region. Therefore, according to the encoding device 100 and the decoding device 200 of the first aspect of the present disclosure, quantization can be efficiently performed on rectangular blocks of various shapes, which increases the likelihood of improving coding efficiency.

[0191] For example, encoding device 100 is an encoding device that encodes moving images by performing quantization, and includes a circuit and a memory. The circuit uses the memory to convert a first quantization matrix for multiple transform coefficients of a square block, thereby generating a second quantization matrix for multiple transform coefficients of a rectangular block from the first quantization matrix, and quantizes the multiple transform coefficients of the rectangular block using the second quantization matrix.

[0192] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be coded. This reduces the amount of coding, improving processing efficiency. Therefore, according to the coding device 100, quantization can be efficiently performed on rectangular blocks.

[0193] For example, in the encoding device 100, the circuit may encode only the first quantization matrix of the first and second quantization matrices into a bitstream.

[0194] This reduces the amount of coding, and therefore the encoding device 100 improves processing efficiency.

[0195] For example, in the encoding device 100, the circuit may generate the second quantization matrix for the multiple transform coefficients of the rectangular block by performing a down-conversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the long side of the rectangular block, which is the block to be processed.

[0196] This allows the encoding device 100 to efficiently generate a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the long side of the rectangular block.

[0197] For example, in the encoding device 100, in the down-conversion process, the circuit may divide the plurality of matrix elements of the first quantization matrix into groups of the same number as the number of matrix elements of the second quantization matrix, and for each of the plurality of groups, the plurality of matrix elements included in the group may be arranged consecutively in the horizontal or vertical direction of the square block, and for each of the plurality of groups, the matrix element located on the lowest frequency side among the plurality of matrix elements included in the group, the matrix element located on the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group may be determined as the matrix element corresponding to the group in the second quantization matrix.

[0198] This allows encoding device 100 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0199] For example, in the encoding device 100, the circuit may generate the second quantization matrix for the multiple transform coefficients of the rectangular block by performing an upconversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the short side of the rectangular block, which is the block to be processed.

[0200] This allows the encoding device 100 to efficiently generate a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the short side of the rectangular block.

[0201] For example, in the encoding device 100, in the upconversion process, the circuit may (i) divide the plurality of matrix elements of the second quantization matrix into groups the number of which is equal to the number of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix elements in the second quantization matrix corresponding to that group by overlapping the plurality of matrix elements included in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements among the plurality of matrix elements of the second quantization matrix.

[0202] This allows encoding device 100 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0203] For example, in the encoding device 100, the circuit may generate the second quantization matrix for the multiple rectangular transform coefficients by switching between a method of performing the down-conversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the long side of the rectangular block, and a method of performing the up-conversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the lengths of the short sides and the long sides of the rectangular block that is the block to be processed.

[0204] This allows encoding device 100 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling generation of a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0205] In addition, the decoding device 200 is a decoding device that decodes moving images by performing inverse quantization, and includes a circuit and a memory. The circuit uses the memory to transform a first quantization matrix for multiple transform coefficients of a square block, thereby generating a second quantization matrix for multiple transform coefficients of a rectangular block from the first quantization matrix, and performs inverse quantization on the multiple quantization coefficients of the rectangular block using the second quantization matrix.

[0206] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be decoded. This allows for a reduction in the amount of code, improving processing efficiency. Therefore, the decoding device 200 can efficiently perform inverse quantization on rectangular blocks.

[0207] For example, in the decoding device 200, the circuit may decode only the first quantization matrix of the first and second quantization matrices from the bitstream.

[0208] This allows the amount of coding to be reduced, and therefore the decoding device 200 improves processing efficiency.

[0209] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the multiple transform coefficients of the rectangular block by performing a down-conversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the long side of the rectangular block, which is the block to be processed.

[0210] This allows the decoding device 200 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the long side of the rectangular block.

[0211] For example, in the decoding device 200, in the down-conversion process, the circuit may divide the plurality of matrix elements of the first quantization matrix into groups of the same number as the number of matrix elements of the second quantization matrix, and for each of the plurality of groups, determine the matrix element located at the lowest frequency side among the plurality of matrix elements included in the group, the transform coefficient located at the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group as the matrix element corresponding to the group in the second quantization matrix.

[0212] This allows the decoding device 200 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0213] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the multiple transform coefficients of the rectangular block by performing an upconversion process from the first quantization matrix for the multiple transform coefficients of the square block having one side the same length as the short side of the rectangular block, which is the block to be processed.

[0214] This allows the decoding device 200 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the short side of the rectangular block.

[0215] For example, in the decoding device 200, in the upconversion process, the circuit may (i) divide the plurality of matrix elements of the second quantization matrix into groups the number of which is equal to the number of the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix elements in the second quantization matrix corresponding to that group by overlapping the plurality of matrix elements included in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements among the plurality of matrix elements of the second quantization matrix.

[0216] This allows the decoding device 200 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0217] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the rectangular transform coefficients by switching between a method of generating the second quantization matrix by performing the down-conversion process from the first quantization matrix for multiple transform coefficients of the square block having one side the same length as the long side of the rectangular block, and a method of generating the second quantization matrix by performing the up-conversion process from the first quantization matrix for multiple transform coefficients of the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the length of the short side and the length of the long side of the rectangular block that is the block to be processed.

[0218] This allows the decoding device 200 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling it to generate a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0219] The encoding method is also an encoding method for encoding a moving image by performing quantization, which converts a first quantization matrix for a plurality of transform coefficients of a square block to generate a second quantization matrix for a plurality of transform coefficients of a rectangular block from the first quantization matrix, and quantizes the plurality of transform coefficients of the rectangular block using the second quantization matrix.

[0220] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be coded. This reduces the amount of coding, improving processing efficiency. Therefore, according to the coding method, quantization can be efficiently performed on rectangular blocks.

[0221] The decoding method is a decoding method for decoding a moving image by performing inverse quantization, which converts a first quantization matrix for a plurality of transform coefficients of a square block to generate a second quantization matrix for a plurality of transform coefficients of a rectangular block from the first quantization matrix, and performs inverse quantization on the plurality of quantization coefficients of the rectangular block using the second quantization coefficients.

[0222] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be decoded. This makes it possible to reduce the amount of code, thereby improving processing efficiency. Therefore, according to the decoding method, dequantization can be efficiently performed on the rectangular block.

[0223] This aspect may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some of the processes, some of the device configurations, and some of the syntax described in the flowcharts of this aspect may be implemented in combination with other aspects.

[0224] [Second mode] The encoding device 100, the decoding device 200, the encoding method, and the decoding method according to the second embodiment of the present disclosure will be described below.

[0225] [Second example of encoding and decoding processes using quantization matrices] 17 is a diagram showing a second example of the encoding process flow using a quantization matrix (QM) in the encoding device 100. Note that the encoding device 100 described here performs encoding processing for each square or rectangular block obtained by dividing a screen.

[0226] First, in step S701, the quantization unit 108 generates a QM corresponding to the size of the effective transform coefficient area for each block size, i.e., square block and rectangular block. In other words, the quantization unit 108 quantizes only a plurality of transform coefficients in a predetermined area on the low-frequency side among a plurality of transform coefficients included in the target block, using a quantization matrix.

[0227] The entropy coding unit 110 describes, in the stream, a signal related to the QM corresponding to the valid transform coefficient region generated in step S701. In other words, the entropy coding unit 110 encodes, into the bitstream, a signal related to a quantization matrix corresponding only to a plurality of transform coefficients within a predetermined range on the low-frequency side. The quantization unit 108 may generate the QM corresponding to the valid transform coefficient region from a value defined by the user and set in the coding device 100, or may adaptively generate the QM using coding information of a picture that has already been coded. The QM corresponding to the valid transform coefficient region may be coded in the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. The QM corresponding to the valid transform coefficient region does not need to be described in the stream. In this case, the quantization unit 108 may use a default value predefined by the standard as the value of the QM corresponding to the valid transform coefficient region.

[0228] The processing of step S701 may be configured to be performed all at once at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processing every time during block-by-block processing. Furthermore, the QM generated in step S701 may be configured to generate multiple types of QMs for blocks of the same block size, such as for luminance blocks / chrominance blocks, for intra-prediction blocks / inter-prediction blocks, and for other conditions.

[0229] In the processing flow shown in FIG. 17, the processing steps other than step S701 are loop processing in units of blocks, and are the same as the processing in the first example described with reference to FIG.

[0230] This allows for encoding of a target block having a block size in which only a portion of the low-frequency region of the multiple transform coefficients included in the target block contains valid transform coefficients without needlessly describing signals related to the QM of invalid regions in the stream. This makes it possible to reduce the amount of code in the header region, thereby increasing the possibility of improving encoding efficiency.

[0231] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0232] Fig. 18 is a diagram showing an example of a decoding process flow using a quantization matrix (QM) in a decoding device 200 corresponding to the encoding device 100 described in Fig. 17. Note that the decoding device 200 described here performs decoding processing for each square or rectangular block obtained by dividing a screen.

[0233] First, in step S801, the entropy decoding unit 202 decodes a signal related to a QM corresponding to a valid transform coefficient region from the stream, and generates a QM corresponding to the valid transform coefficient region using the decoded signal related to the QM corresponding to the valid transform coefficient region. The QM corresponding to the valid transform coefficient region is a QM corresponding to the size of the valid transform coefficient region for each block size of the current block. Note that the QM corresponding to the valid transform coefficient region may be decoded from the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. Alternatively, the QM corresponding to the valid transform coefficient region may not be decoded from the stream. In this case, for example, a default value predefined in the standard may be used as the QM corresponding to the valid transform coefficient region.

[0234] The processing of step S801 may be configured to be performed all at once at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processing every time in block-by-block processing.Furthermore, the QM generated by the entropy decoding unit 202 in step S801 may be configured to generate multiple types of QMs for blocks of the same block size, for luminance blocks / chrominance blocks, for intra-prediction blocks / inter-prediction blocks, and depending on other conditions.

[0235] In the processing flow shown in FIG. 18, the processing flow other than step S801 is a loop process in units of blocks, and is the same as the processing flow of the first example described using FIG.

[0236] This makes it possible to perform decoding processing for a block size in which only a portion of the low-frequency region of the multiple transform coefficients included in the block to be processed is an area containing valid transform coefficients, even if signals related to QM of invalid areas are not written in the stream unnecessarily. Therefore, it is possible to reduce the amount of code in the header region, which increases the possibility of improving coding efficiency.

[0237] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0238] Fig. 19 is a diagram illustrating an example of QM corresponding to the size of the valid transform coefficient area for each block size in step S701 in Fig. 17 and step S801 in Fig. 18. Note that the processing described here is common to the encoding device 100 and the decoding device 200.

[0239] 19(a) shows an example in which the block size of the target block is a square block of 64×64. Only the 32×32 region on the lower frequency side indicated by diagonal lines in the figure is a region containing valid transform coefficients. In the target block, the transform coefficients in regions other than the valid transform coefficient region are forcibly set to 0, that is, the transform coefficients are invalidated, so quantization processing and inverse quantization processing are not required. In other words, the encoding device 100 and the decoding device 200 according to the second aspect of the present disclosure generate only the 32×32 QM corresponding to the 32×32 region on the lower frequency side indicated by diagonal lines in the figure.

[0240] Next, (b) of Fig. 19 shows an example in which the block size of the block to be processed is a rectangular block of 64 × 32. In (b) of Fig. 19, similar to the example in (a) of Fig. 19, the encoding device 100 and the decoding device 200 generate only a 32 × 32 QM corresponding to the 32 × 32 region on the lower frequency side.

[0241] Next, (c) of Fig. 19 shows an example in which the block size of the block to be processed is a rectangular block of 64 × 16. Unlike the example (a) of Fig. 19, (c) of Fig. 19 has a vertical block size of only 16, so the encoding device 100 and the decoding device 200 generate only a 32 × 16 QM corresponding to the 32 × 16 region on the low frequency side.

[0242] In this way, if either the vertical or horizontal side of the block to be processed is greater than 32, the transform coefficients in the area greater than 32 are invalidated, and only the area of ​​32 or less is treated as the valid transform coefficient area and is subjected to quantization and inverse quantization processing, generating the quantization coefficients of the QM, and encoding and decoding the signal related to the QM into a stream.

[0243] This allows encoding and decoding to be performed without needlessly describing QM-related signals in invalid areas in the stream, thereby reducing the amount of code in the header area, which increases the possibility of improving coding efficiency.

[0244] 19 is merely an example, and other valid transform coefficient area sizes may be used. For example, if the block to be processed is a luminance block, an area of ​​up to 32 × 32 may be the valid transform coefficient area, and if the block to be processed is a chrominance block, an area of ​​up to 16 × 16 may be the valid transform coefficient area. Furthermore, if the long side of the block to be processed is 64, an area of ​​up to 32 × 32 may be the valid transform coefficient area, and if the long side of the block to be processed is 128 or 256, an area of ​​up to 62 × 62 may be the valid transform coefficient area.

[0245] Alternatively, a configuration may be adopted in which, using processing similar to the method described in the first aspect, coefficients of a quantization matrix corresponding to all square and rectangular frequency components are generated once, and then a QM corresponding to only the valid transform coefficient region described with reference to Fig. 19 is generated. In this case, the amount of signals related to the QM described in the stream remains the same as in the method described in the first aspect, but it is possible to omit quantization processing outside the valid transform coefficient region while still being able to generate QMs for all squares and rectangles using the method described in the first aspect. This increases the possibility of reducing the amount of processing involved in the quantization processing.

[0246] [Modification of the second example of encoding and decoding processes using quantization matrices] 20 is a diagram showing a modified example of the second example of the encoding process flow using a quantization matrix (QM) in the encoding device 100. Note that the encoding device 100 described here performs encoding processing for each square or rectangular block obtained by dividing the screen.

[0247] In this modified example, the configuration of the second example described in FIG. 17 is combined with the configuration of the first example described in FIG. 11, and processing of steps S1001 and S1002 is performed instead of step S701 in FIG.

[0248] First, in step S1001, the quantization unit 108 generates a QM for a square block. At this time, the QM for the square block is a QM that corresponds to the size of the valid transform coefficient area in the square block. Furthermore, the entropy coding unit 110 describes a signal related to the QM for the square block generated in step S1001 in a stream. At this time, the signal related to the QM described in the stream is a signal related only to quantized coefficients that correspond to the valid transform coefficient area.

[0249] Next, in step S1002, the quantization unit 108 generates a QM for rectangular blocks using the QM for square blocks generated in step S1001. At this time, the entropy coding unit 110 does not describe a signal related to the QM for rectangular blocks in the stream.

[0250] In the processing flow shown in FIG. 20, the processing other than steps S1001 and S1002 is a loop processing in units of blocks, and is the same as the processing in the first example described using FIG.

[0251] This allows encoding processing to be performed even in encoding methods that have rectangular blocks of various shapes by describing in the stream only the QM-related signals corresponding to square blocks, without describing in the stream signals related to QMs corresponding to rectangular blocks of each shape. Furthermore, for a target block of a block size in which only a region including some of the transform coefficients included in the target block is a valid region (i.e., a valid transform coefficient region), encoding processing is possible without unnecessary description in the stream of QM-related signals for invalid regions. Therefore, since QMs can be used for rectangular blocks while reducing the amount of code in the header region, there is a high possibility of improving encoding efficiency.

[0252] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0253] Fig. 21 is a diagram showing an example of a decoding process flow using a quantization matrix (QM) in a decoding device 200 corresponding to the encoding device 100 described in Fig. 20. Note that the decoding device 200 described here performs decoding processing for each square or rectangular block obtained by dividing a screen.

[0254] In this modified example, the configuration of the second example described in FIG. 18 is combined with the configuration of the first example described in FIG. 12, and processing of steps S1101 and S1102 is performed instead of step S801 in FIG.

[0255] First, in step S1101, the entropy decoding unit 202 decodes a signal related to a QM for a square block from the stream and generates a QM for the square block using the decoded signal related to the QM for the square block. At this time, the signal related to the QM for the square block decoded from the stream is a signal related only to quantization coefficients corresponding to the valid transform coefficient area. Therefore, the QM for the square block generated by the entropy decoding unit 202 is a QM corresponding to the size of the valid transform coefficient area.

[0256] Next, in step S1102, the entropy decoding unit 202 generates a QM for a rectangular block using the QM for a square block generated in step S1101. Note that at this time, the entropy decoding unit 202 does not decode a signal related to the QM for a rectangular block from the stream.

[0257] In the processing flow shown in FIG. 21, the processing other than steps S1101 and S1102 is a loop processing in units of blocks, and is the same as the processing in the first example described using FIG.

[0258] As a result, even in a decoding method having rectangular blocks of various shapes, decoding is possible as long as only signals related to QMs corresponding to square blocks are described in the stream, even if signals related to QMs corresponding to rectangular blocks of various shapes are not described in the stream. Furthermore, for a processing target block having a block size in which only a region including some of the transform coefficients included in the processing target block is a valid region (i.e., a valid transform coefficient region), decoding is possible even if signals related to QMs for invalid regions are not unnecessarily described in the stream. Therefore, since the QM can be used for rectangular blocks while reducing the amount of code in the header region, there is a high possibility of improving coding efficiency.

[0259] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0260] [First example of a method for generating a QM for a rectangular block in a modified version of the second example] Fig. 22 is a diagram illustrating a first example of generating a QM for a rectangular block from a QM for a square block in step S1002 in Fig. 20 and step S1102 in Fig. 21. Note that the processing described here is common to the encoding device 100 and the decoding device 200.

[0261] In Figure 22, square blocks ranging in size from 2x2 to 256x256 are shown corresponding to the size of the QM for a rectangular block generated from the QM for the square block of each size. In the example shown in Figure 22, the size of the block to be processed and the size of the valid transform coefficient area of ​​the block to be processed are shown. The numerical value in parentheses indicates the size of the valid transform coefficient area of ​​the block to be processed. Note that for rectangular blocks whose size of the block to be processed and the size of the valid transform coefficient area of ​​the block to be processed are the same as those in the first example described in Figure 13, the numerical value is omitted from the correspondence table shown in Figure 22.

[0262] Here, the length of the long side of each rectangular block is the same as the length of one side of the corresponding square block, and the rectangular blocks are smaller than the square blocks. In other words, the QM for the rectangular blocks is generated by down-converting the QM for the square blocks.

[0263] Note that FIG. 22 does not distinguish between luminance blocks and chrominance blocks, and shows the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each square block QM. The correspondence between QMs for square blocks and QMs for rectangular blocks that is adapted to the actual format used may be derived as appropriate. For example, in the 4:2:0 format, luminance blocks are twice the size of chrominance blocks. Therefore, when referring to luminance blocks in the process of generating a QM for a rectangular block from a QM for a square block, usable QMs for square blocks correspond to square blocks with sizes ranging from 4×4 to 256×256. In this case, only QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used as QMs for rectangular blocks generated from a QM for a square block. Furthermore, when referring to chrominance blocks in the process of generating a QM for a rectangular block from a QM for a square block, only QMs corresponding to rectangular blocks whose short side length is 2 or more and whose long side length is 128 or less are used as the QM for the rectangular block generated from the QM for the square block. Note that the same applies to the 4:4:4 format as explained in FIG. 13.

[0264] In this way, it is advisable to appropriately derive the correspondence between the QM for square blocks and the QM for rectangular blocks depending on the format actually used.

[0265] It should be noted that the size of the effective transform coefficient area shown in FIG. 22 is an example, and an effective transform coefficient area size other than the size exemplified in FIG. 22 may be used.

[0266] Note that the block sizes shown in Fig. 22 are merely examples and are not limiting. For example, block sizes other than those shown in Fig. 22 may be usable, or only some of the block sizes shown in Fig. 22 may be usable.

[0267] FIG. 23 is a diagram for explaining a method for generating the QM for the rectangular block described in FIG. 22 by down-converting from the corresponding QM for the square block.

[0268] In the example of FIG. 23, a QM corresponding to a 32×32 effective transform coefficient area in a 64×32 rectangular block is generated from a QM corresponding to a 32×32 effective transform coefficient area in a 64×64 square block.

[0269] First, as shown in (a) of Figure 23, a QM for an intermediate 64 x 64 square block having a 32 x 64 valid transform coefficient region is generated by extending a gradient in the vertical direction for the quantization coefficients of a QM corresponding to a 32 x 32 valid transform coefficient region. Methods for extending the gradient include, for example, extending the gradient so that the difference between the quantization coefficient on the 31st row and the quantization coefficient on the 32nd row becomes the difference between the adjacent coefficients thereafter, or deriving the amount of change between the difference between the quantization coefficient on the 30th row and the 31st row and the difference between the quantization coefficient on the 31st row and the 32nd row, and extending the difference between the adjacent quantization coefficients thereafter while correcting them with the amount of change.

[0270] Next, as shown in (b) of Figure 23, a QM for a 64x32 rectangular block is generated by down-converting the QM for the intermediate 64x64 square block, which has a 32x64 effective area, using a method similar to that described using Figure 14. In this case, the resulting effective area is the 32x32 area indicated by the diagonal lines in the QM for the 64x32 rectangular block in (c) of Figure 23.

[0271] Note that, although an example has been described here in which a QM for a square block having a valid area is down-converted in the vertical direction to generate a QM for a rectangular block, a method similar to the example in Figure 23 may also be used to down-convert a QM for a square block having a valid area in the horizontal direction to generate a QM for a rectangular block.

[0272] Note that, although an example has been described here in which a QM for a rectangular block is generated in two steps via a QM for an intermediate square block, it is also possible to generate a QM for a rectangular block directly from a QM for a square block with a valid area, without going through a QM for an intermediate square block, by using a transformation formula or the like that leads to processing results similar to those in the example of Figure 23.

[0273] [Second example of a method for generating a QM for a rectangular block in a modified version of the second example] Fig. 24 is a diagram illustrating a second example of generating a QM for a rectangular block from a QM for a square block in step S1002 in Fig. 20 and step S1102 in Fig. 21. Note that the processing described here is common to the encoding device 100 and the decoding device 200.

[0274] In Fig. 24, for square blocks ranging in size from 2x2 to 256x256, the size of the QM for a rectangular block generated from the QM for the square block of each size is shown in correspondence. In the example shown in Fig. 24, the size of the block to be processed and the size of the valid transform coefficient area of ​​the block to be processed are shown. The numerical value in parentheses indicates the size of the valid transform coefficient area of ​​the block to be processed. Note that for rectangular blocks whose size of the block to be processed and the size of the valid transform coefficient area are the same, the processing is the same as in the first example described in Fig. 15, so the numerical value is omitted in the correspondence table shown in Fig. 24.

[0275] Here, the length of the short side of each rectangular block is the same as the length of one side of the corresponding square block, and the rectangular blocks are larger than the square blocks. In other words, the QM for the rectangular blocks is generated by up-converting the QM for the square blocks.

[0276] Note that FIG. 24 does not distinguish between luminance blocks and chrominance blocks, and illustrates the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each square block QM. The correspondence between QMs for square blocks and QMs for rectangular blocks adapted to the actual format used may be derived as appropriate. For example, in the case of the 4:2:0 format, luminance blocks are twice the size of chrominance blocks. Therefore, when referring to luminance blocks in the process of generating a QM for a rectangular block from a QM for a square block, only QMs corresponding to rectangular blocks having a short side length of 4 or more and a long side length of 256 or less are used for the QM for the rectangular block generated from the QM for the square block. Furthermore, when referring to chrominance blocks in the process of generating a QM for a rectangular block from a QM for a square block, only QMs corresponding to rectangular blocks having a short side length of 2 or more and a long side length of 128 or less are used for the QM for the rectangular block generated from the QM for the square block. Note that the 4:4:4 format is the same as the content described in FIG. 13.

[0277] In this way, it is advisable to appropriately derive the correspondence between the QM for square blocks and the QM for rectangular blocks depending on the format actually used.

[0278] It should be noted that the size of the effective transform coefficient area shown in FIG. 24 is an example, and an effective transform coefficient area size other than the size exemplified in FIG. 24 may be used.

[0279] Note that the block sizes shown in Fig. 24 are merely examples and are not limiting. For example, block sizes other than those shown in Fig. 24 may be usable, or only some of the block sizes shown in Fig. 24 may be usable.

[0280] FIG. 25 is a diagram for explaining a method for generating the QM for the rectangular block described in FIG. 24 by up-converting from the QM for the corresponding square block.

[0281] In the example of Figure 25, a QM corresponding to a 32x32 effective transform coefficient region in a 64x32 rectangular block is generated from a QM corresponding to a 32x32 effective transform coefficient region in a 32x32 square block.

[0282] First, as shown in Figure 25(a), a QM for an intermediate 64x32 rectangular block is generated by up-converting a QM for a 32x32 square block using the same method as described with reference to Figure 16. At this time, the valid region is also up-converted to 64x32.

[0283] Next, as shown in (b) of Figure 25, by cutting out only the 32x32 area on the low-frequency side from the 64x32 effective area, a QM for a 64x32 rectangular block having a 32x32 effective area is generated.

[0284] Note that, although an example has been described here in which a QM for a square block is upconverted horizontally to generate a QM for a rectangular block, a method similar to that of the example in Figure 25 may also be used to upconvert a QM for a square block vertically to generate a QM for a rectangular block.

[0285] Note that, although an example has been described here in which a QM for a rectangular block is generated in two steps via a QM for an intermediate rectangular block, a QM for a rectangular block may also be generated directly from a QM for a square block using a transformation formula or the like that leads to processing results similar to those in the example of Figure 25, without going through a QM for an intermediate rectangular block.

[0286] [Other variations of the second example of the encoding process and the decoding process using the quantization matrix] The method for generating a QM for a rectangular block from a QM for a square block may be switched between the first example of the method for generating a QM for a rectangular block described using Figures 22 and 23 and the second example of the method for generating a QM for a rectangular block described using Figures 24 and 25, depending on the size of the rectangular block to be generated. For example, one method is to compare the ratio of the length and width of the rectangular block (the down-conversion or up-conversion ratio) with a threshold, and use the first example if it is greater than the threshold, and the second example if it is less than the threshold. Another method is to write a flag in the stream for each rectangular block size indicating whether to use the first or second example, and switch between them. This makes it possible to switch between down-conversion and up-conversion depending on the size of the rectangular block, thereby enabling the generation of a more appropriate QM for a rectangular block.

[0287] [Second Example of Encoding Process and Decoding Process Using Quantization Matrix and Effects of Modification of Second Example] According to the encoding device 100 and the decoding device 200 of the second aspect of the present disclosure, the configurations described with reference to Figures 17 and 18 enable encoding and decoding of rectangular blocks without needlessly describing signals related to QM of invalid areas in the stream for a target block having a block size in which only areas including some transform coefficients among a plurality of transform coefficients included in the target block are valid areas. Therefore, it is possible to reduce the amount of code in the header area, which increases the possibility of improving encoding efficiency.

[0288] Furthermore, according to the encoding device 100 and the decoding device 200 of the second modification of the present disclosure, the configurations described with reference to Figures 20 and 21 enable encoding and decoding of rectangular blocks even in encoding formats having rectangular blocks of various shapes by describing only QMs corresponding to square blocks in the stream, without describing QMs corresponding to rectangular blocks of each shape in the stream. In other words, according to the encoding device 100 and the decoding device 200 of the second modification of the present disclosure, a QM corresponding to a rectangular block can be generated from a QM corresponding to a square block, thereby enabling the use of an appropriate QM for the rectangular block while reducing the amount of code in the header region. Therefore, according to the encoding device 100 and the decoding device 200 of the second modification of the present disclosure, quantization can be efficiently performed on rectangular blocks of various shapes, which increases the likelihood of improving coding efficiency.

[0289] For example, the encoding device 100 is an encoding device that encodes moving images by performing quantization, and includes a circuit and a memory, and the circuit uses the memory to quantize only a plurality of transform coefficients within a predetermined range on the low-frequency side of a plurality of transform coefficients included in a block to be processed, using a quantization matrix.

[0290] As a result, the target block is quantized using only the quantization matrix corresponding to a predetermined range on the low-frequency side, which has a large impact on the visual sense, so the image quality of the moving image is less likely to deteriorate. Also, because only the quantization matrix corresponding to the predetermined range is encoded, the amount of code is reduced. Therefore, according to the encoding device 100, the image quality of the moving image is less likely to deteriorate and processing efficiency is improved.

[0291] For example, in the encoding device 100, the circuit may encode, into a bitstream, a signal related to the quantization matrix corresponding only to a plurality of transform coefficients within the predetermined range on the low frequency domain side.

[0292] This reduces the amount of coding, and therefore the encoding device 100 improves processing efficiency.

[0293] For example, in the encoding device 100, the block to be processed is a square block or a rectangular block, and the circuit may convert a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block to generate a second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block from the first quantization matrix as the quantization matrix, and encode only the first quantization matrix of the first and second quantization matrices into a bitstream as a signal related to the quantization matrix.

[0294] As a result, encoding device 100 generates a first quantization matrix corresponding to a predetermined range on the low-frequency side of a square block for processing target blocks of many shapes, including rectangular blocks, and encodes only the first quantization matrix, thereby reducing the amount of code. Furthermore, encoding device 100 generates a second quantization matrix corresponding to a predetermined range on the low-frequency side of a rectangular block from the first quantization matrix, thereby making it less likely that the image quality of moving images will deteriorate. Therefore, encoding device 100 makes it less likely that the image quality of moving images will deteriorate, and improves processing efficiency.

[0295] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing a down-conversion process from a first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having one side the same length as the long side of the rectangular block, which is the target block.

[0296] This allows the encoding device 100 to efficiently perform quantization on rectangular blocks.

[0297] For example, in the encoding device 100, in the down-conversion process, the circuit may extrapolate and extend the plurality of matrix elements of the first quantization matrix in a predetermined direction, divide the extended plurality of matrix elements of the first quantization matrix into groups the number of which is equal to the number of matrix elements of the second quantization matrix, and for each of the plurality of groups, determine the matrix element located at the lowest frequency side among the plurality of matrix elements included in the group, the matrix element located at the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group as the matrix element corresponding to the group in the second quantization matrix.

[0298] This allows the encoding device 100 to perform quantization on rectangular blocks more efficiently.

[0299] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing an upconversion process from a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the short side of the rectangular block that is the target block.

[0300] This allows the encoding device 100 to efficiently perform quantization on rectangular blocks.

[0301] For example, in the encoding device 100, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into groups the number of which is equal to the number of matrix elements of the first quantization matrix, and for each of the plurality of groups, extend the first quantization matrix in the predetermined direction by overlapping the plurality of matrix elements included in that group in the predetermined direction, and extract from the extended first quantization matrix the same number of matrix elements as the number of matrix elements of the second quantization matrix, or (ii) extend the first quantization matrix in the predetermined direction by performing linear interpolation between matrix elements of the second quantization matrix that are adjacent in the predetermined direction, and extract from the extended first quantization matrix the same number of matrix elements as the number of matrix elements of the second quantization matrix.

[0302] This allows the encoding device 100 to perform quantization on rectangular blocks more efficiently.

[0303] For example, in the encoding device 100, the circuit may generate the second quantization matrix for the multiple transform coefficients within the specified range in the rectangular block by switching between a method of performing the down-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the long side of the rectangular block, and a method of performing the up-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the length of the short side and the length of the long side of the rectangular block that is the target block to be processed.

[0304] This allows encoding device 100 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling generation of a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0305] In addition, the decoding device 200 is a decoding device that decodes moving images by performing inverse quantization, and is equipped with a circuit and a memory, and the circuit uses the memory to perform inverse quantization using a quantization matrix on only a plurality of quantization coefficients within a predetermined range on the low-frequency region side of a plurality of quantization coefficients included in a block to be processed.

[0306] As a result, the quantization of the target block is performed using only the quantization matrix corresponding to a predetermined range on the low-frequency side of the target block, which has a large impact on the visual sense, so that the image quality of the moving image is less likely to be degraded. Furthermore, because only the quantization matrix corresponding to the predetermined range is encoded, it is possible to reduce the amount of code. Therefore, according to the decoding device 200, the image quality of the moving image is less likely to be degraded and processing efficiency is improved.

[0307] For example, in the decoding device 200, the circuit may decode, from the bitstream, a signal relating to the quantization matrix corresponding to only a plurality of transform coefficients within the predetermined range on the low frequency domain side.

[0308] This allows the amount of coding to be reduced, and therefore the decoding device 200 improves processing efficiency.

[0309] For example, in the decoding device 200, the block to be processed is a square block or a rectangular block, and the circuit may convert a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block to generate a second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block from the first quantization matrix as the quantization matrix, and decode only the first quantization matrix of the first and second quantization matrices from the bitstream as a signal related to the quantization matrix.

[0310] As a result, decoding device 200 generates a first quantization matrix corresponding to a predetermined range on the low-frequency side of a square block for processing target blocks of many shapes, including rectangular blocks, and decodes only the first quantization matrix, thereby enabling a reduction in the amount of code. Furthermore, decoding device 200 generates a second quantization matrix corresponding to a predetermined range on the low-frequency side of a rectangular block from the first quantization matrix, thereby making it less likely that the image quality of moving images will deteriorate. Therefore, decoding device 200 makes it less likely that the image quality of moving images will deteriorate and improves processing efficiency.

[0311] For example, in the decoding device 200, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing a down-conversion process from a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the long side of the rectangular block, which is the target block.

[0312] This allows the decoding device 200 to efficiently perform inverse quantization on rectangular blocks.

[0313] For example, in the decoding device 200, in the down-conversion process, the circuit may extrapolate and extend the plurality of matrix elements of the first quantization matrix in a predetermined direction, divide the extended plurality of matrix elements of the first quantization matrix into groups the number of which is equal to the number of matrix elements of the second quantization matrix, and for each of the plurality of groups, determine the matrix element located at the lowest frequency side among the plurality of matrix elements included in the group, the matrix element located at the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group as the matrix element corresponding to the group in the second quantization matrix.

[0314] This allows the decoding device 200 to perform inverse quantization on rectangular blocks more efficiently.

[0315] For example, in the decoding device 200, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing an upconversion process from a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the short side of the rectangular block, which is the block to be processed.

[0316] This allows the decoding device 200 to efficiently perform inverse quantization on rectangular blocks.

[0317] For example, in the decoding device 200, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into groups whose number is equal to the number of the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, extend the first quantization matrix in the predetermined direction by overlapping the plurality of matrix elements included in that group in the predetermined direction, and extract from the extended first quantization matrix the same number of matrix elements as the number of the plurality of matrix elements of the second quantization matrix, or (ii) extend the first quantization matrix in the predetermined direction by performing linear interpolation between matrix elements of the second quantization matrix that are adjacent in the predetermined direction, and extract from the extended first quantization matrix the same number of matrix elements as the number of the plurality of matrix elements of the second quantization matrix.

[0318] This allows the decoding device 200 to perform inverse quantization on rectangular blocks more efficiently.

[0319] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the multiple transform coefficients within the specified range in the rectangular block by switching between a method of performing the down-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the long side of the rectangular block, and a method of performing the up-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the length of the short side and the length of the long side of the rectangular block that is the target block to be processed.

[0320] This allows the decoding device 200 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling it to generate a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0321] In addition, the encoding method is an encoding method for encoding moving images by performing quantization, and quantization is performed using a quantization matrix on only a plurality of transform coefficients within a predetermined range on the low-frequency region side of a plurality of transform coefficients included in a block to be processed.

[0322] As a result, the quantization of the target block is performed using only the quantization matrix corresponding to a predetermined range on the low-frequency side, which has a large impact on the visual sense, so the image quality of the moving image is less likely to deteriorate. Also, because only the quantization matrix corresponding to the predetermined range is encoded, the amount of code is reduced. Therefore, according to the encoding method, the image quality of the moving image is less likely to deteriorate and processing efficiency is improved.

[0323] In addition, the decoding method is a decoding method for decoding a moving image by performing inverse quantization, and performs inverse quantization using a quantization matrix on only a plurality of quantization coefficients within a predetermined range on the low-frequency region side of a plurality of quantization coefficients included in a block to be processed.

[0324] As a result, the inverse quantization of the target block is performed using only the quantization matrix corresponding to a predetermined range on the low-frequency side of the target block, which has a large impact on the visual sense, so the image quality of the moving image is less likely to deteriorate. Also, because only the quantization matrix corresponding to the predetermined range is decoded, it is possible to reduce the amount of code. Therefore, according to the decoding method, the image quality of the moving image is less likely to deteriorate and processing efficiency is improved.

[0325] This aspect may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some of the processes, some of the device configurations, and some of the syntax described in the flowcharts of this aspect may be implemented in combination with other aspects.

[0326] [Third aspect] The encoding device 100, the decoding device 200, the encoding method, and the decoding method according to the third aspect of the present disclosure will be described below.

[0327] [Third example of encoding and decoding processes using quantization matrices] 26 is a diagram showing a third example of the encoding process flow using a quantization matrix (QM) in the encoding device 100. Note that the encoding device 100 described here performs encoding processing for each square or rectangular block obtained by dividing a screen.

[0328] First, in step S1601, the quantization unit 108 generates a QM corresponding to the diagonal components of the target block (hereinafter also referred to as a QM of only diagonal components), and generates a QM corresponding to the target block from the quantization coefficient values ​​of the QM of only diagonal components for each block size of the target block, which may be a square block or a rectangular block of various shapes, using a common method described below. In other words, the quantization unit 108 generates a quantization matrix for the target block from the diagonal components of the quantization matrix for multiple transform coefficients included in the target block that are consecutively arranged in the diagonal direction of the target block. Note that using a common method means using a common method for all target blocks regardless of the shape and size of the blocks. The diagonal components are, for example, multiple coefficients on a diagonal line extending from the low-pass side to the high-pass side of the target block.

[0329] The entropy coding unit 110 writes signals related to the QM of only the diagonal components generated in step S1601 into a stream. In other words, the entropy coding unit 110 encodes signals related to the diagonal components of the quantization matrix into a bitstream.

[0330] The quantization unit 108 may generate the quantization coefficient values ​​of the QM for only diagonal components from values ​​defined by a user and set in the encoding device 100, or may adaptively generate the quantization coefficients using encoding information of a picture that has already been encoded. The QM for only diagonal components may be coded in the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. The QM for only diagonal components does not need to be described in the stream. In this case, the quantization unit 108 may use a default value predefined by a standard as the QM for only diagonal components.

[0331] The processing of step S1601 may be configured to be performed all at once at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processing every time during block-by-block processing. Furthermore, the QM generated in step S1601 may be configured to generate multiple types of QMs for blocks of the same block size, such as for luminance blocks / chrominance blocks, for intra-prediction blocks / inter-prediction blocks, and for other conditions.

[0332] In the processing flow shown in FIG. 26, the processing steps other than step S1601 are loop processing in units of blocks, and are the same as the processing in the first example described with reference to FIG.

[0333] This makes it possible to encode a target block by describing only the quantization coefficients of the QM of the diagonal components in the stream, without describing all of the quantization coefficients of the QM of the target block of each block size in the stream. Therefore, even in encoding methods that use blocks of many shapes, including rectangular blocks, it is possible to generate and use a QM corresponding to the target block without significantly increasing the amount of code in the header region, which increases the possibility of improving encoding efficiency.

[0334] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0335] Fig. 27 is a diagram showing an example of a decoding process flow using a quantization matrix (QM) in a decoding device 200 corresponding to the encoding device 100 described in Fig. 26. Note that the decoding device 200 described here performs decoding processing for each square or rectangular block obtained by dividing a screen.

[0336] First, in step S1701, the entropy decoding unit 202 decodes a signal related to a QM of only diagonal components from the stream, and generates QMs corresponding to various block sizes of target blocks of various shapes, such as square blocks and rectangular blocks, using the decoded signal related to the QM of only diagonal components, using a common method described below. Note that the QM of only diagonal components may be decoded from the sequence header region, picture header region, slice header region, auxiliary information region, or region storing other parameters of the stream. Alternatively, the QM of only diagonal components may not be decoded from the stream. In this case, for example, a default value predefined in a standard may be used as the QM of only diagonal components.

[0337] The processing of step S1701 may be configured to be performed all at once at the start of sequence processing, picture processing, or slice processing, or may be configured to perform some of the processing every time during block-by-block processing. Furthermore, the QM generated by the entropy decoding unit 202 in step S1701 may be configured to generate multiple types of QMs for blocks of the same block size, such as for luminance blocks / chrominance blocks, for intra-prediction blocks / inter-prediction blocks, and for other conditions.

[0338] In the processing flow shown in FIG. 27, the processing flow other than step S1701 is a loop process in units of blocks, and is the same as the processing flow of the first example described using FIG.

[0339] This makes it possible to decode the block to be processed as long as only the QM quantization coefficients of the diagonal components of the block to be processed are described in the stream, even if all of the QM quantization coefficients of each block size of the block to be processed are not described in the stream. Therefore, it is possible to reduce the amount of code in the header area, which increases the possibility of improving processing efficiency.

[0340] Note that this processing flow is an example, and the order of the processes described may be changed, some of the processes described may be omitted, or processes not described may be added.

[0341] Fig. 28 is a diagram illustrating an example of a method for generating a QM for a current block from quantization coefficient values ​​of QMs of only diagonal components for each block size using a common method described below in step S1601 in Fig. 26 and step S1701 in Fig. 27. Note that the process described here is common to the encoding device 100 and the decoding device 200.

[0342] The encoding device 100 and the decoding device 200 according to the third aspect generate a quantization matrix (QM) for a block to be processed by overlapping each of a plurality of matrix elements of the diagonal components of the block to be processed in the horizontal and vertical directions. More specifically, the encoding device 100 and the decoding device 200 generate a QM for the block to be processed by simply extending the values ​​of the quantization coefficients of the QM of the diagonal components upward and leftward, that is, by arranging the same values ​​consecutively.

[0343] Here, we have explained an example of how to generate a QM for a block to be processed when the block to be processed is a square block. However, when the block to be processed is a rectangular block, the QM for the block to be processed may also be generated from the values ​​of the quantization coefficients of the QM of the diagonal components, as in the example of Figure 28.

[0344] Fig. 29 is a diagram illustrating another example of a method for generating a QM for a current block from quantization coefficient values ​​of only diagonal components of the QM for each block size of the current block, using a common method described below, in step S1601 in Fig. 26 and step S1701 in Fig. 27. Note that the process described here is common to the encoding device 100 and the decoding device 200.

[0345] The encoding device 100 and the decoding device 200 according to the third aspect may generate a quantization matrix for a current block by overlapping each of a plurality of matrix elements of the diagonal components of the current block in a diagonal direction. More specifically, the encoding device 100 and the decoding device 200 generate a QM for the current block by simply stretching the values ​​of the quantization coefficients of the QM for the diagonal components in the lower left and upper right directions, that is, by arranging the same values ​​consecutively.

[0346] In this case, the encoding device 100 and the decoding device 200 may generate a QM for the block to be processed by using not only the quantization coefficients of the diagonal elements but also the quantization coefficients of the neighboring elements of the diagonal elements and overlapping each quantization coefficient in a diagonal direction. In other words, the quantization matrix (QM) for the block to be processed may be generated from multiple matrix elements of the diagonal elements and matrix elements located near the diagonal elements. As a result, even if it is difficult to fill all the quantization coefficients of the block to be processed using only the diagonal elements, all the quantization coefficients can be filled by using the quantization coefficients of the neighboring elements. Note that the neighboring components of the diagonal elements are, for example, components adjacent to any of multiple coefficients on the diagonal line extending from the low-pass side to the high-pass side of the block to be processed.

[0347] For example, the quantization coefficients of neighboring components of the diagonal components are quantization coefficients at positions shown in Fig. 29. The encoding device 100 and the decoding device 200 may set the quantization coefficients of neighboring components of the diagonal components by encoding the signals related to the quantization coefficients of the neighboring components into a stream and decoding the stream, or may derive and set the quantization coefficients by interpolating the values ​​of neighboring quantization coefficients of the diagonal components using linear interpolation or the like without encoding the signals into a stream and decoding them from the stream.

[0348] Here, we have explained an example of a method for generating a QM for a block to be processed when the block to be processed is a square block, but when the block to be processed is a rectangular block, the QM for the block to be processed may also be generated from the values ​​of the quantization coefficients of the QM of the diagonal components, as in the example of Figure 29.

[0349] [Other variations of the third example of encoding and decoding processes using quantization matrices] 28 and 29, a method for generating QM quantization coefficients for the entire block to be processed from the QM quantization coefficients of the diagonal components has been described, but a configuration in which only some of the QM coefficients for the block to be processed are generated from the QM quantization coefficients of the diagonal components may also be used. For example, for a QM corresponding to the low-frequency region of the block to be processed, all quantization coefficients included in the QM may be coded into a stream and decoded from the stream, and only the QM quantization coefficients corresponding to the mid-frequency and high-frequency regions of the block to be processed may be generated from the QM quantization coefficients of the diagonal components of the block to be processed.

[0350] Note that the method for generating a QM for a rectangular block from a QM for a square block may be a combination of the third example described with reference to Figures 26 and 27 and the first example described with reference to Figures 11 and 12. For example, the QM for a square block may be generated using one of the two common methods described above from the quantization coefficient values ​​of the QM for only the diagonal components of the square block as described in the third example, and the QM for a rectangular block may be generated using the QM for a square block generated as described in the first example.

[0351] Note that the method for generating a QM for a rectangular block from a QM for a square block may be a combination of the third example described with reference to Figures 26 and 27 and the second example described with reference to Figures 17 and 18. For example, a QM corresponding to the size of the valid transform coefficient area for each block size of the block to be processed may be generated from the quantized coefficient values ​​of the QM for only the diagonal components of the valid transform coefficient area using one of the common methods described above.

[0352] [Effect of the third example of encoding and decoding processes using quantization matrices] According to the encoding device 100 and the decoding device 200 according to the third aspect of the present disclosure, with the configurations described with reference to Figures 26 and 27, even if all of the QM quantization coefficients of each block size of the block to be processed are not described in the stream, as long as only the QM quantization coefficients of the diagonal components of the block to be processed are described in the stream, encoding and decoding of the block to be processed are possible. Therefore, it is possible to reduce the amount of code in the header region, which increases the possibility of improving encoding efficiency.

[0353] For example, the encoding device 100 is an encoding device that encodes moving images by performing quantization, and includes a circuit and a memory. The circuit uses the memory to generate a quantization matrix for a target block from diagonal components of a quantization matrix for a plurality of transform coefficients included in the target block that are arranged consecutively in a diagonal direction in the target block, and encodes signals related to the diagonal components into a bitstream.

[0354] As a result, even in a coding method using blocks of various shapes, including rectangular blocks, only the diagonal components of the current block are coded, thereby reducing the amount of coding. Therefore, the coding device 100 improves processing efficiency.

[0355] For example, in the encoding device 100, the circuit may generate the quantization matrix by overlapping each of the plurality of matrix elements of the diagonal components in the horizontal and vertical directions.

[0356] As a result, the quantization coefficients included in the target block are generated from the quantization coefficients of the diagonal components of the target block, so it is not necessary to encode the quantization matrices of the entire target block. This reduces the amount of coding, improving processing efficiency. Therefore, the encoding device 100 can efficiently quantize the target block.

[0357] For example, in the encoding device 100, the circuit may generate the quantization matrix by overlapping each of the plurality of matrix elements of the diagonal components in a diagonal direction.

[0358] As a result, the quantization coefficients included in the target block are generated from the quantization coefficients of the diagonal components of the target block, so it is not necessary to encode the quantization matrices of the entire target block. This reduces the amount of coding, improving processing efficiency. Therefore, the encoding device 100 can efficiently quantize the target block.

[0359] For example, in the encoding device 100, the circuit may generate the quantization matrix from a plurality of matrix elements of the diagonal elements and a plurality of matrix elements located in the vicinity of the diagonal elements.

[0360] As a result, the encoding device 100 generates the quantization coefficients included in the target block from the quantization coefficients of the diagonal components of the target block and its neighboring components, thereby making it possible to generate a more appropriate quantization matrix corresponding to the target block.

[0361] For example, in the encoding device 100, the circuit may not encode signals related to a plurality of matrix elements located near the diagonal element into a bit stream, but may generate the plurality of matrix elements located near the diagonal element by interpolating from a plurality of adjacent matrix elements among the plurality of matrix elements of the diagonal element.

[0362] This reduces the amount of coding and improves processing efficiency, allowing the encoding device 100 to efficiently quantize the current block.

[0363] For example, in the encoding device 100, the circuit may quantize a plurality of transform coefficients included in the target block that fall within a predetermined range on the low-frequency side using a quantization matrix, encode signals related to the quantization matrix corresponding only to the plurality of transform coefficients included in the predetermined range on the low-frequency side into a bitstream, and quantize a plurality of transform coefficients included in the target block that fall within a range other than the predetermined range on the low-frequency side using a plurality of matrix elements of the diagonal components.

[0364] As a result, all quantization coefficients within a predetermined range on the low-frequency side of the target block, which has a large impact on vision, are coded, making it difficult for the image quality of the moving image to deteriorate. Furthermore, for multiple transform coefficients outside the predetermined range on the low-frequency side of the target block, quantization is performed using a quantization matrix of the diagonal components of the target block, reducing the amount of code and improving processing efficiency. Therefore, according to the coding device 100, the image quality of the moving image is difficult to deteriorate and processing efficiency is improved.

[0365] For example, in the encoding device 100, the quantization matrix may be a quantization matrix that corresponds to only a plurality of transform coefficients within the specified range on the low-frequency side among a plurality of transform coefficients included in the block to be processed.

[0366] As a result, according to the encoding device 100, a quantization matrix corresponding to a predetermined range on the low-frequency side, which has a large impact on vision, is generated in the block to be processed, so that the image quality of the moving image is less likely to deteriorate.

[0367] For example, in the encoding device 100, the block to be processed is a square block or a rectangular block, and the circuit may convert a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block to generate a second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block from the first quantization matrix, and encode only the first quantization matrix of the first and second quantization matrices into a bitstream as a signal related to the quantization matrix.

[0368] As a result, only the first quantization matrix corresponding to a predetermined range in the square block is coded, thereby reducing the amount of code. Also, since a second quantization matrix corresponding to a predetermined range in the rectangular block is generated from the first quantization matrix, the image quality of the moving image is less likely to deteriorate. Note that the predetermined range is located in the low-frequency region, which has a large impact on the visual sense. Therefore, according to coding device 100, the image quality of the moving image is less likely to deteriorate and processing efficiency is improved.

[0369] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing a down-conversion process from the first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the long side of the rectangular block that is the target block.

[0370] This allows encoding device 100 to efficiently generate a second quantization matrix that corresponds to a predetermined range in a rectangular block.

[0371] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing an upconversion process from the first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the short side of the rectangular block that is the target block.

[0372] This allows encoding device 100 to efficiently generate a second quantization matrix that corresponds to a predetermined range in a rectangular block.

[0373] For example, in the encoding device 100, the circuit may generate the second quantization matrix for the multiple transform coefficients within the specified range in the rectangular block by switching between a method of performing the down-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the long side of the rectangular block, and a method of performing the up-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the length of the short side and the length of the long side of the rectangular block that is the target block to be processed.

[0374] This allows the encoding device 100 to switch between down-conversion and up-conversion, thereby generating a more appropriate second quantization matrix corresponding to a specified range of rectangular blocks from a first quantization matrix corresponding to a specified range of square blocks.

[0375] Furthermore, the decoding device 200 is a decoding device that decodes moving images by performing inverse quantization, and includes a circuit and a memory. The circuit uses the memory to generate a quantization matrix for a target block from diagonal components of a quantization matrix for a plurality of quantization coefficients included in the target block that are arranged consecutively in a diagonal direction in the target block, and decodes signals related to the diagonal components from the bitstream.

[0376] As a result, even in a decoding method using blocks of various shapes, including rectangular blocks, only the diagonal components of the current block are decoded, which reduces the amount of coding. Therefore, the decoding device 200 improves processing efficiency.

[0377] For example, in the decoding device 200, the circuit may generate the quantization matrix by overlapping each of the plurality of matrix elements of the diagonal components in the horizontal and vertical directions.

[0378] As a result, the quantization coefficients included in the target block are generated from the quantization coefficients of the diagonal components of the target block, so it is not necessary to decode the quantization matrices of the entire target block. This makes it possible to reduce the amount of code, thereby improving processing efficiency. Therefore, the decoding device 200 can efficiently perform inverse quantization on the target block.

[0379] For example, in the decoding device 200, the circuit may generate the quantization matrix by overlapping each of the plurality of matrix elements of the diagonal components in a diagonal direction.

[0380] As a result, the quantization coefficients included in the target block are generated from the quantization coefficients of the diagonal components of the target block, so it is not necessary to decode the quantization matrices of the entire target block. This makes it possible to reduce the amount of code, thereby improving processing efficiency. Therefore, the decoding device 200 can efficiently perform inverse quantization on the target block.

[0381] For example, in the decoding device 200, the circuit may generate the quantization matrix from a plurality of matrix elements of the diagonal components and a plurality of matrix elements located in the vicinity of the diagonal components.

[0382] As a result, the decoding device 200 generates the quantization coefficients included in the target block from the quantization coefficients of the diagonal components of the target block and its neighboring components, thereby making it possible to generate a more appropriate quantization matrix corresponding to the target block.

[0383] For example, in the decoding device 200, the circuit may not decode signals related to multiple matrix elements located near the diagonal elements from the bit stream, but may generate the multiple matrix elements located near the diagonal elements by interpolating from multiple adjacent matrix elements among the multiple matrix elements of the diagonal elements.

[0384] This allows the amount of coding to be reduced, thereby improving processing efficiency, and therefore the decoding device 200 can efficiently perform inverse quantization on the current block.

[0385] For example, in the decoding device 200, the circuit may perform inverse quantization using a quantization matrix on a plurality of quantization coefficients included in the target block that fall within a predetermined range on the low-frequency side, decode a signal related to the quantization matrix corresponding only to a plurality of transform coefficients included in the predetermined range on the low-frequency side from the bitstream, and perform inverse quantization using a plurality of matrix elements of the diagonal components on a plurality of quantization coefficients included in the target block that fall within a range other than the predetermined range on the low-frequency side.

[0386] As a result, all quantization coefficients within a predetermined range on the low-frequency side of the target block, which has a large impact on vision, are decoded, making it difficult for the image quality of the moving image to deteriorate. Furthermore, for multiple quantization coefficients outside the predetermined range on the low-frequency side of the target block, inverse quantization is performed using a quantization matrix of the diagonal components of the target block, making it possible to reduce the amount of code and improving processing efficiency. Therefore, according to the decoding device 200, the image quality of the moving image is difficult to deteriorate and processing efficiency is improved.

[0387] For example, in the decoding device 200, the quantization matrix may be a quantization matrix that corresponds to only a plurality of transform coefficients within the specified range on the low frequency region side among a plurality of transform coefficients included in the block to be processed.

[0388] As a result, according to the decoding device 200, a quantization matrix corresponding to a predetermined range on the low frequency side, which has a large impact on vision, is generated in the block to be processed, so that the image quality of the moving image is less likely to deteriorate.

[0389] For example, in the decoding device 200, the block to be processed is a square block or a rectangular block, and the circuit may convert a first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block to generate a second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block from the first quantization matrix, and decode only the first quantization matrix of the first and second quantization matrices from the bitstream as a signal related to the quantization matrix.

[0390] This allows for a reduction in the amount of code, since only the first quantization matrix corresponding to a predetermined range in the square block is decoded. Furthermore, since a second quantization matrix corresponding to a predetermined range in the rectangular block is generated from the first quantization matrix, the image quality of the moving image is less likely to be degraded. The predetermined range is located in the low-frequency region, which has a large impact on the visual sense. Therefore, according to decoding device 200, the image quality of the moving image is less likely to be degraded, and processing efficiency is improved.

[0391] For example, in the decoding device 200, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing a down-conversion process from the first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the long side of the rectangular block that is the target block.

[0392] This allows decoding device 200 to efficiently generate a second quantization matrix that corresponds to a predetermined range in a rectangular block.

[0393] For example, in the decoding device 200, the circuit may generate the second quantization matrix for a plurality of transform coefficients within the predetermined range in the rectangular block by performing an upconversion process from the first quantization matrix for a plurality of transform coefficients within the predetermined range in the square block having one side the same length as the short side of the rectangular block that is the target block.

[0394] This allows decoding device 200 to efficiently generate a second quantization matrix that corresponds to a predetermined range in a rectangular block.

[0395] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the multiple transform coefficients within the specified range in the rectangular block by switching between a method of performing the down-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the long side of the rectangular block, and a method of performing the up-conversion process from the first quantization matrix for the multiple transform coefficients within the specified range in the square block having one side the same length as the short side of the rectangular block, depending on the ratio between the length of the short side and the length of the long side of the rectangular block that is the target block to be processed.

[0396] This allows the decoding device 200 to switch between down-conversion and up-conversion to generate a more appropriate second quantization matrix corresponding to a specified range of rectangular blocks from a first quantization matrix corresponding to a specified range of square blocks.

[0397] The encoding method is also an encoding method for encoding a moving image by performing quantization, in which a quantization matrix for a target block is generated from diagonal components of a quantization matrix for a plurality of transform coefficients included in the target block that are arranged consecutively in a diagonal direction in the target block, and signals related to the diagonal components are encoded into a bitstream.

[0398] This reduces the amount of code even in coding schemes that use blocks of various shapes, including rectangular blocks, because only the diagonal components of the block being processed are coded. Therefore, the coding method improves processing efficiency.

[0399] The decoding method is a decoding method for decoding a moving image by performing inverse quantization, in which a quantization matrix for a target block is generated from diagonal components of a quantization matrix for a plurality of quantization coefficients included in the target block that are arranged consecutively in a diagonal direction in the target block, and a signal related to the diagonal components is decoded from a bitstream.

[0400] This reduces the amount of information even in a decoding method using blocks of various shapes, including rectangular blocks, because only the diagonal components of the block being processed are decoded. Therefore, the decoding method improves processing efficiency.

[0401] This aspect may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some of the processes, some of the device configurations, and some of the syntax described in the flowcharts of this aspect may be implemented in combination with other aspects.

[0402] [Implementation example] 30 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a circuit 160 and a memory 162. For example, several components of the encoding device 100 shown in FIG. 1 are implemented by the circuit 160 and memory 162 shown in FIG.

[0403] The circuit 160 is an electronic circuit that can access the memory 162 and performs information processing. For example, the circuit 160 is a dedicated or general-purpose electronic circuit that encodes moving images using the memory 162. The circuit 160 may be a processor such as a CPU. Alternatively, the circuit 160 may be a collection of multiple electronic circuits.

[0404] Also, for example, circuit 160 may function as multiple components of encoding device 100 shown in Fig. 1, excluding the components for storing information. That is, circuit 160 may perform the operations described above as the operations of these components.

[0405] The memory 162 is a dedicated or general-purpose memory that stores information for encoding video by the circuit 160. The memory 162 may be an electronic circuit, connected to the circuit 160, or included in the circuit 160.

[0406] Furthermore, memory 162 may be a collection of multiple electronic circuits or may be composed of multiple sub-memories. Furthermore, memory 162 may be a magnetic disk, an optical disk, or the like, and may be expressed as a storage or a recording medium, etc. Furthermore, memory 162 may be a non-volatile memory or a volatile memory.

[0407] For example, the memory 162 may serve as a component for storing information among the multiple components of the encoding device 100 shown in Fig. 1. Specifically, the memory 162 may serve as the block memory 118 and the frame memory 122 shown in Fig. 1.

[0408] The memory 162 may store a video to be encoded, or a bit string corresponding to the encoded video, and may also store a program for the circuit 160 to encode the video.

[0409] Note that not all of the components shown in Fig. 1 need to be implemented in the encoding device 100, and not all of the processes described above need to be performed. Some of the components shown in Fig. 1 may be included in another device, and some of the processes described above may be executed by another device. Then, by implementing some of the components shown in Fig. 1 and performing some of the processes described above in the encoding device 100, a prediction sample set can be appropriately derived.

[0410] Fig. 31 is a flowchart showing an example of the operation of the encoding device 100 shown in Fig. 30. For example, the encoding device 100 shown in Fig. 30 performs the operation shown in Fig. 31 when encoding a moving image. Specifically, the circuit 160 performs the following operation using the memory 162.

[0411] First, the circuit 160 generates a second quantization matrix for a plurality of transform coefficients of a rectangular block from the first quantization matrix by converting the first quantization matrix for a plurality of transform coefficients of a square block (step S201). Next, the circuit 160 quantizes the plurality of transform coefficients of the rectangular block using the second quantization matrix (step S202).

[0412] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be coded. This reduces the amount of coding, improving processing efficiency. Therefore, according to the coding device 100, rectangular blocks can be efficiently quantized.

[0413] For example, the circuit 160 may encode only the first quantization matrix of the first and second quantization matrices into a bitstream.

[0414] This reduces the amount of coding, and therefore the encoding device 100 improves processing efficiency.

[0415] For example, circuit 160 may generate a second quantization matrix for multiple transform coefficients of a rectangular block by performing a down-conversion process from a first quantization matrix for multiple transform coefficients of a square block having one side the same length as the long side of the rectangular block being processed.

[0416] This allows encoding device 100 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the long side of the rectangular block.

[0417] For example, in the down-conversion process, circuit 160 may divide the plurality of matrix elements of the first quantization matrix into groups of the same number as the plurality of matrix elements of the second quantization matrix, and for each of the plurality of groups, the plurality of matrix elements included in the group may be arranged consecutively in the horizontal or vertical direction of a square block, and for each of the plurality of groups, the matrix element located on the lowest frequency side among the plurality of matrix elements included in the group, the matrix element located on the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group may be determined to be the matrix element corresponding to the group in the second quantization matrix.

[0418] This allows encoding device 100 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0419] For example, circuit 160 may generate a second quantization matrix for multiple transform coefficients of a rectangular block by performing an upconversion process from a first quantization matrix for multiple transform coefficients of a square block having one side the same length as the short side of the rectangular block being processed.

[0420] This allows encoding device 100 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having a side with the same length as the shorter side of the rectangular block.

[0421] For example, in the upconversion process, circuit 160 may (i) divide the multiple matrix elements of the second quantization matrix into groups the same number as the number of multiple matrix elements of the first quantization matrix, and for each of the multiple groups, determine the matrix elements in the second quantization matrix corresponding to that group by overlapping the multiple matrix elements included in that group, or (ii) determine the multiple matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements among the multiple matrix elements of the second quantization matrix.

[0422] This allows encoding device 100 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0423] For example, circuit 160 may generate the second quantization matrix for multiple rectangular transform coefficients by switching between a method of performing a down-conversion process from a first quantization matrix for multiple transform coefficients in a square block having one side the same length as the long side of the rectangular block, and a method of performing an up-conversion process from a first quantization matrix for multiple transform coefficients in a square block having one side the same length as the short side of the rectangular block, depending on the ratio between the lengths of the short sides of the rectangular block that is the block to be processed.

[0424] This allows encoding device 100 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling generation of a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0425] 32 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a circuit 260 and a memory 262. For example, several components of the decoding device 200 shown in FIG. 10 are implemented by the circuit 260 and memory 262 shown in FIG.

[0426] The circuit 260 is an electronic circuit that can access the memory 262 and performs information processing. For example, the circuit 260 is a dedicated or general-purpose electronic circuit that decodes video using the memory 262. The circuit 260 may be a processor such as a CPU. Alternatively, the circuit 260 may be a collection of multiple electronic circuits.

[0427] Also, for example, the circuit 260 may function as multiple components of the decoding device 200 shown in Fig. 10 , excluding the components for storing information. That is, the circuit 260 may perform the operations described above as the operations of these components.

[0428] The memory 262 is a dedicated or general-purpose memory that stores information for decoding video by the circuit 260. The memory 262 may be an electronic circuit, connected to the circuit 260, or included in the circuit 260.

[0429] Furthermore, the memory 262 may be a collection of multiple electronic circuits, or may be composed of multiple sub-memories. Furthermore, the memory 262 may be a magnetic disk, an optical disk, or the like, and may be expressed as a storage or a recording medium, etc. Furthermore, the memory 262 may be a non-volatile memory or a volatile memory.

[0430] For example, the memory 262 may serve as a component for storing information among the multiple components of the decoding device 200 shown in Fig. 10. Specifically, the memory 262 may serve as the block memory 210 and the frame memory 214 shown in Fig. 10.

[0431] The memory 262 may store a bit string corresponding to an encoded video, or may store a decoded video, or may store a program for the circuit 260 to decode the video.

[0432] Note that decoding device 200 does not necessarily have to implement all of the components shown in Fig. 10 , and does not necessarily have to perform all of the above-described processes. Some of the components shown in Fig. 10 may be included in another device, and some of the above-described processes may be executed by another device. Then, decoding device 200 may implement some of the components shown in Fig. 10 and perform some of the above-described processes, thereby enabling a prediction sample set to be appropriately derived.

[0433] Fig. 33 is a flowchart showing an example of the operation of the decoding device 200 shown in Fig. 32. For example, the decoding device 200 shown in Fig. 32 performs the operation shown in Fig. 33 when decoding a video. Specifically, the circuit 260 uses the memory 262 to perform the following operation.

[0434] First, the circuit 260 generates a second quantization matrix for a plurality of transform coefficients of a rectangular block from the first quantization matrix by transforming the first quantization matrix for a plurality of transform coefficients of a square block (step S301). Next, the circuit 260 inverse-quantizes the plurality of transform coefficients of the rectangular block using the second quantization matrix (step S302).

[0435] As a result, a quantization matrix corresponding to a rectangular block is generated from a quantization matrix corresponding to a square block, so the quantization matrix corresponding to the rectangular block does not need to be decoded. This allows for a reduction in the amount of code, improving processing efficiency. Therefore, the decoding device 200 can efficiently perform inverse quantization on rectangular blocks.

[0436] For example, the circuit 260 may decode only the first quantization matrix of the first and second quantization matrices from the bitstream.

[0437] This allows the amount of coding to be reduced, and therefore the decoding device 200 improves processing efficiency.

[0438] For example, the circuit 260 may generate a second quantization matrix for multiple transform coefficients of a rectangular block by performing a down-conversion process from a first quantization matrix for multiple transform coefficients of a square block having one side the same length as the long side of the rectangular block being processed.

[0439] This allows the decoding device 200 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the long side of the rectangular block.

[0440] For example, in the down-conversion process, circuit 260 may divide the plurality of matrix elements of the first quantization matrix into groups of the same number as the number of matrix elements of the second quantization matrix, and for each of the plurality of groups, determine the matrix element located at the lowest frequency side among the plurality of matrix elements included in the group, the transformation coefficient located at the highest frequency side among the plurality of matrix elements included in the group, or the average value of the plurality of matrix elements included in the group as the matrix element corresponding to that group in the second quantization matrix.

[0441] This allows the decoding device 200 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0442] For example, the circuit 260 may generate a second quantization matrix for a plurality of transform coefficients in a rectangular block by performing an upconversion process from a first quantization matrix for a plurality of transform coefficients in a square block having one side the same length as the short side of the rectangular block being processed.

[0443] This allows the decoding device 200 to efficiently generate a quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block having one side with the same length as the short side of the rectangular block.

[0444] For example, in the upconversion process, the circuit 260 may (i) divide the plurality of matrix elements of the second quantization matrix into groups the number of which is equal to the number of the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix elements in the second quantization matrix corresponding to that group by overlapping the plurality of matrix elements included in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements among the plurality of matrix elements of the second quantization matrix.

[0445] This allows the decoding device 200 to more efficiently generate quantization matrices corresponding to rectangular blocks from quantization matrices corresponding to square blocks.

[0446] For example, the circuit 260 may generate the second quantization matrix for rectangular transform coefficients by switching between a method of performing a down-conversion process from a first quantization matrix for multiple transform coefficients of a square block having one side the same length as the long side of the rectangular block, and a method of performing an up-conversion process from a first quantization matrix for multiple transform coefficients of a square block having one side the same length as the short side of the rectangular block, depending on the ratio between the lengths of the short sides of the rectangular block that is the block to be processed.

[0447] This allows the decoding device 200 to switch between down-conversion and up-conversion depending on the block size of the block to be processed, thereby enabling it to generate a more appropriate quantization matrix corresponding to a rectangular block from a quantization matrix corresponding to a square block.

[0448] Furthermore, each component may be a circuit, as described above. These circuits may form a single circuit as a whole, or each may be a separate circuit. Furthermore, each component may be realized by a general-purpose processor or a dedicated processor.

[0449] Furthermore, a process performed by a specific component may be performed by another component. The order in which the processes are performed may be changed, or multiple processes may be performed in parallel. Furthermore, the encoding / decoding device may include the encoding device 100 and the decoding device 200.

[0450] Furthermore, the ordinal numbers such as first and second used in the description may be changed as appropriate. Furthermore, new ordinal numbers may be assigned to components or removed.

[0451] Although aspects of the encoding device 100 and the decoding device 200 have been described above based on the embodiments, the aspects of the encoding device 100 and the decoding device 200 are not limited to these embodiments. As long as they do not deviate from the spirit of this disclosure, various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining components of different embodiments may also be included within the scope of the aspects of the encoding device 100 and the decoding device 200.

[0452] This aspect may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some of the processes, some of the device configurations, and some of the syntax described in the flowcharts of this aspect may be implemented in combination with other aspects.

[0453] (Embodiment 2) In each of the above embodiments, each of the functional blocks can typically be realized by an MPU, memory, etc. Furthermore, the processing by each of the functional blocks is typically realized by a program execution unit such as a processor reading and executing software (programs) recorded on a recording medium such as a ROM. The software may be distributed by downloading, etc., or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, each functional block can also be realized by hardware (dedicated circuits).

[0454] Furthermore, the processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. The processor that executes the program may be a single processor or multiple processors. That is, centralized processing or distributed processing may be performed.

[0455] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.

[0456] Furthermore, here, we will explain application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and a system using the same. The system is characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, and an image coding / decoding device that includes both. Other components of the system can be appropriately changed depending on the situation.

[0457] [Usage example] 34 is a diagram showing the overall configuration of a content supply system ex100 that provides a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations, are installed in each cell.

[0458] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be configured to connect a combination of any of the above elements. The devices may be connected to each other directly or indirectly via a telephone network or short-range wireless communication, without using the base stations ex106 to ex110, which are fixed wireless stations. Furthermore, a streaming server ex103 is connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 is connected to a terminal in a hotspot on an airplane ex117, etc., via a satellite ex116.

[0459] Note that wireless access points, hotspots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0460] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that is compatible with mobile communication systems generally known as 2G, 3G, 3.9G, 4G, and 5G.

[0461] The home appliance ex118 is a refrigerator or an appliance included in a home fuel cell cogeneration system.

[0462] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) performs the encoding process described in each of the above embodiments on still images or video content captured by a user using the terminal, multiplexes the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0463] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal on an airplane ex117, or the like, which is capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device functions as an image decoding device according to one aspect of the present disclosure.

[0464] [Distributed processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented as a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server is dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing delays. Furthermore, if an error occurs or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the affected network portion, thereby achieving high-speed and stable distribution.

[0465] In addition to the distributed processing of the distribution itself, the encoding of captured data can be performed on each device, on the server side, or shared among devices. For example, encoding generally involves two processing loops. The first loop detects the image complexity or code size for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a device can perform the first encoding process, and the server that receives the content can perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each device. In this case, if there is a request for near-real-time reception and decoding, the data encoded by a device can be received and played back on another device, enabling more flexible real-time distribution.

[0466] As another example, the camera ex113 or the like extracts features from an image, compresses the data related to the features as metadata, and transmits the data to the server. The server performs compression according to the meaning of the image, for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server recompresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a heavy processing load such as CABAC (context-adaptive binary arithmetic coding).

[0467] As another example, in a stadium, shopping mall, factory, etc., there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.

[0468] Furthermore, since multiple pieces of video data are of nearly the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referenced. Alternatively, the server may receive encoded data from each terminal and change the reference relationships between multiple pieces of data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.

[0469] The server may also perform transcoding to change the encoding format of the video data before distributing it. For example, the server may convert MPEG-based encoding to VP-based encoding, or convert H.264 to H.265.

[0470] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0471] [3D, multi-angle] In recent years, there has been an increasing trend to integrate and use images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are nearly synchronized with each other. The videos taken by each device are integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points included in the videos match.

[0472] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a 3D shape of the scene based on not only the 2D video but also images of the same scene captured from different angles. The server may also separately encode 3D data generated by point clouds, or may select or reconstruct images to be transmitted to the receiving terminal from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0473] In this way, users can enjoy scenes by selecting any video corresponding to each camera device, or can enjoy content in which video from any viewpoint is extracted from 3D data reconstructed using multiple images or videos. Furthermore, like the video, sound may also be collected from multiple different angles, and the server may multiplex and transmit sound from a specific angle or space in accordance with the video.

[0474] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, and may perform encoding that allows reference between the viewpoint images using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0475] In the case of AR images, the server superimposes virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device may acquire or store virtual object information and 3D data, generate a 2D image according to the user's viewpoint movement, and smoothly connect the images to create superimposed data. Alternatively, the decoding device may send the user's viewpoint movement to the server in addition to a request for virtual object information, and the server may create superimposed data based on the viewpoint movement received from the 3D data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data may also have an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the 3D data to 0, etc., to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object.

[0476] Similarly, the decoding of distributed data may be performed by each client terminal, by the server, or by multiple terminals. For example, one terminal may first send a reception request to the server, and then other terminals may receive and decode content according to the request, after which the decoded signal is transmitted to a device with a display. By distributing the processing and selecting appropriate content regardless of the capabilities of the communication terminals themselves, high-quality data can be reproduced. As another example, large-sized image data may be received on a TV or other device, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of ​​responsibility or an area of ​​interest in more detail.

[0477] In the future, it is expected that content will be seamlessly received by switching the appropriate data for the current connection using delivery system standards such as MPEG-DASH in situations where multiple short-, medium-, or long-distance wireless communications are available, both indoors and outdoors. This will allow users to freely select and switch between decoding and display devices, such as their own devices, indoors and outdoors, in real time. Decoding can also be performed by switching between decoding and display devices based on user location information. This will enable users to display map information on the wall or ground of a neighboring building with an embedded display device while traveling to their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as if the encoded data is cached on a server that can be quickly accessed from the receiving device or copied to an edge server in a content delivery service.

[0478] [Scalable Coding] Regarding content switching, we will explain it using a scalable stream, shown in Figure 35, compressed and encoded using the video encoding method described in each of the above embodiments. The server may have multiple streams with the same content but different qualities, but may also switch content by taking advantage of the characteristics of a temporally / spatially scalable stream, achieved by encoding the stream in layers as shown. In other words, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as communication bandwidth, allowing the decoding side to freely switch between low-resolution and high-resolution content. For example, if a user wants to continue watching a video they were watching on their smartphone ex115 while on the go on a device such as an Internet TV after returning home, the device can simply decode the same stream up to different layers, thereby reducing the burden on the server.

[0479] Furthermore, in addition to the above-described scalability configuration in which pictures are coded for each layer and an enhancement layer exists above a base layer, the enhancement layer may include meta-information based on image statistics, etc., and the decoding side may generate high-quality content by super-resolving pictures in the base layer based on the meta-information. Super-resolution may mean either improving the signal-to-noise ratio at the same resolution or increasing the resolution. The meta-information may include information for specifying linear or nonlinear filter coefficients used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least-squares calculation used in the super-resolution process.

[0480] Alternatively, a picture may be divided into tiles or the like according to the meaning of objects in the image, and the decoding side may select tiles to decode and decode only a portion of the area. Furthermore, by storing the object's attributes (such as a person, a car, or a ball) and its position in the video (such as a coordinate position in the same image) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in Figure 36, the meta information is stored using a data storage structure different from that of pixel data, such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0481] Furthermore, meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units, which allows the decoding side to obtain the time when a specific person appears in the video, and by combining this with information in units of pictures, it is possible to identify the picture in which the object exists and the position of the object within the picture.

[0482] [Webpage optimization] FIG. 37 is a diagram showing an example of a web page display screen on a computer ex111 or the like. FIG. 38 is a diagram showing an example of a web page display screen on a smartphone ex115 or the like. As shown in FIGS. 37 and 38, a web page may include multiple link images that are links to image content, and the appearance of the web page may differ depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture contained in each content as a link image, display a video such as a GIF animation using multiple still images or I-pictures, or receive only the base layer and decode and display the video until the user explicitly selects a link image, or until the link image approaches the center of the screen or until the entire link image is within the screen.

[0483] When a link image is selected by a user, the display device decodes the base layer with the highest priority. If the HTML constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. To ensure real-time performance, before a selection is made or when the communication bandwidth is very limited, the display device decodes and displays only forward-referenced pictures (I-pictures, P-pictures, and forward-reference-only B-pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Alternatively, the display device may intentionally ignore the picture reference relationships and roughly decode all B-pictures and P-pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.

[0484] [Autonomous driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0485] In this case, since a vehicle, drone, airplane, etc. including a receiving terminal moves, the receiving terminal can realize seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting the location information of the receiving terminal at the time of a reception request. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, user situation, or communication bandwidth status.

[0486] In this way, in the content supply system ex100, the client can receive, decode, and play back the encoded information sent by the user in real time.

[0487] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distribution companies, but also unicast or multicast distribution of low-quality, short-duration content from individuals. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, with the following configuration.

[0488] During shooting, either in real time or after accumulating the footage, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes (e.g., scenes with lower brightness or out-of-focus compared to other pictures), emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scene.

[0489] In some cases, personal content may contain content that infringes copyright, moral rights, or portrait rights, or may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. The server may also recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons, and the server may replace the specified area with another image or blur the focus. For a person, the server may track the person in the video and replace the image of the face.

[0490] Furthermore, because viewing personal content with small data volumes requires real-time performance, the decoding device first receives the base layer as a top priority, and then decodes and plays it back, depending on the bandwidth. The decoding device may also receive an enhancement layer during this time, and if the content is played back more than twice, such as when playback is looped, it may play back high-quality video, including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smoother and the image quality improves. In addition to scalable encoding, a similar experience can also be provided by configuring a single stream consisting of a rough stream played the first time and a second stream that is encoded with reference to the first video.

[0491] [Other use cases] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 possessed by each terminal. The LSIex500 may be a single chip or may be configured with multiple chips. It is also possible to incorporate video encoding or decoding software into some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by the computer ex111, and perform the encoding or decoding process using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data is data that has been encoded by the LSIex500 possessed by the smartphone ex115.

[0492] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays the content.

[0493] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, the content supply system ex100 is more suited to multicast than the content supply system ex100, which is more suited to unicast, but similar applications are possible with regard to encoding and decoding processes.

[0494] [Hardware configuration] FIG. 39 is a diagram illustrating a smartphone ex115. FIG. 40 is a diagram illustrating an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and other images received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing encoded data or decoded data such as captured video or still images, recorded voice, received video or still images, and email, and a slot unit ex464 that serves as an interface with a SIM ex468 for identifying users and authenticating access to various data, including networks. In addition, an external memory may be used instead of the memory unit ex467.

[0495] In addition, a main control unit ex460 that comprehensively controls the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a bus ex470.

[0496] When the power key is turned on by a user, the power supply circuit unit ex461 supplies power from the battery pack to each unit, thereby starting up the smartphone ex115 into an operational state.

[0497] The smartphone ex115 processes calls, data communications, and other communications under the control of a main control unit ex460, which includes a CPU, ROM, RAM, and the like. During calls, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454, which then undergoes spectrum spread processing by the modulation / demodulation unit ex452, digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, and then transmitted via the antenna ex450. The received data is amplified, frequency-converted, and analog-to-digital converted, then subjected to spectrum despreading processing by the modulation / demodulation unit ex452, and converted into an analog audio signal by the audio signal processing unit ex454, which then outputs the amplified data from the audio output unit ex457. During data communications mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466, etc., of the main unit, and similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compression-encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the video encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 also encodes the audio signal picked up by the audio input unit ex456 while the camera unit ex465 is capturing video, still images, etc., and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451 before transmitting the data via the antenna ex450.

[0498] When receiving video attached to an email or chat, or video linked to a web page, etc., the multiplexed data received via the antenna ex450 is decoded by the multiplexing / separation unit ex453, which separates the multiplexed data into a video data bitstream and an audio data bitstream. The multiplexing / separation unit ex453 then supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and displays the video or still image included in the linked video file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. Note that with the widespread use of real-time streaming, audio playback may be socially inappropriate depending on the user's circumstances. Therefore, a configuration that initially plays only the video data without playing the audio signal is desirable. The audio may be played in synchronization only when the user performs an operation such as clicking on the video data.

[0499] Although the smartphone ex115 has been used as an example, three types of implementation are possible for the terminal: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcasting system, multiplexed data in which audio data and the like are multiplexed onto video data is received or transmitted, but the multiplexed data may also include text data related to the video in addition to audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0500] While the main control unit ex460, which includes a CPU, controls the encoding and decoding processes, devices often also include a GPU. Therefore, a configuration is possible in which a memory shared by the CPU and GPU, or a memory with addresses managed for common use, is used to take advantage of the GPU's performance and process a large area at once. This shortens encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation and quantization processes at a picture level or other unit in the GPU rather than the CPU. [Industrial Applicability]

[0501] The present disclosure is applicable to, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, video conference systems, electronic mirrors, and the like. [Explanation of symbols]

[0502] 100 Encoding device 102 Division 104 Subtraction section 106 Conversion unit 108 Quantization section 110 Entropy coding unit 112, 204 Inverse quantization section 114, 206 Inverse conversion unit 116, 208 Addition section 118, 210 block memory 120, 212 Loop filter section 122, 214 frame memory 124, 216 Intra prediction section 126, 218 Inter prediction section 128, 220 Predictive control unit 160, 260 circuits 162, 262 memory 200 Decryption Device 202 Entropy Decoding Unit

Claims

1. An encoding device that encodes a moving image by performing quantization, The circuit and a memory; The circuit uses the memory to: performing an up-conversion process only in a first direction on a first quantization matrix for square blocks and performing a down-conversion process only in the first direction to generate a second quantization matrix for rectangular blocks; quantizing a plurality of transform coefficients of a target block using the second quantization matrix; encoding the quantized transform coefficients of the current block; the first quantization matrix is ​​a square matrix having a first number of rows and columns; the second quantization matrix is ​​a rectangular matrix having a second number of rows and a third number of columns different from the second number, and the second number or the third number is the same as the first number; In the up-conversion process, the circuit copies matrix elements of the first quantization matrix in the first direction; In the down-conversion process, the circuit thins out matrix elements of the first quantization matrix in the first direction; encoding only information relating to the first quantization matrix of the first and second quantization matrices into a bitstream; The first direction is a vertical direction. Encoding device.

2. A decoding device that decodes a moving image by performing inverse quantization, The circuit and a memory; The circuit uses the memory to: performing an up-conversion process only in a first direction on a first quantization matrix for square blocks and performing a down-conversion process only in the first direction to generate a second quantization matrix for rectangular blocks; performing inverse quantization on a plurality of quantization coefficients of a target block using the second quantization matrix; Decoding the plurality of quantized coefficients of the inversely quantized block; the first quantization matrix is ​​a square matrix having a first number of rows and columns; the second quantization matrix is ​​a rectangular matrix having a second number of rows and a third number of columns different from the second number, and the second number or the third number is the same as the first number; In the up-conversion process, the circuit copies matrix elements of the first quantization matrix in the first direction; In the down-conversion process, the circuit thins out matrix elements of the first quantization matrix in the first direction; Decoding only information relating to the first quantization matrix from the bitstream, of the first quantization matrix and the second quantization matrix; The first direction is a vertical direction. Decryption device.

3. The circuit and a memory; The circuit uses the memory to: performing an up-conversion process only in a first direction on a first quantization matrix for square blocks and performing a down-conversion process only in the first direction to generate a second quantization matrix for rectangular blocks; quantizing a plurality of transform coefficients of a target block using the second quantization matrix; encoding the quantized transform coefficients of the current block; the first quantization matrix is ​​a square matrix having a first number of rows and columns; the second quantization matrix is ​​a rectangular matrix having a second number of rows and a third number of columns different from the second number, and the second number or the third number is the same as the first number; In the up-conversion process, the circuit copies matrix elements of the first quantization matrix in the first direction; In the down-conversion process, the circuit thins out matrix elements of the first quantization matrix in the first direction; generating a bitstream including only information relating to the first quantization matrix out of the first quantization matrix and the second quantization matrix; The first direction is a vertical direction. Bitstream generator.