Encoding apparatus, decoding apparatus, and storage medium

By using the quantization matrix generated by up-conversion and down-conversion in moving images, the problem of low quantization efficiency of rectangular blocks is solved, and more efficient encoding processing is achieved.

CN116233459BActive Publication Date: 2025-11-18PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310279343.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-03-30
Filing Date
2019-03-27
Publication Date
2025-11-18
Estimated Expiration
2039-03-27

AI Technical Summary

Technical Problem

In existing technologies, the quantization efficiency of rectangular blocks of various shapes in motion images is low, which leads to a reduction in coding efficiency.

Method used

The circuit and memory are used to perform up-conversion and down-conversion to generate quantization matrices with different numbers of rows and columns, which are used to quantize rectangular blocks of various shapes in motion images. The quantization efficiency is improved by adjusting the number of rows and columns of the quantization matrix in the horizontal and vertical directions.

Benefits of technology

It improves the quantization efficiency of rectangular blocks of various shapes in moving images, thereby enhancing coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116233459B_ABST
    Figure CN116233459B_ABST
Patent Text Reader

Abstract

The present application provides an encoding device, a decoding device and a storage medium. The encoding device performs up-conversion and down-conversion on a first quantization matrix to generate a second quantization matrix, the first quantization matrix has a first number of rows and a first number of columns equal to the first number of rows to form a square matrix, the second quantization matrix has a second number of rows and a second number of columns different from the second number of rows to form a rectangular matrix; and, using the second quantization matrix to quantize the transform coefficients of the current block, wherein, in the up-conversion, the circuit generates the second quantization matrix by performing up-conversion in a first direction in such a way that one of the second number of rows and the second number of columns is greater than the first number of rows, the first direction is the horizontal direction; and, in the down-conversion, the circuit generates the second quantization matrix by performing down-conversion in a second direction in such a way that the other of the second number of rows and the second number of columns is less than the first number of rows, the second direction is different from the first direction, and the second direction is the vertical direction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on March 27, 2019, with application number 201980020977.0 and entitled "Encoding device, decoding device, encoding method and decoding method". Technical Field

[0002] This invention relates to encoding apparatus for encoding moving images, etc. Background Technology

[0003] Previously, as a specification for encoding moving images, there was H.265, also known as HEVC (High Efficiency Video Coding) (Non-Patent Document 1).

[0004] Existing technical documents

[0005] Non-patent literature

[0006] Non-patent literature 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0007] The problem that the invention aims to solve

[0008] However, coding efficiency is reduced if the various shapes of rectangular blocks contained in the motion picture are not efficiently quantized.

[0009] Therefore, the present invention provides an encoding device capable of efficiently quantizing rectangular blocks of various shapes contained in a moving image.

[0010] Methods for solving problems

[0011] An encoding apparatus according to a technical solution of the present invention includes: a circuit and a memory coupled to the circuit. The circuit uses the memory to perform the following processing: performing upconversion and downconversion on a first quantization matrix to generate a second quantization matrix, wherein the first quantization matrix has a number of rows equal to the number of rows and a number of columns equal to the number of columns to form a square matrix, and the second quantization matrix has a number of rows equal to the number of rows and a number of columns different from the number of columns to form a rectangular matrix; and using the second quantization matrix to quantize the transform coefficients of the current block, wherein, in the upconversion, the circuit generates the second quantization matrix by performing the upconversion in a first direction such that one of the number of rows and the number of columns is greater than the number of rows, and the first direction is a horizontal direction; and, in the downconversion, the circuit generates the second quantization matrix by performing the downconversion in a second direction such that the other of the number of rows and the number of columns is less than the number of rows, and the second direction is different from the first direction, and the second direction is a vertical direction.

[0012] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0013] Invention Effects

[0014] The encoding device and the like of one of the technical solutions of the present invention can efficiently quantize rectangular blocks of various shapes contained in a moving image. Attached Figure Description

[0015] Figure 1 This is a block diagram showing the functional structure of the encoding device according to Embodiment 1.

[0016] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1.

[0017] Figure 3 It is a table representing the transformation basis functions corresponding to each transformation type.

[0018] Figure 4A This is a diagram showing an example of the shape of the filter used in ALF.

[0019] Figure 4B This is another example of the shape of the filter used in ALF.

[0020] Figure 4C This is another example of the shape of the filter used in ALF.

[0021] Figure 5A This is a diagram representing the 67 intra-prediction modes of intra-frame prediction.

[0022] Figure 5B This is a flowchart illustrating the outline of predictive image correction processing based on OBMC processing.

[0023] Figure 5C This is a conceptual diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.

[0024] Figure 5D This is a diagram representing an example of FRUC.

[0025] Figure 6 It is a diagram used to illustrate pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0026] Figure 7 It is a diagram used to illustrate pattern matching (template matching) between a template in the current image and a block in a reference image.

[0027] Figure 8 It is a diagram used to illustrate a model that assumes uniform linear motion.

[0028] Figure 9A It is a diagram used to illustrate the derivation of the motion vectors of sub-block units based on the motion vectors of multiple adjacent blocks.

[0029] Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.

[0030] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.

[0031] Figure 9D This is a diagram used to illustrate an overview of a predictive image generation method that employs LIC-based brightness correction processing.

[0032] Figure 10 This is a block diagram illustrating the functional structure of the decoding device according to Embodiment 1.

[0033] Figure 11 This is a diagram illustrating the first example of the encoding process using a quantization matrix in an encoding device.

[0034] Figure 12 It indicates that it was used in Figure 11 The diagram illustrates an example of the decoding process of the quantization matrix in the decoding device corresponding to the encoding device described in the text.

[0035] Figure 13 It is used to explain in Figure 11Step S102 and Figure 12 In step S202, a diagram of the first example of generating a quantization matrix for a rectangular block from a quantization matrix used for a square block is shown.

[0036] Figure 14 It is used to illustrate that by in Figure 13 The diagram illustrates the method of generating a rectangular block by downconverting the quantization matrix used for the corresponding square block.

[0037] Figure 15 It is used to explain in Figure 11 Step S102 and Figure 12 In step S202, a diagram of the second example of generating a quantization matrix for a rectangular block from a quantization matrix used for a square block is shown.

[0038] Figure 16 It is used to illustrate that by in Figure 15 The diagram illustrates the method of generating a rectangular block by upconverting the quantization matrix used for the corresponding square block.

[0039] Figure 17 This is the second example of a diagram illustrating the encoding process using a quantization matrix in an encoding device.

[0040] Figure 18 It indicates that it was used in Figure 17 The diagram illustrates an example of the decoding process of the quantization matrix in the decoding device corresponding to the encoding device described in the text.

[0041] Figure 19 It is used to explain in Figure 17 Step S701 and Figure 18 A diagram showing an example of the quantization matrix corresponding to the size of the effective transformation coefficient region in each block size in step S801.

[0042] Figure 20 This is a diagram illustrating a variation of the second example of the encoding process using a quantization matrix in an encoding device.

[0043] Figure 21 It indicates that it was used in Figure 20 The diagram illustrates an example of the decoding process of the quantization matrix in the decoding device corresponding to the encoding device described in the text.

[0044] Figure 22 It is used to explain in Figure 20 Step S1002 and Figure 21 In step S1102, a diagram of the first example of generating a quantization matrix for a rectangular block from a quantization matrix for a square block is shown.

[0045] Figure 23 It is used to illustrate that by in Figure 22 The diagram illustrates the method of generating a rectangular block by down-converting the quantization matrix used for the corresponding square block.

[0046] Figure 24 It is used to explain in Figure 20 Step S1002 and Figure 21 In step S1102, a diagram of the second example of generating a quantization matrix for a rectangular block from a quantization matrix used for a square block is shown.

[0047] Figure 25 It is used to illustrate that by in Figure 24 The diagram illustrates the method of generating the quantization matrix for the rectangular block by upconverting the quantization matrix for the corresponding square block.

[0048] Figure 26 This is the third example of a diagram illustrating the encoding process using a quantization matrix in an encoding device.

[0049] Figure 27 It indicates that it was used in Figure 26 The diagram illustrates an example of the decoding process of the quantization matrix in the decoding device corresponding to the encoding device described in the text.

[0050] Figure 28 It is used to explain in Figure 26 Step S1601 and Figure 27 In step S1701, a diagram illustrates an example of a method for generating a quantization matrix using a common approach based on the values ​​of the quantization coefficients of the quantization matrix for only the diagonal components in each block size.

[0051] Figure 29 It is used to explain in Figure 26 Step S1601 and Figure 27 In step S1701, another example of a method for generating a quantization matrix using a common method is shown in the figure, based on the values ​​of the quantization coefficients of the quantization matrix that uses only the diagonal components in the processing object blocks of each block size.

[0052] Figure 30 This is a block diagram showing an example of the installation of an encoding device.

[0053] Figure 31 It means Figure 30 The flowchart shows an example of the operation of the encoding device.

[0054] Figure 32 This is a block diagram illustrating an example of the installation of a decoding device.

[0055] Figure 33 It means Figure 32The flowchart shows an example of the operation of the decoding device.

[0056] Figure 34 This is a diagram showing the overall structure of a content supply system that enables content distribution services.

[0057] Figure 35 This is a diagram illustrating an example of encoding construction in the case of hierarchical encoding.

[0058] Figure 36 This is a diagram illustrating an example of encoding construction in the case of hierarchical encoding.

[0059] Figure 37 This is an example of a web page display.

[0060] Figure 38 This is an example of a web page display.

[0061] Figure 39 This is a diagram illustrating an example of a smartphone.

[0062] Figure 40 This is a block diagram representing a structural example of a smartphone. Detailed Implementation

[0063] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.

[0064] Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangements and connection methods of constituent elements, steps, and sequences of steps shown in the following embodiments are examples and are not intended to limit the scope of the claims. In addition, any constituent elements in the following embodiments that are not described in the independent claim representing the highest-level concept are described as arbitrary constituent elements.

[0065] (Implementation Method 1)

[0066] First, as an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention described later, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding and decoding apparatus for the processing and / or structure described in the various embodiments of the present invention, and the processing and / or structure described in the various embodiments of the present invention can also be implemented in encoding and decoding apparatuses different from Embodiment 1.

[0067] When applying the processing and / or structure described in various aspects of the present invention to Embodiment 1, one of the following may also be performed, for example.

[0068] (1) For the encoding or decoding device of Embodiment 1, the constituent element that corresponds to the constituent element described in each aspect of the present invention is replaced with the constituent element described in each aspect of the present invention.

[0069] (2) For the encoding or decoding device of Embodiment 1, after any modification such as adding, replacing, or deleting any of the constituent elements of the plurality of constituent elements constituting the encoding or decoding device, the constituent elements corresponding to the constituent elements described in each aspect of the present invention are replaced with the constituent elements described in each aspect of the present invention.

[0070] (3) After adding processing to the method implemented by the encoding or decoding device of Embodiment 1, and / or replacing or deleting any of the processing among the multiple processing included in the method, the processing corresponding to the processing described in each aspect of the present invention is replaced with the processing described in each aspect of the present invention.

[0071] (4) A portion of the constituent elements constituting the encoding or decoding apparatus of Embodiment 1 are combined with constituent elements described in various aspects of the present invention, a portion of constituent elements having the functions of constituent elements described in various aspects of the present invention, or a portion of constituent elements implementing the processing performed by constituent elements described in various aspects of the present invention.

[0072] (5) A component having a portion of the functions of a portion of the components constituting the encoding or decoding apparatus of embodiment 1, or a component implementing a portion of the processing performed by a portion of the components constituting the encoding or decoding apparatus of embodiment 1, is combined with the components described in various aspects of the present invention, the components having a portion of the functions of the components described in various aspects of the present invention, or the components implementing a portion of the processing performed by the components described in various aspects of the present invention.

[0073] (6) For the method implemented by the encoding or decoding device of Embodiment 1, the processing that corresponds to the processing described in each aspect of the present invention among the multiple processing included in the method is replaced with the processing described in each aspect of the present invention.

[0074] (7) A portion of the processing included in the method implemented by the encoding or decoding apparatus of Embodiment 1 is combined with the processing described in the various aspects of the present invention.

[0075] Furthermore, the implementation of the processes and / or structures described in the various embodiments of the present invention is not limited to the examples described above. For example, it may be implemented in an apparatus used for a different purpose than the moving image / image encoding apparatus or moving image / image decoding apparatus disclosed in Embodiment 1, or the processes and / or structures described in each embodiment may be implemented individually. In addition, the processes and / or structures described in different embodiments may be combined and implemented.

[0076] [Overview of the encoding device]

[0077] First, an overview of the encoding device for Embodiment 1 will be provided. Figure 1 This is a block diagram illustrating the functional structure of the encoding apparatus 100 according to Embodiment 1. The encoding apparatus 100 is a motion picture / image encoding apparatus that encodes motion pictures / images in block units.

[0078] like Figure 1 As shown, the encoding device 100 is a device for encoding images in block units, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a cyclic filtering unit 120, a frame memory 122, an intra-frame prediction unit 124, an inter-frame prediction unit 126, and a prediction control unit 128.

[0079] The encoding device 100 is implemented, for example, by a general-purpose processor and memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the segmentation unit 102, subtraction unit 104, transform unit 106, quantization unit 108, entropy coding unit 110, inverse quantization unit 112, inverse transform unit 114, addition unit 116, cyclic filtering unit 120, intra-frame prediction unit 124, inter-frame prediction unit 126, and prediction control unit 128.

[0080] The following describes the constituent elements included in the encoding device 100.

[0081] [Divider]

[0082] The segmentation unit 102 divides each image contained in the input moving image into multiple blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first segments the image into fixed-size blocks (e.g., 128×128). These fixed-size blocks may be called coding tree units (CTUs). Furthermore, the segmentation unit 102 divides each fixed-size block into variable-size blocks (e.g., 64×64 or less) based on recursive quadtree and / or binary tree block segmentation. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the image may be used as processing units of CUs, PUs, and TUs.

[0083] Figure 2 This is a diagram illustrating an example of block segmentation in Implementation Method 1. In Figure 2 In the diagram, solid lines represent block boundaries based on quadtree block partitioning, and dashed lines represent block boundaries based on binary tree block partitioning.

[0084] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into 4 square blocks of 64×64 (quadtree block partitioning).

[0085] The 64×64 block in the upper left corner is then vertically divided into two rectangular 32×64 blocks, and the 32×64 block on the left is then vertically divided into two rectangular 16×64 blocks (binary tree block partitioning). As a result, the 64×64 block in the upper left corner is divided into two 16×64 blocks (11 and 12) and a 32×64 block (13).

[0086] The 64×64 block in the upper right corner is horizontally divided into two rectangular 64×32 blocks, 14 and 15 (binary tree block division).

[0087] The 64×64 block in the lower left corner is divided into four 32×32 square blocks (quadtree block partitioning). The upper left and lower right blocks of these four 32×32 blocks are further partitioned. The upper left 32×32 block is vertically divided into two 16×32 rectangular blocks, and the right 16×32 block is horizontally divided into two 16×16 blocks (binary tree block partitioning). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block partitioning). As a result, the lower left 64×64 block is divided into 16×32 block 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.

[0088] The 64×64 block 23 in the lower right corner is not divided.

[0089] As described above, in Figure 2 In the example, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quadtree plus binary tree) partitioning.

[0090] In addition, Figure 2 In this context, a block can be divided into 2 or 4 blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these. For example, a block can also be divided into 3 blocks (ternary tree partitioning). Partitioning including such ternary tree partitioning is sometimes referred to as MBT (multi-type tree) partitioning.

[0091] [Subtraction Section]

[0092] The subtraction unit 104 subtracts the prediction signal (prediction sample) from the original signal (original sample) in block units divided by the segmentation unit 102. That is, the subtraction unit 104 calculates the prediction error (also called residual) of the encoded target block (hereinafter referred to as the current block). Furthermore, the subtraction unit 104 outputs the calculated prediction error to the transformation unit 106.

[0093] The original signal is the input signal of the encoding device 100, which is the signal representing the image of each picture that constitutes the moving image (e.g., luminance signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0094] [Transformation Section]

[0095] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain, and outputs the transformation coefficients vectorization unit 108. Specifically, the transformation unit 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.

[0096] Alternatively, the transform unit 106 can adaptively select a transform type from multiple transform types and use the transform basis function corresponding to the selected transform type to transform the prediction error into transform coefficients. Such a transform is sometimes referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform).

[0097] Several transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 This is a table representing the transformation basis functions corresponding to each transformation type. Figure 3 In this context, N represents the number of input pixels. The choice of transform type from these multiple transform types can depend on the type of prediction (intra-frame prediction and inter-frame prediction) or the intra-frame prediction mode.

[0098] Information indicating whether such EMT or AMT is applied (e.g., referred to as the AMT flag) and information indicating the selected transform type are signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).

[0099] Furthermore, the transform unit 106 can also perform a re-transformation on the transform coefficients (transformation results). Such a re-transformation may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (e.g., a 4×4 sub-block) contained in the block of transform coefficients corresponding to the intra-frame prediction error. Information indicating whether NSST is applied and information related to the transform matrix used in NSST are signaled at the CU level. In addition, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0100] Here, a separable transformation refers to a method of performing multiple transformations in each direction, which is equivalent to the number of dimensions of the input. A non-separable transformation refers to a method of treating two or more dimensions together as one dimension and transforming them together when the input is multidimensional.

[0101] For example, as one example of a non-separable transformation, one can cite the way that when the input is a 4×4 block, it is treated as a permutation of 16 elements, and the permutation is transformed using a 16×16 transformation matrix.

[0102] Furthermore, the Hypercube Givens Transform, which treats a 4×4 input block as a permutation of 16 elements and then performs multiple Givens rotations on that permutation, is also an example of a non-separable transformation.

[0103] [Quantitative Department]

[0104] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scan order and quantizes the transform coefficients based on the quantization parameters (QP) corresponding to the scanned transform coefficients. Furthermore, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0105] The specified order is the order in which the transform coefficients are quantized / inverse quantized. For example, the specified scan order is defined by ascending frequency (from low frequency to high frequency) or descending frequency (from high frequency to low frequency).

[0106] The quantization parameter is a parameter that defines the quantization step size (quantization width). For example, if the value of the quantization parameter increases, the quantization step size also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0107] [Entropy Coding Department]

[0108] The entropy coding unit 110 generates a coded signal (coded bitstream) by performing variable-length coding on the quantization coefficients, which are input from the quantization unit 108. Specifically, the entropy coding unit 110 performs arithmetic coding on the binary signal, for example, by binarizing the quantization coefficients.

[0109] [De-quantization Department]

[0110] The inverse quantization unit 112 performs inverse quantization on the quantization coefficients that are input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantization coefficients of the current block in a predetermined scan order. Furthermore, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0111] [Inverse Transformation Section]

[0112] The inverse transform unit 114 restores the prediction error by performing an inverse transform on the transform coefficients, which are input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Furthermore, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0113] Furthermore, the restored prediction error differs from the prediction error calculated by the subtraction unit 104 because information was lost during quantization. In other words, the restored prediction error includes quantization error.

[0114] [Addition Department]

[0115] The addition unit 116 reconstructs the current block by adding the prediction error, which is input from the inverse transform unit 114, to the prediction sample, which is input from the prediction control unit 128. Furthermore, the addition unit 116 outputs the reconstructed block to the block memory 118 and the cyclic filtering unit 120. The reconstructed block may be referred to as a local decoding block.

[0116] [Block Memory]

[0117] Block memory 118 is a storage unit used to save blocks within the encoded object image (hereinafter referred to as the current image) referenced in intra-frame prediction. Specifically, block memory 118 saves the reconstructed blocks output from addition unit 116.

[0118] [Loop Filtering Section]

[0119] The cyclic filtering unit 120 applies cyclic filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Cyclic filtering refers to filtering used within the encoding loop (in-loop filtering), such as deblocking filtering (DF), sample adaptive offset (SAO), and adaptive cyclic filtering (ALF).

[0120] In ALF, a least-squares error filter is used to remove coding distortion. For example, for each 2×2 sub-block within the current block, one filter is selected from multiple filters based on the direction of the gradient and the activity of the locality.

[0121] Specifically, sub-blocks (e.g., 2×2 sub-blocks) are first classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is based on the direction and activity of the gradient. For example, using the gradient direction value D (e.g., 0–2 or 0–4) and the gradient activity value A (e.g., 0–4), a classification value C (e.g., C = 5D + A) is calculated. Then, based on the classification value C, the sub-blocks are classified into multiple classes (e.g., 15 or 25 classes).

[0122] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, the gradient activity value A is derived, for example, by summing the gradients in multiple directions and quantizing the sum.

[0123] Based on the results of this classification, the filter used for the sub-block is determined from among multiple filters.

[0124] The shape of the filter used in ALF can be, for example, a circular symmetrical shape. Figures 4A to 4C This is a diagram showing several examples of the shapes of filters used in ALF. Figure 4A This indicates a 5×5 diamond-shaped filter. Figure 4BThis indicates a 7×7 diamond-shaped filter. Figure 4C This represents a 9×9 diamond-shaped filter. Information representing the filter's shape is signaled at the image level. However, the signaling of the filter's shape information is not limited to the image level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0125] The on / off state of ALF is determined, for example, at the picture level or the CU level. For instance, regarding luminance, the decision to use ALF is made at the CU level, while regarding chromatic aberration, it is made at the picture level. Information indicating the on / off state of ALF is signaled at the picture level or the CU level. However, the signaling of information indicating the on / off state of ALF is not limited to the picture level or the CU level; it can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0126] The coefficient set of a selectable set of filters (e.g., up to 15 or 25 filters) is signaled at the picture level. Furthermore, the signaling of the coefficient set is not limited to the picture level; it can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0127] [Frame Memory]

[0128] The frame memory 122 is a storage unit used to store reference images used in inter-frame prediction, and is also sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the cyclic filtering unit 120.

[0129] Intra-frame prediction unit

[0130] The intra-frame prediction unit 124 performs intra-frame prediction (also called intra-picture prediction) of the current block by referring to the blocks in the current image stored in the block memory 118, thereby generating a prediction signal (intra-frame prediction signal). Specifically, the intra-frame prediction unit 124 generates an intra-frame prediction signal by performing intra-frame prediction by referring to samples (e.g., luminance value, chrominance value) of blocks adjacent to the current block, and outputs the intra-frame prediction signal to the prediction control unit 128.

[0131] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of predefined intra-prediction modes. The plurality of intra-prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0132] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode as specified by the H.265 / HEVC (High-Efficiency Video Coding) specification (Non-Patent Document 1).

[0133] Multiple directional prediction modes may include, for example, the 33 directional prediction modes specified in the H.265 / HEVC specification. Alternatively, multiple directional prediction modes may also include 32 additional directional prediction modes (a total of 65 directional prediction modes). Figure 5A This diagram represents the 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-frame prediction. Solid arrows indicate the 33 directions specified by the H.265 / HEVC specification, while dashed arrows indicate the additional 32 directions.

[0134] Additionally, in intra-frame prediction of chroma blocks, luma blocks can also be referenced. That is, the chroma components of the current block can be predicted based on the luma components of the current block. Such intra-frame prediction is sometimes referred to as CCLM (cross-component linear model) prediction. This intra-frame prediction mode of chroma blocks referencing luma blocks (e.g., called CCLM mode) can also be added as one of the intra-frame prediction modes for chroma blocks.

[0135] The intra-prediction unit 124 can also correct the intra-predicted pixel values ​​based on the gradient of the reference pixels in the horizontal / vertical directions. Intra-prediction accompanied by such correction is sometimes referred to as PDPC (position-dependent intraprediction combination). Information indicating whether PDPC has been used (e.g., a PDPC flag) is signaled, for example, at the CU level. Furthermore, the signaling of this information is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, or CTU level).

[0136] [Inter-frame prediction department]

[0137] The inter-frame prediction unit 126 performs inter-frame prediction (also called inter-picture prediction) for the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture, thereby generating a prediction signal (inter-frame prediction signal). Inter-frame prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Furthermore, the inter-frame prediction unit 126 uses motion information (e.g., motion vectors) obtained through motion estimation to perform motion compensation, thereby generating the inter-frame prediction signal for the current block or sub-block. Finally, the inter-frame prediction unit 126 outputs the generated inter-frame prediction signal to the prediction control unit 128.

[0138] The motion information used in motion compensation is signaled. A motion vector predictor can also be used in the signaling of motion vectors. That is, the difference between the motion vector and the predicted motion vector can also be signaled.

[0139] Alternatively, the inter-frame prediction signal can be generated using not only the motion information of the current block obtained through motion estimation but also the motion information of neighboring blocks. Specifically, the prediction signal based on the motion information obtained through motion estimation can be weighted and added together with the prediction signal based on the motion information of neighboring blocks, thereby generating the inter-frame prediction signal in sub-block units within the current block. Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0140] In this OBMC mode, information indicating the size of the sub-block used for OBMC (e.g., OBMC block size) is signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is used (e.g., OBMC flag) is signaled at the CU level. However, the signaling level for this information is not limited to the sequence and CU levels; it can also be other levels (e.g., image level, slice level, tile level, CTU level, or sub-block level).

[0141] The OBMC model will be explained in more detail. Figure 5B and Figure 5C This is a flowchart and concept diagram used to illustrate the outline of predictive image correction processing based on OBMC processing.

[0142] First, the predicted image (Pred) obtained through normal motion compensation is obtained using the motion vectors (MV) assigned to the encoded object block.

[0143] Next, the predicted image (Pred_L) is obtained by using the motion vector (MV_L) of the encoded left adjacent block for the encoded object block. The first correction of the predicted image is performed by weighted superposition of the predicted image and Pred_L.

[0144] Similarly, the predicted image (Pred_U) is obtained by using the motion vector (MV_U) of the upper adjacent block of the encoded object block. The predicted image is then corrected a second time by weighting and superimposing the predicted image after the first correction and Pred_U, and this is used as the final predicted image.

[0145] In addition, this describes a two-stage correction method using the left and top adjacent blocks, but it can also be configured to perform more corrections using the right and bottom adjacent blocks than the two-stage method.

[0146] In addition, the area to be overlaid may not be the entire pixel area of ​​the block, but only a part of the area near the block boundary.

[0147] Furthermore, the process of correcting the predicted image based on a single reference image is explained here. However, the same principle applies when correcting the predicted image based on multiple reference images. After obtaining the corrected predicted image based on each reference image, the resulting predicted images are further superimposed to obtain the final predicted image.

[0148] In addition, the processing target block mentioned above can be a prediction block unit or a sub-block unit that further divides the prediction block.

[0149] One method for determining whether to use OBMC processing is to use a signal called obmc_flag. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a motion-complex region. If it does, the obmc_flag is set to 1 and OBMC processing is performed for encoding. If it does not belong to a motion-complex region, the obmc_flag is set to 0, and OBMC processing is not performed for encoding. On the other hand, in a decoding device, decoding is performed by decoding the obmc_flag recorded in the stream and switching between using and not using OBMC processing based on its value.

[0150] Alternatively, motion information can be exported at the decoding device side without being signaled. For example, the merging mode specified by the H.265 / HEVC standard can be used. Furthermore, motion information can also be exported by performing motion estimation at the decoding device side. In this case, motion estimation is performed without using the pixel values ​​of the current block.

[0151] Here, we will explain the motion estimation mode performed on the decoding device side. This motion estimation mode on the decoding device side may be called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0152] exist Figure 5DThe diagram below illustrates an example of FRUC processing. First, referencing the motion vectors of coded blocks spatially or temporally adjacent to the current block, a list of multiple candidates, each with a predicted motion vector, is generated (this list can also be shared with a merge list). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value is calculated for each candidate included in the candidate list, and one candidate is selected based on the evaluation value.

[0153] Furthermore, based on the selected candidate motion vectors, motion vectors for the current block are derived. Specifically, for example, the selected candidate motion vector (best candidate MV) can be derived as is, using it as the motion vector for the current block. Alternatively, for example, motion vectors for the current block can be derived by performing pattern matching in the surrounding region of the position within the reference image corresponding to the selected candidate motion vector. That is, the surrounding region of the best candidate MV can be searched using the same method, and if an MV with a better evaluation value is found, the best candidate MV is updated to the aforementioned MV and used as the final MV for the current block. Alternatively, a structure that does not perform this processing can be implemented.

[0154] The exact same processing can also be performed when processing is done in sub-block units.

[0155] Furthermore, the evaluation value is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region within the reference image corresponding to the motion vector and the specified region. Alternatively, information other than the difference value can be used to calculate the evaluation value.

[0156] As a pattern matching, either pattern matching 1 or pattern matching 2 is used. Pattern matching 1 and pattern matching 2 can be referred to as bilateral matching and template matching, respectively.

[0157] In the first pattern matching, pattern matching is performed between two blocks within two different reference images, along the motion trajectory of the current block. Therefore, in the first pattern matching, the regions within other reference images along the motion trajectory of the current block are used as the defined regions for calculating the candidate evaluation values ​​described above.

[0158] Figure 6 This diagram illustrates an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory. For example... Figure 6As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best matching pair among two blocks in two different reference images (Ref0, Ref1) along the motion trajectory of the current block. Specifically, for the current block, the difference between the reconstructed image at a specified position in the first encoded reference image (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference image (Ref1) specified by the symmetrical MV scaled by the aforementioned candidate MV over the display time interval is derived, and the obtained difference value is used to calculate an evaluation value. The candidate MV with the best evaluation value can be selected as the final MV from among multiple candidate MVs.

[0159] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) indicating two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, in the case where the current image is located between the two reference images in time and the temporal distances from the current image to the two reference images are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0160] In the second pattern matching, pattern matching is performed between the template in the current image (the block adjacent to the current block in the current image (e.g., the upper and / or left adjacent block)) and the block in the reference image. Therefore, in the second pattern matching, the block adjacent to the current block in the current image is used as the defined area for calculating the candidate evaluation value as described above.

[0161] Figure 7 This is an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 7 As shown, in the second pattern matching, the motion vector of the current block is derived by searching within the reference image (Ref0) for the block that best matches the block adjacent to the current block (Cur block) within the current image (Cur Pic). Specifically, for the current block, the difference between the reconstructed images of the encoded regions of the left and top adjacent regions or one of them and the reconstructed image at the same position within the encoded reference image (Ref0) specified by the candidate MV is derived. The obtained difference value is used to calculate the evaluation value, and the candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.

[0162] Information indicating whether FRUC mode is used (e.g., referred to as the FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is used (e.g., when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (e.g., referred to as the FRUC mode flag) is signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0163] This section explains the mode for deriving motion vectors based on a model that assumes uniform linear motion. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.

[0164] Figure 8 This diagram is used to illustrate a model that assumes uniform linear motion. Figure 8 In the middle, (v x v y () represents the velocity vector, and τ0 and τ1 represent the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1), respectively. (MVx0, MVy0) represents the motion vector corresponding to the reference image Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference image Ref1.

[0165] At this time, in the velocity vector (v x v y Under the assumption of constant linear motion of MVx0, MVy0 and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1) respectively, and the following optical flow equation (1) holds.

[0166] [Formula 1]

[0167]

[0168] Here, I (k) This represents the luminance value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation states that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on this optical flow equation combined with Hermite interpolation, the block-unit motion vector obtained from merge lists, etc., is corrected in pixels.

[0169] Alternatively, motion vectors can be derived on the decoding device side using a different method than deriving motion vectors based on a model assuming constant linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0170] Here, we will explain the mode of deriving motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0171] Figure 9A This is a diagram used to illustrate the derivation of sub-block unit motion vectors based on the motion vectors of multiple adjacent blocks. Figure 9A In this context, the current block comprises 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the top-left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the top-right control point of the current block is derived. Furthermore, using the two motion vectors v0 and v1, the motion vectors (v0, v1, v1) of each sub-block within the current block are derived using the following equation (2). x v y ).

[0172] [Formula 2]

[0173]

[0174] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents the pre-set weight coefficient.

[0175] Such an affine motion compensation prediction mode may also include several modes with different methods for deriving the motion vectors of the upper left and upper right control points. Information representing such an affine motion compensation prediction mode (e.g., affine flags) is signaled at the CU level. Furthermore, the signaling of information representing this affine motion compensation prediction mode is not limited to the CU level; it can also be at other levels (e.g., sequence level, image level, slice level, tile level, CTU level, or sub-block level).

[0176] [Forecasting and Control Department]

[0177] The prediction control unit 128 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0178] This section illustrates an example of exporting motion vectors from an encoded object image using a merge mode. Figure 9B This is a diagram used to illustrate the overview of motion vector derivation processing based on the merging mode.

[0179] First, a list of candidate predicted MVs registered with the predicted MVs is generated. Candidate predicted MVs include: spatially adjacent predicted MVs (MVs) belonging to multiple coded blocks spatially surrounding the coded object block; temporally adjacent predicted MVs (MVs) belonging to blocks whose positions in the coded reference image are projected nearby; combined predicted MVs (MVs) generated by combining the MV values ​​of spatially adjacent and temporally adjacent predicted MVs; and zero predicted MVs (MVs with a value of zero).

[0180] Next, the MV for the encoded object block is determined by selecting one predicted MV from the multiple predicted MVs registered in the predicted MV list.

[0181] Furthermore, in the variable-length coding section, merge_idx, which represents the signal that selected which prediction MV was recorded in the stream and encoded.

[0182] In addition, Figure 9B The predicted MVs registered in the predicted MV list described in the figure are one example. They may also be a number different from the number shown in the figure, or a structure that does not include a part of the predicted MVs in the figure, or a structure that adds predicted MVs other than the predicted MVs in the figure.

[0183] Alternatively, the MV of the encoded object block exported through the merge mode can be used for the DMVR processing described later to determine the final MV.

[0184] Here, an example of using DMVR to determine MV is explained.

[0185] Figure 9C This is a conceptual diagram used to illustrate the outline of DMVR processing.

[0186] First, the optimal MVP set for the processing object block is taken as the candidate MV. According to the candidate MV, reference pixels are obtained from the first reference image of the processed image in the L0 direction and the second reference image of the processed image in the L1 direction, respectively. The template is generated by taking the average of each reference pixel.

[0187] Next, using the template described above, the surrounding areas of the candidate music videos (MVs) for the first and second reference images are searched, and the MV with the lowest cost is selected as the final MV. Furthermore, the cost value is calculated using the differences between the pixel values ​​of the template and the pixel values ​​of the search area, as well as the MV value.

[0188] Furthermore, the general outline of the processing described herein is essentially the same in both the encoding and decoding devices.

[0189] In addition, even if it is not the process described here, any other process that can search for the surrounding of candidate MVs and export the final MV can be used.

[0190] Here, the mode of generating predicted images using LIC processing is explained.

[0191] Figure 9D This is a diagram illustrating the outline of a predictive image generation method using LIC-based brightness correction processing.

[0192] First, export the MV used to obtain the reference image corresponding to the encoded object block from the reference image, which is an encoded image.

[0193] Next, for the encoded object block, using the brightness pixel values ​​of the left and top adjacent encoded surrounding reference areas and the brightness pixel values ​​at the same position in the reference image specified by MV, information indicating how the brightness values ​​change in the reference image and the encoded object image is extracted, and brightness correction parameters are calculated.

[0194] By using the aforementioned brightness correction parameters to perform brightness correction processing on the reference image within the reference image specified by MV, a predicted image for the coded object block is generated.

[0195] in addition, Figure 9D The shape of the surrounding reference area mentioned above is one example; other shapes may also be used.

[0196] Furthermore, the process of generating a prediction image based on a single reference image is described here, but the same applies when generating a prediction image based on multiple reference images. The prediction image is generated after performing brightness correction processing on the reference images obtained from each reference image in the same way.

[0197] One method for determining whether to use LIC processing is to use a lic_flag as a signal indicating whether LIC processing is used. Specifically, in an encoding device, it is determined whether the block to be encoded belongs to a region where a brightness change has occurred. If it does, the lic_flag is set to 1, and LIC processing is used for encoding. If it does not belong to a region where a brightness change has occurred, the lic_flag is set to 0, and LIC processing is not used for encoding. On the other hand, in a decoding device, decoding is performed by decoding the lic_flag recorded in the stream and switching between using and not using LIC processing based on its value.

[0198] Other methods for determining whether to use LIC processing include checking whether LIC processing was used in surrounding blocks. As a specific example, when the encoded target block is in merge mode, it is determined whether the surrounding encoded blocks selected during the export of the MV in merge mode processing have been encoded using LIC processing. Based on the result, encoding is switched between using LIC processing and other methods. Furthermore, in this example, the decoding process is exactly the same.

[0199] [Overview of the Decoding Device]

[0200] Next, an outline of a decoding apparatus capable of decoding the encoded signal (encoded bit stream) output from the encoding apparatus 100 will be described. Figure 10 This is a block diagram illustrating the functional structure of the decoding device 200 according to Embodiment 1. The decoding device 200 is a motion image / image decoding device that decodes motion images / images in block units.

[0201] like Figure 10 As shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a cyclic filtering unit 212, a frame memory 214, an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220.

[0202] The decoding device 200 is implemented, for example, by a general-purpose processor and memory. In this case, when the processor executes the software program stored in the memory, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 can also be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the adder 208, the cyclic filter unit 212, the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220.

[0203] The following describes the constituent elements included in the decoding device 200.

[0204] [Entropy Decoding Department]

[0205] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. As a result, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units.

[0206] [De-quantization Department]

[0207] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoded target block (hereinafter referred to as the current block), which is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient of the current block based on the quantization parameter corresponding to that quantization coefficient. Furthermore, the inverse quantization unit 204 outputs the inverse quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0208] [Inverse Transformation Section]

[0209] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients, which are inputs from the inverse quantization unit 204.

[0210] For example, if the information read from the encoded bitstream represents EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information representing the transform type read from the transducer.

[0211] Furthermore, for example, when the information read from the encoded bitstream is represented using NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients.

[0212] [Addition Department]

[0213] The adder 208 reconstructs the current block by adding the prediction error, which is input from the inverse transform 206, to the prediction sample, which is input from the prediction control 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the cyclic filtering 212.

[0214] [Block Memory]

[0215] Block memory 210 is a storage unit used to store blocks within the decoded target image (hereinafter referred to as the current image) that serve as a reference in intra-frame prediction. Specifically, block memory 210 stores the reconstructed blocks output from adder 208.

[0216] [Loop Filtering Section]

[0217] The cyclic filtering unit 212 applies cyclic filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.

[0218] Given that the information indicating the on / off state of the ALF is read from the encoded bitstream, and the ALF is on, one filter is selected from multiple filters based on the direction and activity of the gradient of locality, and the selected filter is applied to the reconstructed block.

[0219] [Frame Memory]

[0220] The frame memory 214 is a storage unit used to store reference images used in inter-frame prediction; it is also sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the cyclic filtering unit 212.

[0221] Intra-frame prediction unit

[0222] The intra-prediction unit 216 performs intra-prediction based on the intra-prediction pattern read from the encoded bitstream, referring to blocks within the current image stored in the block memory 210, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra-prediction unit 216 performs intra-prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, thereby generating an intra-prediction signal, and outputs the intra-prediction signal to the prediction control unit 220.

[0223] In addition, if the intra-prediction mode of the reference luma block is selected in the intra-prediction of the chromatic difference block, the intra-prediction unit 216 can also predict the chromatic difference component of the current block based on the luma component of the current block.

[0224] Furthermore, when the information read from the encoded bitstream represents PDPC, the intra-prediction unit 216 corrects the pixel values ​​after intra-prediction based on the gradient of the reference pixel in the horizontal / vertical direction.

[0225] [Inter-frame prediction department]

[0226] The inter-frame prediction unit 218 refers to a reference image stored in the frame memory 214 and predicts the current block. Prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4×4 blocks). For example, the inter-frame prediction unit 218 uses motion information (e.g., motion vectors) read from the coded bitstream to perform motion compensation, thereby generating an inter-frame prediction signal for the current block or sub-block, and outputs the inter-frame prediction signal to the prediction control unit 220.

[0227] Furthermore, when the information read from the encoded bitstream is represented in OBMC mode, the inter-frame prediction unit 218 uses not only the motion information of the current block obtained through motion estimation, but also the motion information of adjacent blocks to generate the inter-frame prediction signal.

[0228] Furthermore, when the information read from the coded bitstream is in FRUC mode, the inter-frame prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) read from the coded stream, thereby deriving motion information. The inter-frame prediction unit 218 then uses the derived motion information to perform motion compensation.

[0229] Furthermore, when using BIO mode, the inter-frame prediction unit 218 derives motion vectors based on a model assuming constant-velocity linear motion. Additionally, when the information representation read from the encoded bitstream employs affine motion compensation prediction mode, the inter-frame prediction unit 218 derives motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0230] [Forecasting and Control Department]

[0231] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208.

[0232] [Method 1]

[0233] The encoding device 100, decoding device 200, encoding method, and decoding method of the first embodiment of the present invention will be described below.

[0234] [Example 1 of encoding and decoding using quantization matrices]

[0235] Figure 11 This diagram illustrates a first example of an encoding process using the quantization matrix (QM) in the encoding device 100. Furthermore, the encoding device 100 described here performs encoding processing on each block of a square or rectangle that divides the image (hereinafter also referred to as a frame) contained in the moving image.

[0236] First, in step S101, the quantization unit 108 generates a QM for square blocks. The QM for square blocks is a quantization matrix of multiple transform coefficients for square blocks. Hereinafter, the QM for square blocks will also be referred to as the first quantization matrix. Furthermore, the quantization unit 108 can generate the QM for square blocks based on user-defined values ​​set in the encoding device 100, or it can adaptively generate it using the encoding information of already encoded images. Furthermore, the entropy encoding unit 110 can record signals related to the QM for square blocks generated by the quantization unit 108 in the bitstream. At this time, the QM for square blocks can also be encoded as a region storing the sequence header region, image header region, slice header region, auxiliary information region, or other parameters of the stream. Furthermore, the QM for square blocks may not be recorded in the stream. At this time, the quantization unit 108 may also use the default value of the QM for square blocks predefined in the standard. Furthermore, the entropy encoding unit 110 may not record all the coefficients (i.e., quantization coefficients) of the matrix of the QM for square blocks in the stream, and may only record the partial quantization coefficients required to generate the QM in the stream. This reduces the amount of information encoded.

[0237] Next, in step S102, the quantization unit 108 uses the QM for the square block generated in step S101 to generate a QM for the rectangular block. The QM for the rectangular block is a quantization matrix for multiple transform coefficients of the rectangular block. Hereinafter, the QM for the rectangular block will also be referred to as the second quantization matrix. Furthermore, the entropy encoding unit 110 does not record the QM signal for the rectangular block in the stream.

[0238] Furthermore, the processing in steps S101 and S102 can be performed centrally at the start of sequence processing, image processing, or slice processing, or it can be performed in parts at a time during block unit processing. Additionally, the QM generated by the quantization unit 108 in steps S101 and S102 can be configured to generate multiple QMs for blocks of the same block size based on conditions such as luminance block / chrominance block, intra-frame prediction block / inter-frame prediction block, and others.

[0239] Next, the block-unit loop begins. First, in step S103, the intra-frame prediction unit 124 or the inter-frame prediction unit 126 performs prediction processing using intra-frame prediction or inter-frame prediction, etc., on a block-unit basis. In step S104, the transform unit 106 performs transform processing using discrete cosine transform (DCT) or the like on the generated prediction residual image. In step S105, the quantization unit 108 quantizes the generated transform coefficients using the outputs of steps S101 and S102, namely the QM for square blocks and the QM for rectangular blocks. Furthermore, in inter-frame prediction, the block pattern within the image to which the target block belongs can be used together with the block pattern within the image to which the target block belongs. In this case, the QM for inter-frame prediction can be used together, or the QM for intra-frame prediction can be used specifically for the block pattern within the image to which the target block belongs. Furthermore, in step S106, the inverse quantization unit 112 uses the QM for square blocks and the QM for rectangular blocks, which are the outputs of steps S101 and S102, to perform inverse quantization processing on the quantized transform coefficients. In step S107, the inverse transform unit 114 generates a residual (prediction error) image by performing inverse transform processing on the inverse quantized transform coefficients. Next, in step S108, the addition unit 116 generates a reconstructed image by adding the residual image and the prediction image. This series of processing flows is repeated to end the block unit loop.

[0240] Therefore, even in encoding methods with rectangular blocks of various shapes, the QM corresponding to each rectangular block shape is not recorded in the stream; only the QM corresponding to square blocks is recorded in the stream, thus enabling encoding processing. That is, according to the encoding apparatus 100 of the first aspect of the present invention, the QM corresponding to rectangular blocks is not recorded in the stream, thereby reducing the encoding amount in the header region. Furthermore, according to the encoding apparatus 100 of the first aspect of the present invention, the QM corresponding to rectangular blocks can be generated based on the QM corresponding to square blocks, thus allowing the appropriate QM to be used for rectangular blocks without increasing the encoding amount in the header region. Therefore, according to the encoding apparatus 100 of the first aspect of the present invention, it is possible to efficiently quantize rectangular blocks of various shapes, thus increasing the possibility of improving encoding efficiency. Moreover, the QM for square blocks may not be recorded in the stream, or a default value for the QM for square blocks predefined in the standard may be used.

[0241] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0242] Figure 12 It indicates that it was used in Figure 11 This diagram illustrates an example of the decoding process of the quantization matrix (QM) in the decoding device 200 corresponding to the encoding device 100 described herein. Furthermore, in the decoding device 200 described here, decoding processing is performed on each square or rectangular block that divides the image.

[0243] First, in step S201, the entropy decoding unit 202 decodes the signal related to the QM for square blocks from the stream, and uses the signal related to the decoded QM for square blocks to generate the QM for square blocks. Alternatively, the QM for square blocks can be decoded from the sequence header region, image header region, slice header region, auxiliary information region, or region storing other parameters of the stream. Furthermore, the QM for square blocks may not be decoded from the stream. In this case, the default value predefined in the standard can be used as the QM for square blocks. Additionally, the entropy decoding unit 202 may not decode all the quantization coefficients of the matrix of the QM for square blocks from the stream, but only decode a portion of the quantization coefficients required to generate the QM from the stream to generate the QM.

[0244] Next, in step S202, the entropy decoding unit 202 uses the QM signal for the square block generated in step S201 to generate a QM signal for the rectangular block. Furthermore, the entropy decoding unit 202 does not decode the QM signal for the rectangular block from the stream.

[0245] Furthermore, the processing in steps S201 and S202 can be performed centrally at the start of sequence processing, image processing, or slice processing, or it can be performed in parts at a time during block unit processing. Additionally, the QM generated by the entropy decoding unit 202 in steps S201 and S202 can be configured to generate multiple QMs for blocks of the same block size based on conditions such as luminance block / chrominance block, intra-frame prediction block / inter-frame prediction block, and others.

[0246] Next, the block-unit loop begins. First, in step S203, the intra-frame prediction unit 216 or the inter-frame prediction unit 218 performs prediction processing using intra-frame prediction or inter-frame prediction, etc., on a block-unit basis. In step S204, the inverse quantization unit 204 performs inverse quantization processing on the quantized transform coefficients (i.e., quantization coefficients) after decoding from the stream, using the QM for square blocks and the QM for rectangular blocks, which are the outputs of steps S201 and S202. Furthermore, in inter-frame prediction, the block pattern within the image to which the target block belongs can be used together with the block pattern within the image to which the target block belongs. At this time, the QM for inter-frame prediction can be used for both patterns, or the QM for intra-frame prediction can be used for the block pattern within the image to which the target block belongs. Next, in step S205, the inverse transform unit 206 generates a residual (prediction error) image by performing inverse transform processing on the inverse quantized transform coefficients. Next, in step S206, the addition unit 208 generates a reconstructed image by adding the residual image and the predicted image. This series of processing steps is repeated to end the block unit loop.

[0247] Therefore, even in decoding methods involving rectangular blocks of various shapes, even if the QM corresponding to each shape of rectangular block is not recorded in the stream, decoding can be performed as long as the QM corresponding only to square blocks is recorded in the stream. That is, according to the decoding apparatus 200 of the first aspect of the present invention, since the QM corresponding to rectangular blocks is not recorded in the stream, the amount of encoding in the header region can be reduced. In addition, according to the decoding apparatus 200 of the first aspect of the present invention, since the QM corresponding to rectangular blocks can be generated based on the QM corresponding to square blocks, the appropriate QM can be used for rectangular blocks without increasing the amount of encoding in the header region. Therefore, according to the decoding apparatus 200 of the first aspect of the present invention, it is possible to efficiently quantize rectangular blocks of various shapes, thus increasing the possibility of improving encoding efficiency. Furthermore, the QM for square blocks may not be recorded in the stream, or a default QM for square blocks predefined in the standard may be used.

[0248] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0249] [The first example of the QM generation method used for the rectangular block in the first example]

[0250] Figure 13 It is used to explain in Figure 11 Step S102 and Figure 12 In step S202, a diagram showing the first example of generating a QM for a rectangular block based on the QM for a square block is provided. Furthermore, the process described here is a common process in both the encoding device 100 and the decoding device 200.

[0251] exist Figure 13 In this process, for each square block with dimensions ranging from 2×2 to 256×256, a corresponding QM for the rectangular block generated from the QM of the square block of each size will be established and recorded. Figure 13 In the example shown, the characteristic is that the length of the long side of each rectangular block is the same as the length of the 1-sided side of the corresponding square block. In other words, in this example, the characteristic is that the size of the rectangular block, which is the object of processing, is smaller than the size of the square block. That is, the encoding device 100 and the decoding device 200 according to the first aspect of the present invention generate the QM for the rectangular block by down-converting the QM for the square block, which has a 1-sided side having the same length as the long side of the rectangular block, which is the object of processing.

[0252] In addition, Figure 13 The document records the correspondence between QMs for square blocks of various sizes and QMs for rectangular blocks generated from each square block QM, without distinguishing between luma and chroma blocks. The correspondence between square block QMs and rectangular block QMs can also be appropriately derived to adapt to the actual format used. For example, in a 4:2:0 format, the luma block is twice the size of the chroma block. Therefore, when referring to a luma block in the process of generating a rectangular block QM from a square block QM, the usable square block QMs correspond to square blocks of sizes from 4×4 to 256×256. In this case, among the rectangular block QMs generated from a square block QM, only QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used. Furthermore, when referring to a chroma block in the process of generating a rectangular block QM from a square block QM, the usable square block QMs correspond to square blocks of sizes from 2×2 to 128×128. At this point, in the QM for rectangular blocks generated from the QM for square blocks, only the QM corresponding to rectangular blocks with a length of 2 or more for the short side and a length of 128 or less for the long side is used.

[0253] Additionally, for example, in the 4:4:4 format, the luma block is the same size as the chroma block. Therefore, in the process of generating the QM for rectangular blocks, the case of referencing the chroma block is the same as the case of referencing the luma block, and the QM for square blocks that can be used corresponds to square blocks of sizes from 4×4 to 256×256.

[0254] In this way, the correspondence between the QM used for square blocks and the QM used for rectangular blocks can be appropriately derived according to the actual format used.

[0255] in addition, Figure 13 The block size described is one example, and is not limited to this. For example, it can also be used... Figure 13 QM for block sizes other than those shown can also be used only. Figure 13 QM is used for square blocks of a portion of the block sizes shown in the examples.

[0256] Figure 14 It is used to illustrate that by in Figure 13 The diagram illustrates the method of generating rectangular blocks by down-converting the QM used for the corresponding square blocks.

[0257] exist Figure 14 In the example, the QM used to generate 8×4 rectangular blocks is used from 8×8 square blocks.

[0258] In the down-conversion process, multiple matrix elements of the QM used for square blocks can also be divided into the same number of groups as the multiple matrix elements of the QM used for rectangular blocks. For each of the multiple groups, the multiple matrix elements contained in the group are arranged continuously in the horizontal or vertical direction of the square block. For each of the multiple groups, the matrix element located on the lowest domain side of the multiple matrix elements contained in the group is determined as the matrix element corresponding to the group in the QM used for rectangular blocks.

[0259] For example, in Figure 14 In the diagram, multiple matrix elements of QM are used to enclose a specified number of 8×8 square blocks, each enclosed by a thick line. The specified number of matrix elements enclosed by this thick line constitute one group. Figure 14 In the illustrated downconversion process, the 8×8 square block QM is segmented in such a way that the number of these groups is the same as the number of matrix elements (also called quantization coefficients) of the rectangular block QM generated from the QM used for 8×8 square blocks. Figure 14 In the example, two quantization coefficients adjacent in the vertical direction form a group. Then, in the QM used for the 8×8 square block, the lowest domain side in each group is selected (in... Figure 14In the example, the quantization coefficient (top side) is used as the QM value for an 8×4 rectangular block.

[0260] Furthermore, the method for selecting one quantization coefficient from each group as the QM value for the rectangular block is not limited to the examples above, and other methods can also be used. For example, the quantization coefficient located at the highest domain within a group can be used as the QM value for the rectangular block, or the quantization coefficient located in the middle domain can be used as the QM value for the rectangular block. Alternatively, the average, minimum, maximum, or median value of all or a portion of the quantization coefficients within a group can be used. Additionally, if the calculated result of these values ​​results in a decimal, it can be rounded up, down, or to the nearest integer using methods such as rounding.

[0261] Furthermore, the method of selecting one quantization coefficient from each group in the QM used for the square block can also be switched according to the frequency domain in which each group is located in the QM used for the square block. For example, the quantization coefficient located on the lowest domain side of the group can be selected in the low domain group, the quantization coefficient located on the highest domain side of the group can be selected in the high domain group, and the quantization coefficient located in the middle domain of the group can be selected in the middle domain group.

[0262] Alternatively, the generated rectangular block may not use the lowest domain component of QM ( Figure 14 The top left quantization coefficient in the example is derived from the QM used for the square block, but the structure can be directly set from the stream and described in the stream. In this case, the amount of information described in the stream increases, so the amount of encoding in the header region increases, but since the quantization coefficient of the QM of the lowest domain component that has the greatest impact on image quality can be directly controlled, the possibility of improving image quality becomes higher.

[0263] Additionally, this example illustrates the case where a QM for square blocks is down-converted vertically to generate a QM for rectangular blocks. However, in the case where a QM for square blocks is down-converted horizontally to generate a QM for rectangular blocks, the same method can also be used. Figure 14 The same method is used for the example.

[0264] [This is the second example of the QM generation method used in the first example of the rectangular block]

[0265] Figure 15 It is used to explain in Figure 11 Step S102 and Figure 12 The diagram for the second example of generating a QM for a rectangular block based on the QM for square blocks in step S202. Furthermore, the process described here is a common process in both the encoding device 100 and the decoding device 200.

[0266] exist Figure 15 In this process, for each square block with dimensions ranging from 2×2 to 256×256, a corresponding QM for the rectangular block generated from the QM of the square block of each size will be established and recorded. Figure 15 In the example shown, the characteristic is that the length of the short side of each rectangular block is the same as the length of the 1-sided side of the corresponding square block. In other words, in this example, the characteristic is that the size of the rectangular block, which is the object of processing, is larger than the size of the square block. That is, in the encoding apparatus 100 and decoding apparatus 200 according to the first aspect of the present invention, the QM for the rectangular block is generated by upconverting the QM for the square block, which has a 1-sided side having the same length as the short side of the rectangular block, which is the object of processing.

[0267] In addition, Figure 15 The document records the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each square block QM, without distinguishing between luma and chroma blocks. The correspondence between square block QMs and rectangular block QMs can also be appropriately derived to adapt to the actual format used. For example, in the 4:2:0 format, when referring to luma blocks in the process of generating rectangular block QMs from square block QMs, only QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used in the generated rectangular block QMs. Similarly, when referring to chroma blocks in the process of generating rectangular block QMs from square block QMs, only QMs corresponding to rectangular blocks with a short side length of 2 or more and a long side length of 128 or less are used in the generated rectangular block QMs. Furthermore, for the 4:4:4 format... Figure 13 The content described herein is the same, therefore the explanation is omitted here.

[0268] In this way, the correspondence between the QM used for square blocks and the QM used for rectangular blocks can be appropriately derived according to the actual format used.

[0269] in addition, Figure 15 The block size described is one example, and is not limited to this. For example, it can also be used... Figure 15 For square blocks with block sizes other than those shown, QM can also be used only. Figure 15 QM is used for square blocks of a portion of the block sizes shown in the examples.

[0270] Figure 16 It is used to illustrate that by in Figure 15 The diagram illustrates the method of generating rectangular blocks by upconverting the QM used for the corresponding square blocks.

[0271] exist Figure 16 In the example, QM is used to generate 8×4 rectangular blocks from 4×4 square blocks.

[0272] In the upconversion process, it is also possible to (i) divide the multiple matrix elements of the QM for rectangular blocks into the same number of groups as the multiple matrix elements of the QM for square blocks, and for each of the multiple groups, determine the matrix element in the QM for rectangular blocks corresponding to that group by repeating the multiple matrix elements contained in that group, or (ii) determine the multiple matrix elements of the QM for rectangular blocks by performing linear interpolation between adjacent matrix elements in the multiple matrix elements of the QM for rectangular blocks.

[0273] For example, in Figure 16 In the diagram, multiple matrix elements of the QM used for an 8×4 rectangular block are enclosed by thick lines in predetermined numbers. The predetermined number of matrix elements enclosed by these thick lines form a group. Figure 16 In the illustrated upconversion process, the QM for an 8×4 rectangular block is segmented in such a way that the number of these groups is the same as the number of matrix elements (also called quantization coefficients) of the QM used for the corresponding 4×4 square block. Figure 16 In the example, two quantization coefficients adjacent in the left and right directions form a group. Then, in the QM used for 8×4 rectangular blocks, the quantization coefficients that make up each group are selected from the QM used for square blocks corresponding to that group and filled in that group, thus becoming the QM value used for 8×4 rectangular blocks.

[0274] Furthermore, the method for deriving the quantization coefficients within each group of the QM used for the rectangular block is not limited to the examples mentioned above; other methods can also be used. For instance, linear interpolation can be performed by referring to the values ​​of the quantization coefficients in the adjacent frequency domain to derive continuous values ​​for the quantization coefficients within each group. Additionally, if the calculated results of these values ​​result in decimals, they can be rounded up, down, or to the nearest integer using methods such as rounding.

[0275] Furthermore, the method for deriving the quantization coefficients within each group in the QM used for rectangular blocks can be switched according to the frequency domain of each group in the QM used for rectangular blocks. For example, in groups located in the lower domain, the values ​​of each quantization coefficient within the group can be derived as small as possible, while in groups located in the higher domain, the values ​​of each quantization coefficient within the group can be derived as large as possible.

[0276] Alternatively, the generated rectangular block may not use the lowest domain component of QM ( Figure 16The top left quantization coefficient in the example is derived from the QM used for the square block, but the structure can be directly set from the stream and described in the stream. In this case, the amount of information described in the stream increases, so the amount of encoding in the header region increases, but since the quantization coefficient of the QM of the lowest domain component that has the greatest impact on image quality can be directly controlled, the possibility of improving image quality becomes higher.

[0277] Additionally, this example illustrates the case where a square block QM is vertically up-converted to generate a rectangular block QM. However, in the case where a square block QM is horizontally up-converted to generate a rectangular block QM, the same method can also be used. Figure 16 The same method is used for the example.

[0278] [Other variations of the first example of encoding and decoding using quantization matrices]

[0279] As a method for generating rectangular blocks from square blocks using QM, the appropriate QM can be switched depending on the size of the generated rectangular blocks. Figure 13 and Figure 14 The first example illustrating the rectangular block generation method using QM, and the use of... Figure 15 and Figure 16 The illustration provides a second example of a method for generating QM for rectangular blocks. For instance, there exists a method that compares the ratio of the rectangle's width to its height (the down-conversion or up-conversion ratio) to a threshold; if the ratio is greater than the threshold, the first example is used; if it's less, the second example is used. Alternatively, there exists a method that records and switches between flags indicating which method (first or second example) is used for each size of the rectangle in the stream. Thus, it's possible to switch down-conversion and up-conversion processing based on the size of the rectangle, thereby generating a more appropriate QM for the rectangle.

[0280] Alternatively, instead of switching up and down conversion for each size of the rectangular block, you can combine up and down conversion for a single rectangular block. For example, you can upconvert the QM for a 32×32 square block horizontally to generate the QM for a 32×64 rectangular block, and then downconvert the QM for the 32×64 rectangular block vertically to generate the QM for a 16×64 rectangular block.

[0281] Alternatively, a rectangular block can be upconverted in two directions. For example, a QM block for a 16×16 square block can be upconverted vertically to generate a QM block for a 32×16 rectangle block. Then, the QM block for the 32×16 rectangle block can be upconverted horizontally to generate a QM block for a 32×64 rectangle block.

[0282] Alternatively, a rectangular block can be down-converted in two directions. For example, a QM block for a 64×64 square block can be down-converted horizontally to generate a QM block for a 64×32 rectangle block. Then, the QM block for the 64×32 rectangle block can be down-converted vertically to generate a QM block for a 16×32 rectangle block.

[0283] [The effect of the first example of encoding and decoding using a quantization matrix]

[0284] According to the first aspect of the present invention, the encoding device 100 and the decoding device 200, by using Figure 11 and Figure 12 The described structure, even in encoding methods with rectangular blocks of various shapes, does not record the QM corresponding to each rectangular block shape in the stream, but only records the QM corresponding to square blocks, thereby enabling encoding and decoding of rectangular blocks. That is, since the encoding apparatus 100 and decoding apparatus 200 according to the first aspect of the present invention do not record the QM corresponding to rectangular blocks in the stream, the amount of encoding in the header region can be reduced. Furthermore, the encoding apparatus 100 and decoding apparatus 200 according to the first aspect of the present invention can generate the QM corresponding to rectangular blocks based on the QM corresponding to square blocks, thus allowing the appropriate QM to be used for rectangular blocks without increasing the amount of encoding in the header region. Therefore, the encoding apparatus 100 and decoding apparatus 200 according to the first aspect of the present invention can efficiently quantize rectangular blocks of various shapes, thus increasing the possibility of improving encoding efficiency.

[0285] For example, the encoding device 100 is an encoding device that performs quantization to encode a moving image, and includes circuitry and a memory. The circuitry uses the memory to transform a first quantization matrix for a plurality of transform coefficients of a square block, thereby generating a second quantization matrix for a plurality of transform coefficients of a rectangular block based on the first quantization matrix, and using the second quantization matrix to quantize the plurality of transform coefficients of the rectangular block.

[0286] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, it is also possible to avoid encoding the quantization matrix corresponding to the rectangular block. Thus, processing efficiency is improved due to the reduction in encoding. Therefore, the encoding device 100 can efficiently quantize rectangular blocks.

[0287] For example, in the encoding device 100, the circuit may encode only the first quantization matrix from the first quantization matrix and the second quantization matrix into the bit stream.

[0288] This reduces the amount of code. Therefore, according to the encoding device 100, processing efficiency is improved.

[0289] For example, in the encoding device 100, the circuit may perform downconversion processing based on the first quantization matrix to generate the second quantization matrix for multiple transform coefficients of the rectangular block, wherein the first quantization matrix is ​​for multiple transform coefficients of the square block having one side of the same length as the long side of the rectangular block which is the processing target block.

[0290] Thus, the encoding device 100 can efficiently generate a quantization matrix corresponding to a square block, the square block having one side of the same length as the long side of the rectangular block.

[0291] For example, in the encoding device 100, the circuit may, during the downconversion process, divide the plurality of matrix elements of the first quantization matrix into the same number of groups as the plurality of matrix elements of the second quantization matrix. For each of the plurality of groups, the plurality of matrix elements contained in the group are arranged continuously in the horizontal or vertical direction of the square block. For each of the plurality of groups, the matrix element located on the lowest domain side of the plurality of matrix elements contained in the group, the matrix element located on the highest domain side of the plurality of matrix elements contained in the group, or the average value of the plurality of matrix elements contained in the group is determined in the second quantization matrix as the matrix element corresponding to the group.

[0292] Therefore, the encoding device 100 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0293] For example, in the encoding device 100, the circuit may perform an upconversion process on the first quantization matrix for the multiple transform coefficients of the square block to generate the second quantization matrix for the multiple transform coefficients of the rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0294] Thus, the encoding device 100 can efficiently generate a quantization matrix corresponding to a square block, the square block having one side of the same length as the short side of the rectangular block.

[0295] For example, in the encoding device 100, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix element in the second quantization matrix corresponding to that group by repeating the plurality of matrix elements contained in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements in the plurality of matrix elements of the second quantization matrix.

[0296] Therefore, the encoding device 100 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0297] For example, in the encoding device 100, the circuit may generate the second quantization matrix of multiple transform coefficients for the rectangle by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block, which is the processing target block: a method of generating the matrix by performing the downconversion process based on the first quantization matrix of multiple transform coefficients for the square block having one side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing the upconversion process based on the first quantization matrix of multiple transform coefficients for the square block having one side with the same length as the short side of the rectangular block.

[0298] Therefore, by switching downconversion and upconversion according to the block size of the object block being processed, the encoding device 100 is able to generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0299] In addition, the decoding device 200 is a decoding device that performs inverse quantization to decode a moving image. It includes circuitry and a memory. The circuitry uses the memory to transform a first quantization matrix for multiple transform coefficients of a square block, thereby generating a second quantization matrix for multiple transform coefficients of a rectangular block based on the first quantization matrix. The second quantization matrix is ​​then used to perform inverse quantization on the multiple quantization coefficients of the rectangular block.

[0300] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, decoding of the quantization matrix corresponding to the rectangular block is not required. This reduces the amount of coding, thus improving processing efficiency. Therefore, according to the decoding device 200, the rectangular block can be efficiently dequantized.

[0301] For example, in the decoding device 200, the circuit may decode only the first quantization matrix from the bit stream and the first quantization matrix from the second quantization matrix.

[0302] This reduces the amount of code. Therefore, according to the decoding device 200, processing efficiency is improved.

[0303] For example, in the decoding device 200, the circuit may perform downconversion processing based on the first quantization matrix to generate the second quantization matrix for multiple transform coefficients of the rectangular block, wherein the first quantization matrix is ​​for multiple transform coefficients of the square block having one side of the same length as the long side of the rectangular block which is the processing target block.

[0304] Thus, the decoding device 200 can efficiently generate a quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to the square block, wherein the square block has one side of the same length as the long side of the rectangular block.

[0305] For example, in the decoding device 200, the circuit may, during the downconversion process, divide the plurality of matrix elements of the first quantization matrix into the same number of groups as the plurality of matrix elements of the second quantization matrix. For each of the plurality of groups, the matrix element located on the lowest domain side of the plurality of matrix elements contained in the group, the transform coefficient located on the highest domain side of the plurality of matrix elements contained in the group, or the average value of the plurality of matrix elements contained in the group is determined as the matrix element in the second quantization matrix corresponding to that group.

[0306] Therefore, the decoding device 200 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0307] For example, in the decoding device 200, the circuit may generate the second quantization matrix for the multiple transform coefficients of the rectangular block by performing an upconversion process on the first quantization matrix for the multiple transform coefficients of the square block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0308] Thus, the decoding device 200 can efficiently generate a quantization matrix corresponding to a rectangular block from the quantization matrix corresponding to the square block, the square block having one side of the same length as the short side of the rectangular block.

[0309] For example, in the decoding device 200, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix element in the second quantization matrix corresponding to that group by repeating the plurality of matrix elements contained in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements in the plurality of matrix elements of the second quantization matrix.

[0310] Therefore, the decoding device 200 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0311] For example, in the decoding device 200, the circuit may generate the second quantization matrix of multiple transform coefficients for the rectangle by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangle as the processing object block: a method of generating the matrix by performing the downconversion process based on the first quantization matrix of multiple transform coefficients for the square block having one side with the same length as the long side of the rectangle; and a method of generating the matrix by performing the upconversion process based on the first quantization matrix of multiple transform coefficients for the square block having one side with the same length as the short side of the rectangle.

[0312] Therefore, by switching between downconversion and upconversion according to the block size of the block being processed, the decoding device 200 can generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0313] In addition, the encoding method is an encoding method for encoding motion images by performing quantization. It transforms a first quantization matrix for multiple transform coefficients of a square block, generates a second quantization matrix for multiple transform coefficients of a rectangular block based on the first quantization matrix, and uses the second quantization matrix to quantize the multiple transform coefficients of the rectangular block.

[0314] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, it is also possible to avoid encoding the quantization matrix corresponding to the rectangular block. Thus, processing efficiency is improved due to the reduction in encoding. Therefore, based on the encoding method, rectangular blocks can be quantized efficiently.

[0315] In addition, the decoding method is a decoding method that performs inverse quantization to decode the motion image. It transforms a first quantization matrix for multiple transform coefficients of a square block, generates a second quantization matrix for multiple transform coefficients of a rectangular block based on the first quantization matrix, and uses the second quantization coefficients to perform inverse quantization on the multiple quantization coefficients of the rectangular block.

[0316] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, decoding of the quantization matrix corresponding to the rectangular block is unnecessary. This reduces the amount of coding, thus improving processing efficiency. Therefore, based on the decoding method, the rectangular block can be efficiently dequantized.

[0317] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.

[0318] [Method 2]

[0319] The encoding device 100, decoding device 200, encoding method, and decoding method of the second aspect of the present invention will be described below.

[0320] [Example 2 of encoding and decoding using quantization matrices]

[0321] Figure 17 This is a second example of an encoding process using the quantization matrix (QM) in the encoding device 100. Furthermore, in the encoding device 100 described here, each block of a square or rectangle that divides the screen is encoded.

[0322] First, in step S701, the quantization unit 108 generates a QM corresponding to the size of the effective transform coefficient region in each block size of the square block and the rectangular block. In other words, the quantization unit 108 uses a quantization matrix to quantize only the multiple transform coefficients in a specified region on the low-frequency domain side of the multiple transform coefficients contained in the processing target block.

[0323] The entropy encoding unit 110 records the signal related to the QM corresponding to the valid transform coefficient region generated in step S701 in the stream. In other words, the entropy encoding unit 110 encodes the signal related to the quantization matrix corresponding only to a plurality of transform coefficients within a specified range in the low-frequency domain into the bit stream. Furthermore, the quantization unit 108 can generate the QM corresponding to the valid transform coefficient region based on user-defined values ​​set in the encoding device 100, or it can adaptively generate it using the encoding information of already encoded images. Additionally, the QM corresponding to the valid transform coefficient region can also be encoded into the sequence header region, image header region, slice header region, auxiliary information region, or other parameter region of the stored stream. Alternatively, the QM corresponding to the valid transform coefficient region may not be recorded in the stream. In this case, the quantization unit 108 may also use a default value predefined in the standard as the value of the QM corresponding to the valid transform coefficient region.

[0324] Furthermore, the processing in step S701 can be a structure that is performed centrally at the start of sequence processing, or at the start of image processing, or at the start of slice processing, or it can be a structure that performs partial processing each time in block unit processing. In addition, the QM generated in step S701 can also be configured to generate multiple QMs for blocks of the same block size based on luminance block / chrominance block, intra-frame prediction block / inter-frame prediction block, and other conditions.

[0325] In addition, Figure 17 In the processing flow shown, the processing steps other than step S701 are block-based loop processing, which is consistent with the use of... Figure 11 The same applies to the first example described.

[0326] Therefore, for a processing object block that sets only a portion of the low-frequency domain side of the multiple transform coefficients contained in the processing object block to a block size containing only the region of valid transform coefficients, it is possible to encode the QM-related signals of the invalid regions in the stream without waste. Thus, the amount of coding in the header region can be reduced, thereby increasing the potential for improved coding efficiency.

[0327] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0328] Figure 18 It indicates that it was used in Figure 17 This diagram illustrates an example of the decoding process of the quantization matrix (QM) in the decoding device 200 corresponding to the encoding device 100 described herein. Furthermore, in the decoding device 200 described here, decoding processing is performed on each square or rectangular block that divides the image.

[0329] First, in step S801, the entropy decoding unit 202 decodes the QM-related signal corresponding to the valid transform coefficient region from the stream, and uses the decoded QM-related signal to generate the QM corresponding to the valid transform coefficient region. The QM corresponding to the valid transform coefficient region is the QM corresponding to the size of the valid transform coefficient region in each block size of the processing object block. Alternatively, the QM corresponding to the valid transform coefficient region can also be decoded from the sequence header region, image header region, slice header region, auxiliary information region, or other parameter region of the stored stream. Alternatively, the QM corresponding to the valid transform coefficient region may not be decoded from the stream. In this case, for example, a default value predefined in the standard can be used as the QM corresponding to the valid transform coefficient region.

[0330] Furthermore, the processing in step S801 can be performed centrally at the start of sequence processing, image processing, or slice processing, or it can be performed in parts at a time during block unit processing. Additionally, the QM generated by the entropy decoding unit 202 in step S801 can be configured to generate multiple QMs for blocks of the same block size based on conditions such as luminance block / chrominance block, intra-frame prediction block / inter-frame prediction block, and others.

[0331] In addition, Figure 18 In the processing flow shown, the processing flow other than step S801 is a block-based loop processing, which is similar to using... Figure 12 The processing procedure for the first example described is the same.

[0332] Therefore, for a block size that includes only a portion of the low-frequency domain side of the multiple transform coefficients contained in the processing block as the region containing valid transform coefficients, decoding can be performed even without wasting the QM-related signals of invalid regions recorded in the stream. Thus, the amount of coding in the header region can be reduced, thereby increasing the potential for improved coding efficiency.

[0333] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0334] Figure 19 It is used to explain in Figure 17 Step S701 and Figure 18 In step S801, a diagram showing an example of the QM corresponding to the size of the effective transform coefficient region in each block size is provided. Furthermore, the process described here is a common process in both the encoding device 100 and the decoding device 200.

[0335] Figure 19(a) represents an example of a 64×64 square block being processed. Only the 32×32 region on the lower domain side, indicated by the diagonal line in the figure, contains valid transform coefficients. In the processed block, the regions outside this valid transform coefficient region are forcibly made to have 0 transform coefficients, i.e., the transform coefficients are invalidated, thus eliminating the need for quantization and inverse quantization. That is, the encoding apparatus 100 and decoding apparatus 200 according to the second aspect of the present invention only generate a 32×32 QM corresponding to the 32×32 region on the lower domain side indicated by the diagonal line in the figure.

[0336] then, Figure 19 (b) represents an example of the case where the object block being processed is a rectangular block with a size of 64×32. Figure 19 (b) and Figure 19 Similarly, in example (a), the encoding device 100 and the decoding device 200 generate only a 32×32 QM corresponding to the 32×32 region on the low-domain side.

[0337] then, Figure 19 (c) represents an example of processing a rectangular block with a block size of 64×16. Figure 19 (c) and Figure 19 Unlike example (a), the vertical block size is only 16. Therefore, the encoding device 100 and the decoding device 200 only generate a 32×16 QM corresponding to the 32×16 region on the low domain side.

[0338] In this way, if any side of the block being processed is larger than 32, the transform coefficients of the region larger than 32 are set to invalid, and only the region below 32 is set to the region of valid transform coefficients. This region is then set as the processing object for quantization and inverse quantization, and the quantization coefficients of QM are generated, as well as the encoding and decoding of the signal stream related to QM are performed.

[0339] Therefore, encoding and decoding of QM-related signals that become invalid regions can be recorded in the stream without waste, thus reducing the amount of coding in the header region. This increases the potential for improved coding efficiency.

[0340] In addition, Figure 19The size of the valid transformation coefficient region described is one example; other valid transformation coefficient region sizes can also be used. For example, when processing a luma block, a region up to 32×32 can be set as the valid transformation coefficient region; when processing a chroma block, a region up to 16×16 can be set as the valid transformation coefficient region. Furthermore, when the long side of the processed object block is 64, a region up to 32×32 can be set as the valid transformation coefficient region; when the long side of the processed object block is 128 or 256, a region up to 62×62 can be set as the valid transformation coefficient region.

[0341] Alternatively, it can be configured to use the same processing as described in the first method, generating only the coefficients of the quantization matrix corresponding to all frequency components of the square and rectangle after generating the coefficients of the quantization matrix corresponding to all frequency components of the square and rectangle. Figure 19 The QM corresponding to the effective transform coefficient region is described. In this case, the amount of signal associated with this QM recorded in the stream remains unchanged from the method described in the first method, and all square and rectangular QMs can be generated directly using the method described in the first method, omitting quantization processing outside the effective transform coefficient region. Therefore, the possibility of reducing the amount of processing related to quantization is increased.

[0342] [A variation of the second example of encoding and decoding using quantization matrices]

[0343] Figure 20 This diagram is a modified example of the second example of the encoding process using the quantization matrix (QM) in the encoding device 100. Furthermore, in the encoding device 100 described here, each block of a square or rectangle that divides the screen is encoded.

[0344] In this variation, for Figure 17 The structure of the second example described herein is a combination of... Figure 11 The structure of the first example described in the text is replaced. Figure 17 The process proceeds from step S701 to steps S1001 and S1002.

[0345] First, in step S1001, the quantization unit 108 generates a QM for the square block. At this time, the QM for the square block is the QM corresponding to the size of the effective transform coefficient region within the square block. Furthermore, the entropy encoding unit 110 records in the stream the signal related to the QM of the square block generated in step S1001. At this time, the QM-related signal recorded in the stream is the signal related only to the quantization coefficients corresponding to the effective transform coefficient region.

[0346] Next, in step S1002, the quantization unit 108 uses the QM for the square block generated in step S1001 to generate the QM for the rectangular block. Furthermore, at this time, the entropy encoding unit 110 does not record the signal related to the QM for the rectangular block in the stream.

[0347] In addition, Figure 20 In the processing flow shown, the processes other than steps S1001 and S1002 are block-based cyclic processes, similar to those using... Figure 11 The same applies to the first example described.

[0348] Therefore, in encoding methods for rectangular blocks of various shapes, the signals related to the QM corresponding to each rectangular block shape are not recorded in the stream; only the signals related to the QM corresponding to square blocks are recorded in the stream, thus enabling encoding processing. Furthermore, for processing blocks with block sizes where only a portion of the transform coefficients contained in the processing block are designated as valid regions (i.e., valid transform coefficient regions), encoding processing can be performed without waste by recording the signals related to the QM of invalid regions in the stream. Therefore, it is possible to use QM for rectangular blocks while reducing the encoding amount of the header region, thereby increasing the possibility of improving encoding efficiency.

[0349] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0350] Figure 21 It indicates that it was used in Figure 20 This diagram illustrates an example of the decoding process of the quantization matrix (QM) in the decoding device 200 corresponding to the encoding device 100 described herein. Furthermore, in the decoding device 200 described here, decoding processing is performed on each square or rectangular block that divides the image.

[0351] In this variation, for in Figure 18 The structure of the second example described in the text combines the elements of... Figure 12 The structure of the first example described in the text is replaced. Figure 18 The process proceeds from step S801 to steps S1101 and S1102.

[0352] First, in step S1101, the entropy decoding unit 202 decodes the signal related to the QM for square blocks from the stream, and uses the signal related to the decoded QM for square blocks to generate a QM for square blocks. At this time, the signal related to the QM for square blocks decoded from the stream is a signal that is only related to the quantization coefficients corresponding to the effective transform coefficient region. Therefore, the QM for square blocks generated by the entropy decoding unit 202 is a QM that corresponds to the size of the effective transform coefficient region.

[0353] Next, in step S1102, the entropy decoding unit 202 uses the QM for the square block generated in step S1101 to generate a QM for the rectangular block. Furthermore, at this time, the entropy decoding unit 202 does not decode the signal related to the QM for the rectangular block from the stream.

[0354] In addition, Figure 21 In the processing flow shown, the processes other than steps S1101 and S1102 are block-based cyclic processes, which are similar to those using... Figure 12 The same applies to the first example described.

[0355] Therefore, even in encoding schemes with rectangular blocks of various shapes, and even if no QM-related signals corresponding to each rectangular block shape are recorded in the stream, decoding can be performed as long as only the QM-related signals corresponding to square blocks are recorded in the stream. Furthermore, for processing blocks whose block size uses only a portion of the transform coefficients contained in the processing block as valid regions (i.e., valid transform coefficient regions), decoding can be performed even if the QM-related signals for invalid regions are recorded in the stream without waste. Therefore, the possibility of improving encoding efficiency increases because the QM can be used for rectangular blocks while reducing the encoding amount of the header region becomes more apparent.

[0356] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0357] [The first example of using QM's generation method for the rectangular block in the second variation]

[0358] Figure 22 It is used to explain in Figure 20 Step S1002 and Figure 21 In step S1102, a diagram showing the first example of generating a QM for a rectangular block based on the QM for a square block is provided. Furthermore, the process described here is a common process in both the encoding device 100 and the decoding device 200.

[0359] exist Figure 22In this process, for each square block with dimensions ranging from 2×2 to 256×256, a corresponding QM for the rectangular block generated from the QM of the square block of each size will be established and recorded. Figure 22 The example shown illustrates the size of the processing object block and the size of the effective transformation coefficient region within the processing object block. The values ​​within parentheses indicate the size of the effective transformation coefficient region within the processing object block. Furthermore, for a rectangular block whose size is equal to the size of the processing object block and the size of the effective transformation coefficient region within the processing object block, it becomes... Figure 13 The same treatment as the first example described herein, therefore in Figure 22 The corresponding table shown omits the numerical values.

[0360] Here, the key feature is that the length of the longest side of each rectangular block is the same as the length of the first side of the corresponding square block, and the rectangular blocks are smaller than the square blocks. That is, the QM for rectangular blocks is generated by down-converting the QM used for square blocks.

[0361] In addition, Figure 22 The document records the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each square block QM, without distinguishing between luma and chroma blocks. The correspondence between square block QMs and rectangular block QMs can also be appropriately derived to adapt to the actual format being used. For example, in a 4:2:0 format, the luma block is twice the size of the chroma block. Therefore, when referencing the luma block in the process of generating rectangular block QMs from square block QMs, the usable square block QMs correspond to square blocks of sizes from 4×4 to 256×256. In this case, among the rectangular block QMs generated from square block QMs, only the QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used. Additionally, when referencing color difference blocks during the process of generating a QM for rectangular blocks from a QM for square blocks, only the QM corresponding to rectangular blocks with a short side length of 2 or more and a long side length of 128 or less is used in the QM for rectangular blocks generated from the QM for square blocks. Furthermore, the 4:4:4 format case is different from... Figure 13 The content described in the text is the same.

[0362] In this way, the correspondence between the QM used for square blocks and the QM used for rectangular blocks can be appropriately derived according to the actual format used.

[0363] in addition, Figure 22 The size of the effective transformation coefficient region described in the document is one example; it can also be used... Figure 22 The size of the effective transformation coefficient region other than the illustrated size.

[0364] in addition, Figure 22 The block size described is one example, and is not limited to this. For example, it can also be used... Figure 22 Block sizes other than those shown can also be used simply. Figure 22 A portion of the block sizes shown.

[0365] Figure 23 It is used to illustrate that by in Figure 22 The diagram illustrates the method of generating rectangular blocks by down-converting the QM used for the corresponding square blocks.

[0366] exist Figure 23 In the example, based on the QM corresponding to the 32×32 valid transform coefficient region in the 64×64 square block, a QM corresponding to the 32×32 valid transform coefficient region in the 64×32 rectangular block is generated.

[0367] First, such as Figure 23 As shown in (a), by extending the quantization coefficients of the QM corresponding to the effective transform coefficient region of 32×32 in the vertical direction, a QM for generating a 64×64 square block in the middle of the effective region of 32×64 is generated. Methods for extending the quantization coefficients as described above include: extending the difference between the quantization coefficients of row 31 and row 32 in such a way that it becomes the difference between the adjacent coefficients thereafter; or deriving the change in the difference between the quantization coefficients of row 30 and row 31 and the difference between the quantization coefficients of row 31 and row 32, and extending the QM while correcting the difference between the adjacent quantization coefficients thereafter with the change in the change in the same way.

[0368] Next, as Figure 23 As shown in (b), use and use Figure 14 The same method is used to down-transform the QM used for the 64×64 square block in the middle of the 32×64 valid region, thereby generating the QM used for the 64×32 rectangular block. At this point, the resulting valid region becomes... Figure 23 The 64×32 rectangular block of (c) is a 32×32 region indicated by the diagonal lines in QM.

[0369] Furthermore, this example illustrates the case where a QM for square blocks with valid areas is down-converted vertically to generate a QM for rectangular blocks. However, in the case where a QM for square blocks with valid areas is down-converted horizontally to generate a QM for rectangular blocks, the same method can also be used. Figure 23 The same method is used for the example.

[0370] Additionally, this example illustrates the generation of a rectangular block QM via a QM using the central square block in a two-stage process. However, it is also possible to use the guide and... Figure 23 Examples of similar processing results, such as transformation formulas, directly generate QM for rectangular blocks based on QM for square blocks with valid regions.

[0371] [The second example of a variation using the QM generation method for the rectangular block]

[0372] Figure 24 It is used to explain in Figure 20 Step S1002 and Figure 21 The diagram shows a second example of generating a QM for a rectangular block based on the QM for a square block in step S1102. Furthermore, the process described here is a common process in both the encoding device 100 and the decoding device 200.

[0373] exist Figure 24 In this process, for each square block with dimensions ranging from 2×2 to 256×256, a corresponding QM for the rectangular block generated from the QM of the square block of each size will be established and recorded. Figure 24 The example shown illustrates the size of the processing object block and the size of the effective transformation coefficient region within the processing object block. The values ​​within parentheses indicate the size of the effective transformation coefficient region within the processing object block. Furthermore, for a rectangular block whose size is equal to the size of the processing object block and the size of the effective transformation coefficient region, it becomes... Figure 15 The same treatment as the first example described herein, therefore in Figure 24 The corresponding table shown omits the numerical values.

[0374] Here, the key feature is that the length of the short side of each rectangular block is the same as the length of one side of the corresponding square block, and the rectangular blocks are larger than the square blocks. That is, the QM used for the square blocks is generated by upconverting the QM used for the square blocks.

[0375] In addition, Figure 24The document records the correspondence between QMs for square blocks of various block sizes and QMs for rectangular blocks generated from each square block QM, without distinguishing between luma and chroma blocks. The correspondence between square block QMs and rectangular block QMs can also be appropriately derived to adapt to the actual format used. For example, in the 4:2:0 format, the luma block is twice the size of the chroma block. Therefore, when referring to the luma block in the process of generating rectangular block QMs from square block QMs, only the QMs corresponding to rectangular blocks with a short side length of 4 or more and a long side length of 256 or less are used in the rectangular block QMs generated from the square block QMs. Similarly, when referring to the chroma block in the process of generating rectangular block QMs from square block QMs, only the QMs corresponding to rectangular blocks with a short side length of 2 or more and a long side length of 128 or less are used in the rectangular block QMs generated from the square block QMs. Furthermore, for the 4:4:4 format... Figure 13 The content described in the text is the same.

[0376] In this way, the correspondence between the QM used for square blocks and the QM used for rectangular blocks can be appropriately derived according to the actual format used.

[0377] in addition, Figure 24 The size of the effective transformation coefficient region described in the document is one example; it can also be used... Figure 24 The size of the effective transformation coefficient region other than the illustrated size.

[0378] in addition, Figure 24 The block size described is one example, and is not limited to this. For example, it can also be used... Figure 24 Block sizes other than those shown can also be used simply. Figure 24 A portion of the block sizes shown.

[0379] Figure 25 It is used to illustrate that by in Figure 24 The diagram illustrates the method of generating rectangular blocks by upconverting the QM used for the corresponding square blocks.

[0380] exist Figure 25 In the example, based on the QM corresponding to the 32×32 valid transform coefficient region in the 32×32 square block, a QM corresponding to the 32×32 valid transform coefficient region in the 64×32 rectangular block is generated.

[0381] First, such as Figure 25 As shown in (a), use and use Figure 16The same method is used to upconvert the QM used for the 32×32 square block, thereby generating the QM used for the central 64×32 rectangular block. At this point, the valid area is also upconverted to 64×32.

[0382] Next, as Figure 25 As shown in (b), a QM is generated for a 64×32 rectangular block with a 32×32 effective region by only cutting the 32×32 on the low-domain side of the 64×32 effective region.

[0383] Additionally, this example illustrates the case where a QM for square blocks is up-converted horizontally to generate a QM for rectangular blocks. However, in the case where a QM for square blocks is up-converted vertically to generate a QM for rectangular blocks, the same method can also be used. Figure 25 The same method is used for the example.

[0384] Furthermore, this example illustrates the generation of the QM for the rectangular block in a two-stage process via the QM for the intermediate rectangular block, but it is also possible to use the guide and... Figure 25 Examples of similar processing results, such as transformation formulas, directly generate QM for rectangular blocks based on QM used for square blocks.

[0385] [Other variations of the second example of encoding and decoding using quantization matrices]

[0386] The method for generating rectangular blocks from square blocks using QM can also be switched based on the size of the generated rectangular blocks. Figure 22 and Figure 23 The first example illustrating the rectangular block generation method using QM, and the use of... Figure 24 and Figure 25 This is the second example of a method for generating a QM for rectangular blocks. For instance, there exists a method that compares the ratio of the rectangle's width to its height (the down-conversion or up-conversion ratio) to a threshold; if the ratio is greater than the threshold, the first example is used; if it's less than the threshold, the second example is used. Alternatively, there exists a method that records and switches between flags indicating which method (first or second example) is used for each size of the rectangle in the stream. Thus, it's possible to switch down-conversion and up-conversion processing based on the size of the rectangle, thereby generating a more appropriate QM for rectangular blocks.

[0387] [The second example of encoding and decoding using quantization matrices, and the effect of a variation of the second example]

[0388] According to the second aspect of the present invention, the encoding device 100 and the decoding device 200, by using Figure 17 and Figure 18 The described structure allows for the encoding and decoding of rectangular blocks where only a portion of the transform coefficients within the processing block are designated as valid regions. This enables the encoding and decoding of QM-related signals associated with invalid regions to be recorded in the stream without waste. Therefore, the amount of coding in the header region can be reduced, potentially increasing the coding efficiency.

[0389] Furthermore, according to the second embodiment of the present invention, the encoding device 100 and the decoding device 200, by using Figure 20 and Figure 21 The described structure, even in encoding methods with rectangular blocks of various shapes, does not record the QM corresponding to each rectangular block shape in the stream, but only records the QM corresponding to square blocks, thereby enabling encoding and decoding of rectangular blocks. That is, the encoding apparatus 100 and decoding apparatus 200 of the second embodiment of the present invention can generate the QM corresponding to the rectangular blocks based on the QM corresponding to the square blocks, thus enabling the use of appropriate QMs for the rectangular blocks while reducing the encoding amount in the head region. Therefore, the encoding apparatus 100 and decoding apparatus 200 of the second embodiment of the present invention can efficiently quantize rectangular blocks of various shapes, thus increasing the possibility of improving encoding efficiency.

[0390] For example, encoding device 100 is an encoding device that performs quantization to encode a moving image, and includes circuitry and a memory. The circuitry uses the memory to quantize only a specified range of multiple transform coefficients in the low-frequency domain of a plurality of transform coefficients contained in the processing object block using a quantization matrix.

[0391] Therefore, by using only the quantization matrix corresponding to a specified range of the low-frequency domain side that has a significant impact on visual perception within the processing object block, the image quality of the moving image is less likely to degrade. Furthermore, since only the quantization matrix corresponding to this specified range is encoded, the amount of encoding is reduced. Thus, according to the encoding device 100, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0392] For example, in the encoding device 100, the circuit may encode a signal associated with the quantization matrix of a plurality of transform coefficients within the predetermined range that corresponds only to the low-frequency domain side into a bitstream.

[0393] This reduces the amount of code. Therefore, according to the encoding device 100, processing efficiency is improved.

[0394] For example, in the encoding device 100, the processing target block may be a square block or a rectangular block. The circuit transforms a first quantization matrix for a plurality of transform coefficients within a specified range in the square block, generates a second quantization matrix for a plurality of transform coefficients within the specified range in the rectangular block based on the first quantization matrix, and encodes only the first quantization matrix from the first quantization matrix and the second quantization matrix as a signal associated with the quantization matrix into a bit stream.

[0395] Therefore, for a processing object block containing multiple shapes including rectangular blocks, the encoding device 100 generates a first quantization matrix corresponding to a defined range in the low-frequency domain of the square block, and encodes only the first quantization matrix, thus reducing the amount of encoding. Furthermore, the encoding device 100 generates a second quantization matrix based on the first quantization matrix, corresponding to a defined range in the low-frequency domain of the rectangular block, so the image quality of the moving image is less likely to degrade. Therefore, according to the encoding device 100, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0396] For example, in the encoding device 100, the circuit may perform downconversion processing based on a first quantization matrix for a plurality of transform coefficients within a specified range in the square block to generate the second quantization matrix for the plurality of transform coefficients within the specified range in the rectangular block, wherein the square block has one side of the same length as the long side of the rectangular block that is the processing target block.

[0397] Thus, the encoding device 100 is able to efficiently quantize rectangular blocks.

[0398] For example, in the encoding device 100, during the downconversion process, the circuit may extend the plurality of matrix elements of the first quantization matrix by extrapolating in a predetermined direction, and divide the plurality of matrix elements of the extended first quantization matrix into a number of groups equal to the number of matrix elements of the second quantization matrix. For each of the plurality of groups, the matrix element located on the lowest domain side of the plurality of matrix elements contained in the group, the matrix element located on the highest domain side of the plurality of matrix elements contained in the group, or the average value of the plurality of matrix elements contained in the group is determined as the matrix element in the second quantization matrix corresponding to that group.

[0399] Therefore, the encoding device 100 can quantize rectangular blocks more efficiently.

[0400] For example, in the encoding device 100, the circuit may perform upconversion processing based on a first quantization matrix for a plurality of transform coefficients within a specified range in the square block to generate a second quantization matrix for a plurality of transform coefficients within a specified range in the rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0401] Thus, the encoding device 100 is able to efficiently quantize rectangular blocks.

[0402] For example, in the encoding device 100, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, extend the first quantization matrix in the predetermined direction by repeating the plurality of matrix elements contained in the group in the predetermined direction, and extract the same number of matrix elements as the plurality of matrix elements of the second quantization matrix from the extended first quantization matrix, or (ii) extend the first quantization matrix in the predetermined direction by performing linear interpolation between matrix elements adjacent in the predetermined direction among the plurality of matrix elements of the second quantization matrix, and extract the same number of matrix elements as the plurality of matrix elements of the second quantization matrix from the extended first quantization matrix.

[0403] Therefore, the encoding device 100 can quantize rectangular blocks more efficiently.

[0404] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within a predetermined range in the rectangular block by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block, which is the processing target block: a method of generating the matrix by performing the downconversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing the upconversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the short side of the rectangular block.

[0405] Therefore, by switching downconversion and upconversion according to the block size of the object block being processed, the encoding device 100 is able to generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0406] In addition, the decoding device 200 is a decoding device that performs inverse quantization to decode a moving image. It includes circuitry and a memory. The circuitry uses the memory to perform inverse quantization only on a specified range of multiple quantization coefficients in the low-frequency domain of the multiple quantization coefficients contained in the processing object block using a quantization matrix.

[0407] Therefore, by using only the quantization matrix corresponding to a specified range in the low-frequency domain that has a significant impact on visual perception within the processing object block, the image quality of the moving image is less likely to degrade. Furthermore, since only the quantization matrix corresponding to this specified range is encoded, the amount of coding can be reduced. Thus, according to the decoding device 200, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0408] For example, in the decoding device 200, the circuit may decode the signal associated with the quantization matrix corresponding only to a plurality of transform coefficients within the predetermined range of the low-frequency domain side from the bitstream.

[0409] This reduces the amount of code. Therefore, according to the decoding device 200, processing efficiency is improved.

[0410] For example, in the decoding device 200, the processing target block may be a square block or a rectangular block. The circuit transforms a first quantization matrix for a plurality of transform coefficients within a specified range in the square block, generates a second quantization matrix for a plurality of transform coefficients within the specified range in the rectangular block based on the first quantization matrix, and decodes the signal associated with only the first quantization matrix among the first quantization matrix and the second quantization matrix from the bitstream.

[0411] Therefore, the decoding device 200 generates a first quantization matrix corresponding to a defined range in the low-frequency domain of the square block for processing object blocks of multiple shapes including rectangular blocks, and decodes only the first quantization matrix, thus reducing the amount of coding. Furthermore, the decoding device 200 generates a second quantization matrix corresponding to a defined range in the low-frequency domain of the rectangular block based on the first quantization matrix, so the image quality of the moving image is less likely to degrade. Therefore, according to the decoding device 200, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0412] For example, in the decoding device 200, the circuit may perform downconversion processing based on a first quantization matrix for a plurality of transform coefficients within a specified range in the square block to generate a second quantization matrix for a plurality of transform coefficients within a specified range in the rectangular block, wherein the square block has one side of the same length as the long side of the rectangular block that is the processing target block.

[0413] Thus, the decoding device 200 is able to efficiently inverse quantize the rectangular block.

[0414] For example, in the decoding device 200, the circuit may extend the plurality of matrix elements of the first quantization matrix by extrapolating in a predetermined direction during the downconversion process, and divide the plurality of matrix elements of the extended first quantization matrix into a number of groups equal to the number of matrix elements of the second quantization matrix. For each of the plurality of groups, the matrix element located on the lowest domain side of the plurality of matrix elements contained in the group, the matrix element located on the highest domain side of the plurality of matrix elements contained in the group, or the average value of the plurality of matrix elements contained in the group is determined as the matrix element in the second quantization matrix corresponding to that group.

[0415] Therefore, the decoding device 200 is able to perform inverse quantization on rectangular blocks more efficiently.

[0416] For example, in the decoding device 200, the circuit may perform upconversion processing based on a first quantization matrix of a plurality of transform coefficients within a predetermined range in the square block to generate a second quantization matrix for the plurality of transform coefficients within the predetermined range in the rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0417] Thus, the decoding device 200 is able to efficiently inverse quantize the rectangular block.

[0418] For example, in the decoding device 200, the circuit may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, extend the first quantization matrix in the predetermined direction by repeating the plurality of matrix elements contained in the group in the predetermined direction, and extract the same number of matrix elements as the plurality of matrix elements of the second quantization matrix from the extended first quantization matrix, or (ii) extend the first quantization matrix in the predetermined direction by performing linear interpolation between matrix elements adjacent in the predetermined direction among the plurality of matrix elements of the second quantization matrix, and extract the same number of matrix elements as the plurality of matrix elements of the second quantization matrix from the extended first quantization matrix.

[0419] Therefore, the decoding device 200 is able to perform inverse quantization on rectangular blocks more efficiently.

[0420] For example, in the decoding device 200, the circuit may generate the second quantization matrix corresponding to a plurality of transform coefficients within a specified range in the rectangular block by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block, which is the processing target block: a method of generating the matrix by performing the down-conversion process based on the first quantization matrix of a square block having a side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing the up-conversion process based on the first quantization matrix of a square block having a side with the same length as the short side of the rectangular block.

[0421] Therefore, by switching between downconversion and upconversion according to the block size of the block being processed, the decoding device 200 can generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0422] In addition, the encoding method is a method of encoding motion images by quantization, which uses a quantization matrix to quantize only multiple transform coefficients within a specified range of the low-frequency domain side of the multiple transform coefficients contained in the processing object block.

[0423] Therefore, by using only the quantization matrix corresponding to a specified range of the low-frequency domain side of the processing block that has a significant impact on visual perception, the image quality of moving images is less likely to degrade. Furthermore, since only the quantization matrix corresponding to this specified range is encoded, the amount of coding is reduced. Thus, according to this encoding method, the image quality of moving images is less likely to degrade, and processing efficiency is improved.

[0424] In addition, the decoding method is a decoding method that performs inverse quantization to decode the motion image, and only performs inverse quantization on multiple quantization coefficients within a specified range of the low-frequency domain side of the multiple quantization coefficients contained in the processing object block using a quantization matrix.

[0425] Therefore, by using only the quantization matrix corresponding to a specified range in the low-frequency domain that has a significant impact on visual perception within the processing block for inverse quantization, the image quality of moving images is less likely to degrade. Furthermore, since decoding is performed only on the quantization matrix corresponding to this specified range, the amount of coding can be reduced. Thus, according to this decoding method, the image quality of moving images is less likely to degrade, and processing efficiency is improved.

[0426] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.

[0427] [Method 3]

[0428] The encoding device 100, decoding device 200, encoding method, and decoding method of the third aspect of the present invention will be described below.

[0429] [Example 3 of encoding and decoding using quantization matrices]

[0430] Figure 26 This is a diagram illustrating the third example of an encoding process using the quantization matrix (QM) in the encoding device 100. Furthermore, in the encoding device 100 described here, each block of a square or rectangle that divides the screen is encoded.

[0431] First, in step S1601, the quantization unit 108 generates a QM (hereinafter referred to as the QM of diagonal components only) corresponding to the diagonal components of the processing target block. For each block size of the square block or rectangular block of various shapes that serves as the processing target block, the QM corresponding to the processing target block is generated using a common method described below, based on the values ​​of the quantization coefficients of the QM of diagonal components only. In other words, the quantization unit 108 generates a quantization matrix for the processing target block based on the diagonal components of the quantization matrix of the multiple transformation coefficients contained in the processing target block, specifically for multiple transformation coefficients continuously arranged diagonally in the processing target block. Furthermore, using a common method means using a common method for all processing target blocks regardless of the shape and size of the block. Additionally, diagonal components refer to, for example, multiple coefficients along the diagonal from the low-domain side to the high-domain side of the processing target block.

[0432] The entropy encoding unit 110 records the signal related to the diagonal component of the QM generated in step S1601 into the stream. In other words, the entropy encoding unit 110 encodes the signal related to the diagonal component of the quantization matrix into the bit stream.

[0433] Furthermore, the quantization unit 108 can generate the quantization coefficient values ​​of the diagonal component QM based on user-defined values ​​set in the encoding device 100, or it can adaptively generate them using the encoding information of already encoded images. Additionally, the diagonal component QM can be encoded into the sequence header region, image header region, slice header region, auxiliary information region, or other parameter region of the stored stream. Furthermore, the diagonal component QM may not be recorded in the stream. In this case, the quantization unit 108 can also use a default value predefined in the standard as the diagonal component QM.

[0434] Furthermore, the processing in step S1601 can be a structure that is performed centrally at the start of sequence processing, or at the start of image processing, or at the start of slice processing, or it can be a structure that performs partial processing each time in block unit processing. In addition, the QM generated in step S1601 can also be configured to generate multiple QMs for blocks of the same block size based on luminance block use / chrominance block use, intra-frame prediction block use / inter-frame prediction block use, and other conditions.

[0435] In addition, Figure 26 In the processing flow shown, the processing steps other than step S1601 are block-level loop processing, which is consistent with the use of... Figure 11 The same applies to the first example described.

[0436] Therefore, instead of recording all the quantization coefficients of the QM for each block size of the processing target block in the stream, only the quantization coefficients of the diagonal component of the QM are recorded in the stream, thus enabling the encoding processing of the processing target block. Therefore, even when using an encoding method that includes a large number of blocks of various shapes containing rectangular blocks, it is possible to generate and use the QM corresponding to the processing target block without significantly increasing the amount of encoding in the header region, thereby increasing the potential for improved encoding efficiency.

[0437] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0438] Figure 27 This indicates that the use of and Figure 26 This diagram illustrates an example of the decoding process of the quantization matrix (QM) in the decoding device 200 corresponding to the encoding device 100 described herein. Furthermore, in the decoding device 200 described here, decoding processing is performed on each square or rectangular block that divides the image.

[0439] First, in step S1701, the entropy decoding unit 202 decodes the signal related to the QM of the diagonal component only from the stream. Using the decoded signal related to the QM of the diagonal component only, and employing the common method described below, it generates QMs corresponding to the block sizes of various shaped processing object blocks, such as square blocks and rectangular blocks. Alternatively, the QM of the diagonal component only can also be decoded from the sequence header region, image header region, slice header region, auxiliary information region, or other parameter region of the stored stream. Alternatively, the QM of the diagonal component only may not be decoded from the stream. In this case, for example, a default value predefined in the standard can be used as the QM of the diagonal component only.

[0440] Furthermore, the processing in step S1701 can be performed in a concentrated manner at the start of sequence processing, image processing, or slice processing, or it can be performed in a partial manner each time in block unit processing. Additionally, the QM generated by the entropy decoding unit 202 in step S1701 can also be configured to generate multiple QMs for blocks of the same block size based on conditions such as luminance block use / chrominance block use, intra-frame prediction block use / inter-frame prediction block use, and other factors.

[0441] In addition, Figure 27 In the processing flow shown, the processing flow other than step S1701 is a block-based loop processing, which is similar to using... Figure 12 The processing procedure for the first example described is the same.

[0442] Therefore, even if not all the quantization coefficients of the QM for each block size of the object block are recorded in the stream, as long as the quantization coefficients of the QM for only the diagonal components of the object block are recorded in the stream, decoding of the object block can be performed. Thus, the amount of decoding from the beginning region can be reduced, thereby increasing the potential for improved processing efficiency.

[0443] Furthermore, this process is just one example; the order of the recorded processes can be changed, a portion of the recorded processes can be removed, or unrecorded processes can be added.

[0444] Figure 28 It is used to explain in Figure 26 Step S1601 and Figure 27 In step S1701, a diagram illustrating an example of a method for generating the QM of the target block is generated using a common method described below, based on the quantization coefficients of the QM for only the diagonal components in each block size. Furthermore, the processing described here is a common process in both the encoding device 100 and the decoding device 200.

[0445] The encoding device 100 and decoding device 200 of the third method generate the quantization matrix (QM) of the processing object block by repeating multiple matrix elements of the diagonal components of the processing object block in the horizontal and vertical directions, respectively. More specifically, the encoding device 100 and decoding device 200 directly stretch the values ​​of the quantization coefficients of the QM of the diagonal components in the upward and leftward directions, that is, they generate the QM of the processing object block by continuously configuring the same values.

[0446] Additionally, this example illustrates the method for generating the QM (Quality Manager) for the object block when it is a square block. However, the method for generating the QM for the object block when it is a rectangular block can also be described using the same approach. Figure 28 Similarly, the QM of the processed object block is generated based on the quantization coefficients of the QM of the diagonal components.

[0447] Figure 29 It is used to explain in Figure 26 Step S1601 and Figure 27 In step S1701, a diagram illustrating another example of the method for generating the QM of the processed object block using the common method described below, based on the values ​​of the quantization coefficients of the QM for only the diagonal components in each block size of the processed object block. Furthermore, the processing described here is a common process in both the encoding device 100 and the decoding device 200.

[0448] The encoding device 100 and decoding device 200 in the third method can also generate the quantization matrix of the processing object block by repeating multiple matrix elements of the diagonal components in the diagonal direction. More specifically, the encoding device 100 and decoding device 200 directly stretch the values ​​of the quantization coefficients of the QM of the diagonal components in the lower left and upper right directions, that is, by continuously configuring the same values ​​to generate the QM of the processing object block.

[0449] At this point, in addition to the quantization coefficients of the diagonal components, the encoding device 100 and the decoding device 200 can also use the quantization coefficients of the nearby components of the diagonal components to repeat each quantization coefficient in the tilt direction, thereby generating the QM of the processing target block. In other words, the quantization matrix (QM) of the processing target block can be generated based on multiple matrix elements of the diagonal components and matrix elements located near the diagonal components. Thus, even if it is difficult to fill all the quantization coefficients of the processing target block using only the diagonal components, it is possible to fill all the quantization coefficients by using the quantization coefficients of the nearby components. Furthermore, the nearby components of the diagonal components are, for example, components adjacent to any one of the multiple coefficients on the diagonal from the low-domain side to the high-frequency side of the processing target block.

[0450] For example, the quantization coefficient of the components near the diagonal component is Figure 29 The quantization coefficients at the indicated positions. The encoding device 100 and the decoding device 200 can be configured to encode and decode signals related to the quantization coefficients of nearby components of the diagonal component into and from the stream, or they can be configured not to encode and decode the signals into and from the stream, but to derive the signals by interpolating the values ​​of adjacent quantization coefficients in the quantization coefficients of the diagonal component using linear interpolation or the like.

[0451] Additionally, this example illustrates the method for generating the QM (Quality Manager) for the object block when it is a square block. However, the method for generating the QM for the object block when it is a rectangular block can also be described using the same approach. Figure 29 Similarly, the QM of the processed object block is generated based on the quantization coefficients of the QM of the diagonal components.

[0452] [Other variations of the third example of encoding and decoding using quantization matrices]

[0453] exist Figure 28 and Figure 29 The example illustrates a method for generating the quantization coefficients of the QM for the entire processing object block based on the quantization coefficients of the diagonal components of the QM. However, it can also be structured to generate coefficients for only a portion of the QM for the processing object block based on the quantization coefficients of the diagonal components of the QM. For example, for the QM corresponding to the region on the lower domain side of the processing object block, all quantization coefficients contained in that QM can be encoded into and decoded from the stream, while only the quantization coefficients of the QM corresponding to the regions on the middle and upper domains of the processing object block can be generated based on the quantization coefficients of the diagonal components of the QM.

[0454] Alternatively, the method for generating a rectangular block's QM from a square block's QM can also be achieved by using... Figure 26 and Figure 27 The third example and its use Figure 11 and Figure 12 The structure of the first example combination is explained. For example, the QM for the square block can be generated using either of the two common methods described above, based on the quantization coefficients of the QM for the diagonal components of the square block, as explained in the third example. The QM for the rectangular block can be generated using the generated QM for the square block, as explained in the first example.

[0455] Additionally, the method for generating a rectangular block from a square block's QM can also be set to use... Figure 26 and Figure 27 The third example and its use Figure 17 and Figure 18 The structure of the second example combination is illustrated. For example, the QM corresponding to the size of the effective transform coefficient region in each block size of the processed object block can also be generated using any of the common methods described above, based on the quantization coefficient values ​​of the QM of the diagonal components of the effective transform coefficient region.

[0456] [The effect of the third example of encoding and decoding using a quantization matrix]

[0457] According to the third aspect of the present invention, the encoding device 100 and the decoding device 200, by using Figure 26 and Figure 27 The described structure allows for encoding and decoding of the target block even if not all the QM quantization coefficients for each block size are recorded in the stream, as long as the QM quantization coefficients for the diagonal components of the target block are recorded in the stream. Therefore, it is possible to reduce the amount of code in the header region, thus increasing the potential for improved encoding efficiency.

[0458] For example, encoding device 100 is an encoding device that performs quantization to encode a moving image, and includes circuitry and a memory. The circuitry uses the memory to generate a quantization matrix for the processing object block based on the diagonal component of a quantization matrix for the processing object block, which is composed of multiple transform coefficients arranged consecutively in the diagonal direction of the processing object block, and encodes the signal associated with the diagonal component into a bitstream.

[0459] Therefore, even in encoding methods that use blocks of numerous shapes containing rectangular blocks, only the diagonal components of the block being processed are encoded, thus reducing the amount of encoding. Consequently, processing efficiency is improved according to the encoding device 100.

[0460] For example, in the encoding device 100, the circuit may generate the quantization matrix by repeating multiple matrix elements of the diagonal component in both the horizontal and vertical directions.

[0461] Therefore, the quantization coefficients contained in the processing object block are generated based on the quantization coefficients of the diagonal components of the processing object block, so the entire quantization matrix of the processing object block may not need to be encoded. Thus, processing efficiency is improved due to the reduction in coding complexity. Therefore, according to the encoding device 100, the processing object block can be quantized efficiently.

[0462] For example, in the encoding device 100, the circuit may generate the quantization matrix by repeating multiple matrix elements of the diagonal component in the tilt direction.

[0463] Therefore, the quantization coefficients contained in the processing object block are generated based on the quantization coefficients of the diagonal components of the processing object block, so the entire quantization matrix of the processing object block may not need to be encoded. Thus, processing efficiency is improved due to the reduction in coding complexity. Therefore, according to the encoding device 100, the processing object block can be quantized efficiently.

[0464] For example, in the encoding device 100, the circuit may generate the quantization matrix based on a plurality of matrix elements located in the diagonal component and a plurality of matrix elements in the vicinity of the diagonal component.

[0465] Therefore, the encoding device 100 generates the quantization coefficients contained in the processing object block based on the quantization coefficients of the diagonal components and the nearby components of the processing object block, and thus can generate a more appropriate quantization matrix corresponding to the processing object block.

[0466] For example, in the encoding device 100, the circuit may not encode the signals associated with the multiple matrix elements located near the diagonal component into the bit stream, but instead generate the multiple matrix elements located near the diagonal component by interpolating based on the multiple adjacent matrix elements among the multiple matrix elements of the diagonal component.

[0467] This reduces the amount of coding, thus improving processing efficiency. Therefore, the encoding device 100 is able to efficiently quantize the block of data to be processed.

[0468] For example, in the encoding apparatus 100, the circuit may use a quantization matrix to quantize multiple transform coefficients within a specified range on the low-frequency domain side of the multiple transform coefficients contained in the processing object block, encode the signal associated with the quantization matrix corresponding only to the multiple transform coefficients within the specified range on the low-frequency domain side into a bit stream, and use multiple matrix elements of the diagonal component to quantize multiple transform coefficients outside the specified range on the low-frequency domain side of the multiple transform coefficients contained in the processing object block.

[0469] Therefore, in the processing object block, all quantization coefficients within a specified range in the low-frequency domain that have a significant impact on visual perception are encoded, thus minimizing the degradation of motion image quality. Furthermore, in the processing object block, for multiple transform coefficients outside the specified range in the low-frequency domain, quantization is performed using the quantization matrix of the diagonal component of the processing object block, thereby reducing the amount of coding and improving processing efficiency. Therefore, according to the encoding device 100, the motion image quality is less likely to degrade, and processing efficiency is improved.

[0470] For example, in the encoding apparatus 100, the quantization matrix may be a quantization matrix that corresponds only to multiple transform coefficients within the predetermined range on the low-frequency domain side of the multiple transform coefficients contained in the processing object block.

[0471] Therefore, according to the encoding device 100, a quantization matrix corresponding to a specified range of the low-frequency domain side that has a large impact on vision is generated in the processing object block, so the image quality of the moving image is not easily degraded.

[0472] For example, in the encoding device 100, the processing target block may be a square block or a rectangular block. The circuit transforms a first quantization matrix for a plurality of transform coefficients within a specified range in the square block, generates a second quantization matrix for a plurality of transform coefficients within the specified range in the rectangular block based on the first quantization matrix, and encodes only the first quantization matrix from the first quantization matrix and the second quantization matrix as a signal associated with the quantization matrix into a bit stream.

[0473] Therefore, only the first quantization matrix corresponding to the specified range within the square block is encoded, thus reducing the amount of encoding. Furthermore, since a second quantization matrix corresponding to the specified range within the rectangular block is generated based on the first quantization matrix, the image quality of the moving image is less likely to degrade. Additionally, the specified range is located in the low-frequency domain, which has a significant impact on visual perception. Therefore, according to the encoding device 100, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0474] For example, in the encoding device 100, the circuit may perform downconversion processing on the first quantization matrix for the first quantization matrix for the first quantization matrix within the specified range of the square block to generate the second quantization matrix for the first quantization matrix within the specified range of the square block, the square block having one side of the same length as the long side of the rectangular block which is the processing target block.

[0475] Thus, the encoding device 100 is able to efficiently generate a second quantization matrix corresponding to a specified range in the rectangular block.

[0476] For example, in the encoding device 100, the circuit may perform upconversion processing based on a first quantization matrix for a plurality of transform coefficients within a specified range in the square block to generate a second quantization matrix for a plurality of transform coefficients within a specified range in the rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0477] Thus, the encoding device 100 is able to efficiently generate a second quantization matrix corresponding to a specified range in the rectangular block.

[0478] For example, in the encoding device 100, the circuit may generate the second quantization matrix for a plurality of transform coefficients within a predetermined range in the rectangular block by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block, which is the processing target block: a method of generating the matrix by performing the downconversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing the upconversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the short side of the rectangular block.

[0479] Thus, by switching down-conversion and up-conversion, the encoding device 100 can generate a more appropriate second quantization matrix corresponding to the specified range of the rectangular block based on the first quantization matrix corresponding to the specified range of the square block.

[0480] Additionally, the decoding device 200 is a decoding device that performs inverse quantization to decode a moving image. It includes circuitry and a memory. The circuitry uses the memory to generate a quantization matrix for the processing object block based on the diagonal component of a quantization matrix for the processing object block, which is composed of a plurality of quantization coefficients arranged consecutively in the diagonal direction of the processing object block. It then decodes the signal associated with the diagonal component from the bitstream.

[0481] Therefore, even in decoding methods that use blocks of numerous shapes containing rectangular blocks, only the diagonal components of the target block are decoded, thus reducing the amount of encoding. Consequently, processing efficiency is improved according to the decoding apparatus 200.

[0482] For example, in the decoding device 200, the circuit can also generate the quantization matrix by repeating multiple matrix elements of the diagonal component in both the horizontal and vertical directions.

[0483] Therefore, by generating the quantization coefficients of the diagonal components of the processing object block based on these coefficients, it is possible to avoid decoding the entire quantization matrix of the processing object block. This reduces the amount of coding required, thus improving processing efficiency. Therefore, the decoding apparatus 200 can efficiently perform inverse quantization on the processing object block.

[0484] For example, in the decoding device 200, the circuit may generate the quantization matrix by repeating multiple matrix elements of the diagonal component in the tilt direction.

[0485] Therefore, by generating the quantization coefficients of the diagonal components of the processing object block based on these coefficients, it is possible to avoid decoding the entire quantization matrix of the processing object block. This reduces the amount of coding required, thus improving processing efficiency. Therefore, the decoding apparatus 200 can efficiently perform inverse quantization on the processing object block.

[0486] For example, in the decoding device 200, the circuit may generate the quantization matrix from a plurality of matrix elements located in the diagonal component and a plurality of matrix elements in the vicinity of the diagonal component.

[0487] Therefore, the decoding device 200 generates the quantization coefficients contained in the processing object block based on the quantization coefficients of the diagonal components and the nearby components of the processing object block, and thus can generate a more appropriate quantization matrix corresponding to the processing object block.

[0488] For example, in the decoding device 200, the circuit may not decode the signal associated with the multiple matrix elements located near the diagonal component from the bit stream, but instead interpolate from the multiple matrix elements of the multiple matrix elements of the diagonal component to generate the multiple matrix elements located near the diagonal component.

[0489] This reduces the amount of coding, thus improving processing efficiency. Therefore, the decoding device 200 can efficiently perform inverse quantization on the block being processed.

[0490] For example, in the decoding device 200, the circuit may use a quantization matrix to inverse quantize a plurality of quantization coefficients within a specified range in the low-frequency domain of a plurality of quantization coefficients contained in the processing object block, decode from the bitstream a signal that is only related to the quantization matrix corresponding to a plurality of transform coefficients within the specified range in the low-frequency domain, and use a plurality of matrix elements of the diagonal component to inverse quantize a plurality of quantization coefficients outside the specified range in the low-frequency domain of a plurality of quantization coefficients contained in the processing object block.

[0491] Therefore, in the processing object block, all quantization coefficients within the specified range of the low-frequency domain side, which have a significant impact on visual perception, are decoded, thus minimizing the degradation of motion image quality. Furthermore, in the processing object block, for multiple quantization coefficients outside the specified range of the low-frequency domain side, inverse quantization is performed using the quantization matrix of the diagonal component of the processing object block, thereby reducing the amount of coding and improving processing efficiency. Therefore, according to the decoding apparatus 200, the image quality of motion images is less likely to degrade, and processing efficiency is improved.

[0492] For example, in the decoding device 200, the quantization matrix may be a quantization matrix that corresponds to only the multiple transform coefficients in the predetermined range on the low-frequency domain side of the multiple transform coefficients contained in the processing object block.

[0493] Therefore, according to the decoding device 200, a quantization matrix corresponding to a specified range of the low-frequency domain side that has a large impact on vision is generated in the processing object block, so the image quality of the moving image is not easily degraded.

[0494] For example, in the decoding device 200, the processing target block may be a square block or a rectangular block. The circuit transforms a first quantization matrix of multiple transform coefficients within the specified range in the square block, generates a second quantization matrix of multiple transform coefficients within the specified range in the rectangular block based on the first quantization matrix, and decodes only the first quantization matrix from the bitstream as a signal associated with the quantization matrix.

[0495] Therefore, decoding is performed only on the first quantization matrix corresponding to the specified range within the square block, thus reducing the amount of coding. Furthermore, since a second quantization matrix corresponding to the specified range within the rectangular block is generated based on the first quantization matrix, the image quality of the moving image is less likely to degrade. Additionally, the specified range is located in the low-frequency domain, which has a significant impact on visual perception. Therefore, according to the decoding device 200, the image quality of the moving image is less likely to degrade, and processing efficiency is improved.

[0496] For example, in the decoding device 200, the circuit may perform downconversion processing based on the first quantization matrix for a plurality of transform coefficients within the specified range in the square block to generate the second quantization matrix for a plurality of transform coefficients within the specified range in the rectangular block, the square block having one side of the same length as the long side of the rectangular block which is the processing target block.

[0497] Thus, the decoding device 200 is able to efficiently generate a second quantization matrix corresponding to a specified range in the rectangular block.

[0498] For example, in the decoding device 200, the circuit may perform upconversion processing based on a first quantization matrix for a plurality of transform coefficients within a specified range in the square block to generate a second quantization matrix for a plurality of transform coefficients within a specified range in the rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0499] Thus, the decoding device 200 is able to efficiently generate a second quantization matrix corresponding to a specified range in the rectangular block.

[0500] For example, in the decoding device 200, the circuit may generate the second quantization matrix for a plurality of transform coefficients within a predetermined range in the rectangular block by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block, which is the processing target block: a method of generating the matrix by performing the down-conversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing the up-conversion process based on the first quantization matrix for a plurality of transform coefficients within a predetermined range in the square block having a side with the same length as the short side of the rectangular block.

[0501] Thus, by switching down-conversion and up-conversion, the decoding device 200 can generate a more appropriate second quantization matrix corresponding to the specified range of the rectangular block based on the first quantization matrix corresponding to the specified range of the square block.

[0502] In addition, the encoding method is an encoding method for encoding motion images by performing quantization. Based on the diagonal components of the quantization matrix of the multiple transform coefficients contained in the processing object block, which are continuously arranged diagonally in the processing object block, the quantization matrix for the processing object block is generated, and the signal associated with the diagonal components is encoded into the bit stream.

[0503] Therefore, even in encoding methods that use blocks of numerous shapes containing rectangular blocks, only the diagonal components of the target block are encoded, thus reducing the amount of encoding. Consequently, processing efficiency is improved depending on the encoding method.

[0504] In addition, the decoding method is a decoding method that performs inverse quantization to decode the motion image. Based on the diagonal component of the quantization matrix of the multiple quantization coefficients contained in the processing object block, which is arranged continuously in the diagonal direction of the processing block, the quantization matrix for the processing object block is generated, and the signal related to the diagonal component is decoded from the bit stream.

[0505] Therefore, even in decoding methods that use blocks of numerous shapes containing rectangular blocks, only the diagonal components of the target block are decoded, thus reducing the amount of information. Consequently, processing efficiency is improved depending on the decoding method.

[0506] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.

[0507] [Installation Example]

[0508] Figure 30 This is a block diagram showing an example of the installation of the encoding device 100. The encoding device 100 includes circuitry 160 and a memory 162. For example, Figure 1 The multiple components of the encoding device 100 shown are composed of Figure 30 The circuit 160 and memory 162 shown are installed.

[0509] Circuit 160 is an electronic circuit capable of accessing memory 162 and performing information processing. For example, circuit 160 is a dedicated or general-purpose electronic circuit that uses memory 162 to encode moving images. Circuit 160 can also be a processor like a CPU. Alternatively, circuit 160 can be an assembly of multiple electronic circuits.

[0510] Additionally, circuit 160 can also serve, for example. Figure 1 The encoding device 100 shown above includes multiple components other than those used for storing information. That is, the circuit 160 can also perform the aforementioned operations as an operation of these components.

[0511] Memory 162 is a dedicated or general-purpose memory used by storage circuit 160 to encode moving images. Memory 162 can be an electronic circuit, connected to circuit 160, or included in circuit 160.

[0512] Furthermore, the memory 162 can be an assembly of multiple electronic circuits, or it can be composed of multiple sub-memories. Additionally, the memory 162 can be a disk or optical disc, or it can be a storage medium or recording medium. Furthermore, the memory 162 can be either non-volatile or volatile memory.

[0513] For example, memory 162 can also serve as Figure 1 The encoding device 100 shown here has multiple components, including the component for storing information. Specifically, the memory 162 can also serve as... Figure 1 The functions of the block memory 118 and frame memory 122 shown are illustrated.

[0514] Furthermore, the memory 162 can store the encoded motion image, or it can store the bit string corresponding to the encoded motion image. Additionally, the memory 162 can also store the program used by the circuit 160 to encode the motion image.

[0515] Alternatively, the encoding device 100 may not require installation. Figure 1 All of the multiple constituent elements shown can also be processed without the aforementioned multiple processing steps. Figure 1 Some of the multiple components shown may be included in other devices, or some of the aforementioned processes may be performed by other devices. Furthermore, in the encoding device 100, an assembly is installed... Figure 1 One of the constituent elements shown can be appropriately derived into a prediction sample set by performing one of the aforementioned processes.

[0516] Figure 31 It means Figure 30 A flowchart illustrating an example of the operation of the encoding device 100. For example, Figure 30 The encoding device 100 shown performs the following when encoding moving images: Figure 31 The actions shown are as follows. Specifically, circuit 160 uses memory 162 to perform the following actions.

[0517] First, circuit 160 transforms the first quantization matrix for multiple transform coefficients of the square block, and generates a second quantization matrix for multiple transform coefficients of the rectangular block based on the first quantization matrix (step S201). Next, circuit 160 uses the second quantization matrix to quantize the multiple transform coefficients of the rectangular block (step S202).

[0518] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, it is also possible to avoid encoding the quantization matrix corresponding to the rectangular block. Thus, processing efficiency is improved due to the reduction in encoding. Therefore, the encoding device 100 can efficiently quantize rectangular blocks.

[0519] For example, circuit 160 could encode only the first quantization matrix from the first quantization matrix and the second quantization matrix into the bitstream.

[0520] This reduces the amount of code. Therefore, according to the encoding device 100, processing efficiency is improved.

[0521] For example, circuit 160 may perform downconversion processing based on a first quantization matrix of multiple transform coefficients for a square block to generate a second quantization matrix of multiple transform coefficients for a rectangular block, the square block having one side of the same length as the long side of the rectangular block which is the block to be processed.

[0522] Thus, the encoding device 100 can efficiently generate a quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to the square block, wherein the square block has one side of the same length as the long side of the rectangular block.

[0523] For example, in the downconversion process, circuit 160 may divide the multiple matrix elements of the first quantization matrix into the same number of groups as the multiple matrix elements of the second quantization matrix. For each of the multiple groups, the multiple matrix elements contained in the group are arranged continuously in the horizontal or vertical direction of the square block. For each of the multiple groups, the matrix element located on the lowest domain side of the multiple matrix elements contained in the group, the matrix element located on the highest domain side of the multiple matrix elements contained in the group, or the average value of the multiple matrix elements contained in the group is determined as the matrix element corresponding to the group in the second quantization matrix.

[0524] Therefore, the encoding device 100 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0525] For example, circuit 160 may perform upconversion processing based on a first quantization matrix of multiple transform coefficients for a square block to generate a second quantization matrix of multiple transform coefficients for a rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the block to be processed.

[0526] Thus, the encoding device 100 can efficiently generate a quantization matrix corresponding to a rectangular block from the quantization matrix corresponding to the square block, the square block having one side of the same length as the short side of the rectangular block.

[0527] For example, in the upconversion process, circuit 160 may (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix element in the second quantization matrix corresponding to that group by repeating the plurality of matrix elements contained in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements in the plurality of matrix elements of the second quantization matrix.

[0528] Therefore, the encoding device 100 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0529] For example, circuit 160 may generate a second quantization matrix of multiple transform coefficients for a rectangular block by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block being processed: a method of generating the matrix by performing down-conversion processing on a first quantization matrix of multiple transform coefficients for a square block having one side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing up-conversion processing on a first quantization matrix of multiple transform coefficients for a square block having one side with the same length as the short side of the rectangular block.

[0530] Therefore, by switching downconversion and upconversion according to the block size of the object block being processed, the encoding device 100 is able to generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0531] Figure 32 This is a block diagram showing an example of the installation of the decoding device 200. The decoding device 200 includes circuitry 260 and a memory 262. For example, Figure 10 The multiple components of the decoding device 200 shown are transmitted through Figure 32 The circuit 260 and memory 262 shown are installed.

[0532] Circuit 260 is an electronic circuit capable of accessing memory 262 and performing information processing. For example, circuit 260 is a dedicated or general-purpose electronic circuit that uses memory 262 to decode moving images. Circuit 260 can also be a processor like a CPU. Alternatively, circuit 260 can be an assembly of multiple electronic circuits.

[0533] Additionally, for example, circuit 260 can also serve as Figure 10 The decoding device 200 shown includes multiple components, excluding those used for storing information. That is, the circuit 260 can also perform the aforementioned operations as an operation of these components.

[0534] The memory 262 is a dedicated or general-purpose memory used by the storage circuit 260 to decode moving images. The memory 262 can be an electronic circuit, connected to the circuit 260, or contained within the circuit 260.

[0535] Furthermore, the memory 262 can be an assembly of multiple electronic circuits, or it can be composed of multiple sub-memories. Additionally, the memory 262 can be a disk or optical disk, or it can be a memory or recording medium. Furthermore, the memory 262 can be either non-volatile or volatile memory.

[0536] For example, memory 262 can also serve as Figure 10 The decoding device 200 shown here has multiple components, including the component for storing information. Specifically, the memory 262 can also serve as... Figure 10 The functions of the block memory 210 and frame memory 214 shown are illustrated.

[0537] Additionally, memory 262 can store the bit string corresponding to the encoded motion image, or it can store the decoded motion image. Furthermore, memory 262 can also store the program used by circuit 260 to decode the motion image.

[0538] Alternatively, it is not necessary to install in the decoding device 200. Figure 10 All of the multiple constituent elements shown can also be processed without the aforementioned multiple processing steps. Figure 10 Some of the multiple components shown may be included in other devices, or some of the aforementioned processes may be performed by other devices. Furthermore, in the decoding device 200, a portion of the components is installed... Figure 10 One of the constituent elements shown can be appropriately derived into a prediction sample set by performing one of the aforementioned processes.

[0539] Figure 33 It means Figure 32 A flowchart illustrating an example of the operation of the decoding device 200 is shown. For example, Figure 32 The decoding device 200 shown performs the following when decoding moving images: Figure 33 The actions shown are as follows. Specifically, circuit 260 uses memory 262 to perform the following actions.

[0540] First, circuit 260 transforms the first quantization matrix for multiple transform coefficients of the square block, and generates a second quantization matrix for multiple transform coefficients of the rectangular block based on the first quantization matrix (step S301). Next, circuit 260 uses the second quantization matrix to inverse quantize the multiple transform coefficients of the rectangular block (step S302).

[0541] Therefore, since the quantization matrix corresponding to the rectangular block is generated based on the quantization matrix corresponding to the square block, decoding of the quantization matrix corresponding to the rectangular block is not required. This reduces the amount of coding, thus improving processing efficiency. Therefore, according to the decoding device 200, the rectangular block can be efficiently dequantized.

[0542] For example, circuit 260 could decode only the first quantization matrix from the first quantization matrix and the second quantization matrix in the bit stream.

[0543] This reduces the amount of code. Therefore, according to the decoding device 200, processing efficiency is improved.

[0544] For example, circuit 260 may perform downconversion processing based on a first quantization matrix of multiple transform coefficients for a square block to generate a second quantization matrix of multiple transform coefficients for a rectangular block, the square block having one side of the same length as the long side of the rectangular block which is the block to be processed.

[0545] Thus, the decoding device 200 can efficiently generate a quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to the square block, wherein the square block has one side of the same length as the long side of the rectangular block.

[0546] For example, in the downconversion process, circuit 260 may divide the multiple matrix elements of the first quantization matrix into the same number of groups as the multiple matrix elements of the second quantization matrix. For each of the multiple groups, the matrix element located on the lowest domain side of the multiple matrix elements contained in the group, the transformation coefficient located on the highest domain side of the multiple matrix elements contained in the group, or the average value of the multiple matrix elements contained in the group is determined as the matrix element in the second quantization matrix corresponding to that group.

[0547] Therefore, the decoding device 200 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0548] For example, circuit 260 may perform upconversion processing based on a first quantization matrix for multiple transform coefficients of a square block to generate a second quantization matrix for multiple transform coefficients of a rectangular block, the square block having one side of the same length as the short side of the rectangular block which is the processing target block.

[0549] Thus, the decoding device 200 can efficiently generate a quantization matrix corresponding to a rectangular block from the quantization matrix corresponding to the square block, the square block having one side of the same length as the short side of the rectangular block.

[0550] For example, circuit 260 may, in the upconversion process, (i) divide the plurality of matrix elements of the second quantization matrix into the same number of groups as the plurality of matrix elements of the first quantization matrix, and for each of the plurality of groups, determine the matrix element in the second quantization matrix corresponding to that group by repeating the plurality of matrix elements contained in that group, or (ii) determine the plurality of matrix elements of the second quantization matrix by performing linear interpolation between adjacent matrix elements in the plurality of matrix elements of the second quantization matrix.

[0551] Therefore, the decoding device 200 can generate the quantization matrix corresponding to the rectangular block more efficiently based on the quantization matrix corresponding to the square block.

[0552] For example, circuit 260 may generate a second quantization matrix of transform coefficients for a rectangle by switching between the following methods based on the ratio of the length of the short side to the length of the long side of the rectangular block being processed: a method of generating the matrix by performing a down-conversion process based on a first quantization matrix of multiple transform coefficients for a square block having one side with the same length as the long side of the rectangular block; and a method of generating the matrix by performing an up-conversion process based on a first quantization matrix of multiple transform coefficients for a square block having one side with the same length as the short side of the rectangular block.

[0553] Therefore, by switching between downconversion and upconversion according to the block size of the block being processed, the decoding device 200 can generate a more appropriate quantization matrix corresponding to a rectangular block based on the quantization matrix corresponding to a square block.

[0554] Furthermore, as mentioned above, each component can also be a circuit. These circuits can either form a single circuit as a whole or be separate circuits. Additionally, each component can be implemented using either a general-purpose processor or a dedicated processor.

[0555] Alternatively, other components may perform the processing performed by a specific component. Furthermore, the order of processing can be changed, and multiple processes can be performed simultaneously. Alternatively, the encoding / decoding apparatus may include an encoding device 100 and a decoding device 200.

[0556] Additionally, the ordinal numbers 1 and 2 used in the description can be replaced appropriately. Furthermore, new ordinal numbers can be provided for constituent elements, or ordinal numbers can be removed.

[0557] The above description illustrates the configurations of the encoding device 100 and the decoding device 200 based on the embodiments, but the configurations of the encoding device 100 and the decoding device 200 are not limited to this embodiment. As long as they do not depart from the spirit of the invention, various modifications that can be conceived by those skilled in the art to this embodiment, and configurations constructed by combining the constituent elements of different embodiments, can also be included within the scope of the configurations of the encoding device 100 and the decoding device 200.

[0558] This method can also be implemented in combination with at least a portion of other methods in this invention. Furthermore, a portion of the processing described in the flowchart of this method, a portion of the device structure, a portion of the syntax, etc., can be combined with other methods to implement it.

[0559] (Implementation Method 2)

[0560] In the above embodiments, each functional block is typically implemented using an MPU and memory. Furthermore, the processing of each functional block is usually achieved by a program execution unit such as a processor reading and executing software (programs) recorded in a recording medium such as ROM. This software can be distributed via download or by recording it in a recording medium such as semiconductor memory. Alternatively, each functional block can also be implemented in hardware (dedicated circuitry).

[0561] Furthermore, the processing described in each embodiment can be implemented either centrally using a single device (system) or distributedly using multiple devices. Additionally, the processor executing the above-described program can be either single or multiple. That is, it can be either centrally processed or distributed.

[0562] The present invention is not limited to the above embodiments, and various modifications can be made, which are also included within the scope of the present invention.

[0563] Furthermore, examples of applications of the motion picture encoding method (image encoding method) or motion picture decoding method (image decoding method) shown in the above embodiments and systems using them will be described here. The system is characterized by having an image encoding apparatus using the image encoding method, an image decoding apparatus using the image decoding method, and an image encoding / decoding apparatus possessing both. Other structures within the system can be appropriately modified as needed.

[0564] [Usage Example]

[0565] Figure 34 This is a diagram showing the overall structure of the content delivery system ex100 that implements content distribution services. The provision of communication services is divided into desired sizes, and each unit is equipped with base stations ex106, ex107, ex108, ex109, and ex110, which serve as fixed wireless stations.

[0566] In this content delivery system ex100, various devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. This content delivery system ex100 can also combine some of the above elements for connection. Alternatively, the devices can be directly or indirectly interconnected via telephone networks or short-range wireless connections without using base stations ex106 to ex110, which are fixed wireless stations. Furthermore, a streaming media server ex103 is connected to the computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via the Internet ex101, etc. Additionally, the streaming media server ex103 is connected to terminals within a hotspot inside an aircraft ex117 via satellite ex116.

[0567] Alternatively, it can replace base stations ex106 to ex110 by using wireless access points or hotspots. Furthermore, the streaming media server ex103 can connect directly to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or it can connect directly to the aircraft ex117 without going through the satellite ex116.

[0568] The camera ex113 is a digital camera or similar device capable of capturing still and moving images. Additionally, the smartphone ex115 refers to a smartphone, mobile phone, or PHS (Personal Handyphone System) corresponding to mobile communication systems commonly referred to as 2G, 3G, 3.9G, 4G, and the future 5G.

[0569] Home appliance EX118 refers to refrigerators or other equipment included in a home fuel cell combined heat and power system.

[0570] In the content supply system ex100, terminals with photography capabilities are connected to the streaming media server ex103 via a base station ex106, enabling on-site distribution. During on-site distribution, terminals (such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, and terminals within airplanes ex117) perform encoding processing on still or moving images captured by the user using these terminals, as described in the above embodiments. The encoded image data and the encoded audio data are multiplexed, and the resulting data is sent to the streaming media server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present invention.

[0571] On the other hand, the streaming media server ex103 distributes the content data sent by requesting clients. Clients are terminals such as computers ex111, game consoles ex112, cameras ex113, home appliances ex114, smartphones ex115, or airplanes ex117 capable of decoding the encoded data. Each device receiving the distributed data decodes and reproduces the received data. That is, each device functions as an image decoding device according to one aspect of the present invention.

[0572] [Distributed processing]

[0573] Furthermore, the streaming media server ex103 can also be multiple servers or multiple computers, distributing data through decentralized processing or recording. For example, the streaming media server ex103 can also be implemented by a CDN (Contents Delivery Network), which distributes content by connecting many edge servers scattered around the world. In a CDN, physically nearby edge servers are dynamically allocated based on the client. Furthermore, by caching and distributing content to these edge servers, latency can be reduced. Moreover, in the event of an error or a change in communication status due to increased traffic, processing can be distributed across multiple edge servers, the distribution entity can be switched to other edge servers, or the faulty part of the network can be bypassed to continue distribution, thus achieving high-speed and stable distribution.

[0574] Furthermore, beyond the decentralized processing of distribution itself, the encoding processing of the captured data can be performed by each terminal, on the server side, or shared among them. For example, the encoding process typically involves two processing loops. In the first loop, the complexity or code size of the image per frame or scene unit is detected. In the second loop, processing is performed to maintain image quality while improving encoding efficiency. For instance, by having the terminal perform the first encoding processing and the server receiving the content perform the second, the processing load on each terminal can be reduced while improving both content quality and efficiency. In this case, if there is a request for near real-time reception and decoding, the data encoded in the first loop by the terminal can be received and reproduced by other terminals, enabling more flexible real-time distribution.

[0575] As another example, cameras like the ex113 extract features from images, compress the data about these features as metadata, and send it to a server. The server, for instance, adjusts the quantization precision based on the features to determine the importance of the target, performing compression that corresponds to the meaning of the image. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, while more demanding encoding methods like CABAC (Context Adaptive Binary Arithmetic Coding) can be used by the server.

[0576] As another example, in stadiums, shopping malls, or factories, there may be multiple images of roughly the same scene captured by multiple terminals. In such cases, the data is distributed and encoded separately using the terminals that captured the images, as well as other terminals and servers that did not capture images, as needed. This can be done by assigning encoding and processing data to different units, such as GOP (Group of Pictures), image units, or tile units obtained by segmenting images. This reduces latency and improves real-time performance.

[0577] Furthermore, since multiple image datasets depict roughly the same scene, the server can manage and / or instruct the data to be cross-referenced between images captured by different terminals. Alternatively, the server can receive encoded data from each terminal and change the reference relationships between the multiple datasets, or modify or replace the images themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each data item.

[0578] In addition, the server can also transcode the image data by changing its encoding method before distributing it. For example, the server can convert MPEG encoding to VP encoding, or H.264 to H.265.

[0579] In this way, encoding processing can be performed by a terminal or one or more servers. Therefore, the terms "server" or "terminal" will be used below to refer to the main body performing the processing, but it is also possible to perform part or all of the processing performed by the server by the terminal, or vice versa. Furthermore, the same applies to decoding processing.

[0580] [3D, Multi-angle]

[0581] In recent years, there has been an increase in the use of images or videos captured by multiple cameras (ex113 and / or smartphones (ex115)) at roughly the same time, capturing different scenes or capturing the same scene from different angles. These images are then merged based on the relative positions of the cameras or regions containing consistent feature points within the images.

[0582] The server not only encodes 2D moving images, but can also automatically or at user-specified times encode still images based on scene analysis of moving images and send them to the receiving terminal. Furthermore, when the server can obtain the relative positions between shooting terminals, it can generate 3D shapes of scenes not only from 2D moving images, but also from images of the same scene captured from different angles. Additionally, the server can separately encode 3D data generated from point clouds, and can also select or reconstruct images from multiple terminals based on the results of identifying or tracking people or targets using 3D data to generate images to be sent to the receiving terminal.

[0583] In this way, users can freely select images corresponding to each shooting terminal to appreciate the scene, and can also appreciate the content of images extracted from any viewpoint from 3D data reconstructed using multiple images or videos. Furthermore, similar to the images, sound can also be collected from multiple different angles, and the server can match it with the images to multiplex and send sound from a specific angle or space with the images.

[0584] Furthermore, in recent years, content that establishes a correspondence between the real world and the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has become increasingly popular. In the case of VR images, the server creates separate viewpoint images for the right and left eyes. This can be achieved through Multi-View Coding (MVC) to allow reference between the viewpoint images, or by encoding them as separate streams without reference to each other. During the decoding of these different streams, they can be synchronously reproduced according to the user's viewpoint to recreate a virtual three-dimensional space.

[0585] In the case of AR images, the server can overlay virtual object information in virtual space onto camera information in real space based on the 3D position or the user's viewpoint movement. The decoding device acquires or holds the virtual object information and 3D data, generates a 2D image based on the user's viewpoint movement, and creates overlay data by smoothly connecting them. Alternatively, the decoding device can send the user's viewpoint movement to the server in addition to the virtual object information. The server creates overlay data by matching the received viewpoint movement with the 3D data held on the server, encodes the overlay data, and distributes it to the decoding device. Furthermore, the overlay data has an α value representing transmittance in addition to RGB. The server sets the α value of the portion outside the target created from the 3D data to 0, etc., and encodes the portion in a state of transmittance. Alternatively, the server can set a predetermined RGB value as the background, similar to a chroma key, to generate data where the portion outside the target is set as the background color.

[0586] Similarly, the decoding of distributed data can be performed by the individual terminals acting as clients, on the server side, or distributed among them. For example, one terminal could first send a receive request to the server, other terminals could receive the content corresponding to that request, decode it, and then send the decoded signal to a device with a display. By distributing the processing independently of the capabilities of the communicating terminals and selecting appropriate content, data with better image quality can be reproduced. Furthermore, as another example, large-format image data can be received by a TV, and the viewer's personal terminal can decode and display a segmented area of ​​the image, such as tiles. This allows for the sharing of the overall image while simultaneously allowing the viewer to identify their own area of ​​responsibility or areas they wish to examine in more detail.

[0587] Furthermore, it is envisioned that in the future, with the availability of multiple short-range, medium-range, or long-range wireless communications both indoors and outdoors, content can be seamlessly received while switching appropriate data between connected communications using distribution system standards such as MPEG-DASH. This would allow users to switch in real-time not only using their own terminals but also freely choosing decoding or display devices such as monitors installed indoors or outdoors. Additionally, decoding can be performed by switching between decoding and display terminals based on the user's location information. This would also allow movement towards a destination while displaying map information on a portion of the wall or ground of a building adjacent to a display device. Furthermore, the bit rate of the received data can be switched based on the ease of accessing encoded data to the network, such as when the encoded data is cached on a server accessible only briefly from the receiving terminal or copied to an edge server of the content distribution service.

[0588] [Hyper-level coding]

[0589] Regarding content switching, use Figure 35 The scalable stream, which is compressed using the motion picture coding method described in the above embodiments, will be explained. For the server, there can be multiple streams with the same content but different qualities, or the content structure can be switched using the characteristics of a temporally / spatially scalable stream achieved through layered coding, as shown in the illustration. That is, by having the decoding side determine which layer to decode to based on intrinsic factors such as performance and extrinsic factors such as the state of the communication band, the decoding side can freely switch between decoding low-resolution and high-resolution content. For example, if one wants to watch a follow-up video viewed on a smartphone (ex115) while on the go, and then watch it at home on an internet TV or similar device, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.

[0590] Furthermore, in addition to the hierarchical structure described above, where images are encoded layer by layer and enhancement layers exist above the base layers, the enhancement layer can also contain metadata such as image statistics. The decoding side then uses this metadata to perform super-resolution on the base layer images to generate high-quality content. Super-resolution can be either an increase in the signal-to-noise ratio (SN ratio) at the same resolution or an increase in resolution. The metadata includes information used to determine the linear or nonlinear filtering coefficients used in the super-resolution process, or information determining the parameter values ​​for filtering, machine learning, or least-squares operations used in the super-resolution process.

[0591] Alternatively, the image can be segmented into tiles based on the meaning of targets within it, and the decoding side can decode only a portion of the region by selecting the tiles to be decoded. Furthermore, by storing the target's attributes (people, cars, balls, etc.) and its position within the image (coordinates within the same image, etc.) as metadata, the decoding side can determine the desired target's location based on this metadata and decide which tiles to include that target. For example, ... Figure 36 As shown, metadata is stored using data storage structures different from pixel data, such as SEI messages in HEVC. This metadata may represent, for example, the position, size, or color of the main target.

[0592] Furthermore, metadata can be stored in units consisting of multiple images, such as streams, sequences, or random access units. This allows the decoder to obtain information such as the time a specific person appears in the image, and by matching this information with the image unit information, it can determine the image containing the target and the target's location within that image.

[0593] [Web page optimization]

[0594] Figure 37 This is an example of a web page display screen in a computer such as ex111. Figure 38 This is an example image showing the display screen of a web page in a smartphone such as the ex115. Figure 37 and Figure 38 As shown, in cases where a web page contains multiple linked images that serve as links to image content, their visibility varies depending on the viewing device. When multiple linked images are visible on the screen, before the user explicitly selects a linked image, or before the linked image is near the center of the screen, or before the entire linked image enters the screen, the display device (decoding device) displays still images or I-images of each content as linked images, or displays images like GIF animations using multiple still images or I-images, or only receives the basic layer and decodes and displays the image.

[0595] When a user selects a linked image, the display device prioritizes decoding the base layer. Additionally, if the HTML constituting the webpage contains information indicating tiered content, the display device can also decode up to the enhancement layer. Furthermore, in situations where real-time performance is ensured before selection or when communication bandwidth is extremely limited, the display device can reduce the delay between decoding and displaying the first image (the delay from the start of content decoding to the start of display) by decoding and displaying only the preceding reference images (I images, P images, and B images only used for preceding reference). Alternatively, the display device can forcibly ignore image reference relationships and coarsely decode all B and P images as preceding references, performing normal decoding as more images are received over time.

[0596] [Autonomous Driving]

[0597] Furthermore, when receiving and transmitting still images or video data such as two-dimensional or three-dimensional map information for the purpose of autonomous driving or driving assistance, the receiving terminal can also receive information such as weather or construction as metadata, in addition to image data belonging to more than one layer, and decode them by establishing correspondences. Moreover, metadata can belong to a layer or be multiplexed only with image data.

[0598] In this scenario, since the vehicle, drone, or aircraft containing the receiving terminal is in motion, the receiving terminal can seamlessly receive and decode data by switching between base stations ex106 to ex110 by sending its location information when a request is received. Furthermore, the receiving terminal can dynamically adjust the level of metadata reception or map information updates based on user selection, user status, or the status of the communication frequency band.

[0599] As described above, in the content delivery system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0600] Distribution of personal content

[0601] Furthermore, the content delivery system ex100 not only handles high-quality, long-duration content provided by video distribution providers, but also enables unicast or multicast distribution of low-quality, short-duration content provided by individuals. Moreover, it's conceivable that such personal content will increase in the future. To improve the quality of personal content, the server can also perform encoding processing after editing. This can be achieved, for example, through a structure like the following.

[0602] During or after capturing images, the server performs image processing such as shooting errors, scene search, meaning analysis, and target detection based on the original images or encoded data. Furthermore, based on the recognition results, the server manually or automatically corrects focus deviations or camera shake, deletes scenes of lower importance (e.g., scenes with lower brightness or out-of-focus areas), emphasizes target edges, or adjusts tones. The server then encodes the edited data. Additionally, recognizing that longer shooting times reduce audiovisual quality, the server can automatically limit not only low-importance scenes as described above but also scenes with minimal movement to specific timeframes based on image processing results. Alternatively, the server can generate and encode summaries based on the meaning analysis results of the scenes.

[0603] Furthermore, personal content may contain elements that infringe on copyrights, the author's moral rights, or portrait rights in its original state, or the sharing scope may exceed the intended scope, causing inconvenience to the individual. Therefore, for example, the server could forcibly encode images such as faces of people in the periphery of a scene or a home, replacing them with out-of-focus images. Additionally, the server could identify whether a face different from a pre-registered person is captured within the image to be encoded, and if so, apply processing such as mosaic to the face. Alternatively, as pre- or post-processing for encoding, from a copyright perspective, the user can specify the person or background area to be processed, and the server can replace the specified area with another image or blur the focus. If it is a person, the face image can be replaced while tracking the person in a moving image.

[0604] Furthermore, personal content with small data volumes has strong real-time requirements for audiovisual presentation. Therefore, although bandwidth also plays a role, the decoding device prioritizes receiving, decoding, and reproducing the base layer. The decoding device can also receive enhancement layers during this process, including them in high-quality image reproduction if the playback is looped or repeated more than twice. In this way, if the stream is scalably encoded, it can provide an experience where the motion images are initially coarse but gradually become smoother and the image quality improves. Besides scalable encoding, the same experience can be provided when the first, coarser stream and a second stream encoded based on the first motion image are combined into a single stream.

[0605] [Other Use Cases]

[0606] Furthermore, these encoding or decoding processes are typically handled within the LSIex500 chip present in each terminal. The LSIex500 can be a single chip or a multi-chip structure. Alternatively, software for motion image encoding or decoding can be installed on a recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer such as the ex111, and the encoding and decoding processes can be performed using this software. Furthermore, when the ex115 smartphone has a camera, motion image data acquired by that camera can also be transmitted. This motion image data is encoded using the LSIex500 chip present in the ex115 smartphone.

[0607] Alternatively, the LSIex500 can also be a structure that downloads and activates application software. In this case, the terminal first determines whether it corresponds to the content's encoding method or whether it has the capability to execute a specific service. If the terminal does not correspond to the content's encoding method or does not have the capability to execute a specific service, the terminal downloads the codec or application software, and then retrieves and reproduces the content.

[0608] Furthermore, not limited to the content delivery system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of the above embodiments can also be assembled in a digital broadcasting system. Since multiplexed data that multiplexes images and sound is carried and transmitted and received using radio waves for broadcasting via satellites or the like, it is more suitable for multicasting than the unicast-friendly structure of the content delivery system ex100, but the same applications can be performed for encoding and decoding processing.

[0609] [Hardware Structure]

[0610] Figure 39 This is a diagram representing the EX115 smartphone. Additionally, Figure 40This diagram illustrates a structural example of a smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with a base station ex110, a camera unit ex465 for capturing images and still images, and a display unit ex458 for displaying decoded data such as images captured by the camera unit ex465 and images received by the antenna ex450. The smartphone ex115 also includes an operation unit ex466, such as a touch panel; a sound output unit ex457, such as a speaker, for outputting sound or audio; a sound input unit ex456, such as a microphone, for inputting sound; a memory unit ex467 capable of storing encoded or decoded data such as captured images or still images, recorded audio, received images or still images, and emails; and a slot unit ex464 serving as an interface with a SIM ex468 used to identify the user and authenticate access to various data sources, such as the network. Alternatively, an external memory may be used instead of the memory unit ex467.

[0611] Furthermore, the main control unit ex460, which performs integrated control of the display unit ex458 and the operation unit ex466, is interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the image signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the bus ex470.

[0612] If the power button is turned on by the user, the power circuit section ex461 will supply power to each part from the battery pack, thus activating the smartphone ex115 into an operational state.

[0613] The smart phone ex115 is controlled by a main control unit ex460, which includes a CPU, ROM, and RAM, for call and data communication processing. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454. This digital signal is then subjected to spectral diffusion processing by the modulation / demodulation unit ex452, and finally, digital-to-analog conversion and frequency conversion processing by the transmitting / receiving unit ex451 before being transmitted via the antenna ex450. Similarly, received data is amplified and subjected to frequency conversion and analog-to-digital conversion processing. The data is then subjected to inverse spectral diffusion processing by the modulation / demodulation unit ex452, converted into an analog audio signal by the audio signal processing unit ex454, and output from the audio output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 through the operation unit ex466 of the main unit, and are also processed for transmission and reception. In data communication mode, when transmitting images, still images, or images and sound, the image signal processing unit ex455 compresses and encodes the image signal stored in the memory unit ex467 or the image signal input from the camera unit ex465 using the moving image encoding method described in the above embodiments, and sends the encoded image data to the multiplexing / demultiplexing unit ex453. Furthermore, the sound signal processing unit ex454 encodes the sound signal collected by the sound input unit ex456 during the capture of images, still images, etc., by the camera unit ex465, and sends the encoded sound data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded image data and encoded sound data in a prescribed manner, performs modulation and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmitting / receiving unit ex451, and transmits the data via the antenna ex450.

[0614] When receiving images attached to emails or chat tools, or images linked to web pages, the multiplexing / demultiplexing unit ex453, in order to decode the multiplexed data received via antenna ex450, separates the multiplexed data into a bitstream of image data and a bitstream of audio data. The encoded image data is supplied to the image signal processing unit ex455 via the synchronization bus ex470, and the encoded audio data is supplied to the audio signal processing unit ex454. The image signal processing unit ex455 decodes the image signal using a motion picture decoding method corresponding to the motion picture encoding method described in the above embodiments, and displays the image or still image contained in the linked motion picture file from the display unit ex458 via the display control unit ex459. Furthermore, the audio signal processing unit ex454 decodes the audio signal and outputs audio from the audio output unit ex457. Additionally, since real-time streaming is becoming increasingly common, depending on the user's situation, the reproduction of audio may be unsuitable for certain situations. Therefore, as an initial consideration, a structure that reproduces only the image data and not the audio signal is preferred. It can also reproduce sound synchronously only when the user performs actions such as clicking on the image data.

[0615] Furthermore, while the example given here is the ex115 smartphone, as a terminal, three installation methods can be considered: a transmitting terminal with both an encoder and a decoder, a transmitting terminal with only an encoder, and a receiving terminal with only a decoder. Moreover, in a digital broadcasting system, the example described involves receiving and transmitting multiplexed data, such as audio data, multiplexed within video data. However, in addition to audio data, multiplexed data can also multiplex character data associated with the video, or the video data itself can be received or transmitted without using multiplexed data.

[0616] Furthermore, while the description assumes the CPU's main control unit (ex460) controls the encoding or decoding process, many terminals also possess GPUs. Therefore, a structure can be implemented that utilizes GPU performance to process larger regions simultaneously, using shared memory between the CPU and GPU, or memory that manages addresses in a shared manner. This reduces encoding time, ensures real-time performance, and achieves low latency. In particular, it is even more effective if motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization processing are performed on the GPU at the image level, rather than by the CPU.

[0617] Industrial availability

[0618] This invention can be used in, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, video conferencing systems, or electronic mirrors.

[0619] Label Explanation

[0620] 100 encoding device

[0621] 102 Division

[0622] 104 Subtraction Section

[0623] 106 Transformer

[0624] Quantitative Department 108

[0625] 110 Entropy Coding Department

[0626] 112, 204 Inverse Quantization Section

[0627] Inverse Transformation Units 114 and 206

[0628] 116, 208 Addition Department

[0629] 118 and 210 memory blocks

[0630] 120, 212 Circular Filter Section

[0631] 122, 214 frame memory

[0632] Intra-frame prediction units 124 and 216

[0633] Inter-frame prediction units 126 and 218

[0634] Predictive Control Department 128, 220

[0635] 160, 260 circuits

[0636] 162, 262 memory

[0637] 200 decoding device

[0638] 202 Entropy Decoding Department

Claims

1. An encoding device comprising: Circuits; and The memory, coupled to the circuit described above, wherein, The circuit described above uses the memory described above to perform the following processing: The first quantization matrix is ​​up-transformed and down-transformed to generate the second quantization matrix. The first quantization matrix has a first row and a column equal to the first row to form a square matrix. The second quantization matrix has a second row and a column different from the second row to form a rectangular matrix. The transformation coefficients of the current block are quantized using the second quantization matrix described above, where, In the aforementioned upconversion, the circuit performs the upconversion in the first direction such that one of the second row number and the second column number is greater than the first row number, generating the second quantization matrix. The first direction is the horizontal direction. In the downconversion described above, the circuit performs the downconversion in the second direction in such a way that the number of the second row and the number of the second column are both less than the number of the first row, thereby generating the second quantization matrix. The second direction is different from the first direction and is a vertical direction.

2. A decoding device, comprising: Circuits; and The memory, coupled to the circuit described above, wherein, The circuit described above uses the memory described above to perform the following processing: The first quantization matrix is ​​up-transformed and down-transformed to generate the second quantization matrix. The first quantization matrix has a first row and a column equal to the first row to form a square matrix. The second quantization matrix has a second row and a column different from the second row to form a rectangular matrix. The second quantization matrix described above is used to inversely quantize the quantization coefficients of the current block, where, In the aforementioned upconversion, the circuit performs the upconversion in the first direction such that one of the second row number and the second column number is greater than the first row number, generating the second quantization matrix. The first direction is the horizontal direction. In the downconversion described above, the circuit performs the downconversion in the second direction in such a way that the number of the second row and the number of the second column are both less than the number of the first row, thereby generating the second quantization matrix. The second direction is different from the first direction and is a vertical direction.

3. A non-transitory computer-readable medium for storing a bit stream and a computer program containing executable instructions, The aforementioned bitstream contains the quantization coefficients of the current block, which are subjected to inverse quantization processing by the decoding device according to the aforementioned instructions, wherein, In the above inverse quantization process: The first quantization matrix is ​​obtained, which has a first row number and a column number equal to the first row number to form a square matrix; The second quantization matrix is ​​generated by performing up and down transformations on the first quantization matrix. This second quantization matrix has a second row and a second column (different from the number of rows) to form a rectangular matrix; and... The inverse quantization transform coefficients of the current block are generated from the quantization coefficients using the second quantization matrix, wherein, The first quantization matrix is ​​upconverted in the first direction, which is the horizontal direction, such that either the number of the second row or the number of the second column is greater than the number of the first row. The first quantization matrix is ​​down-converted in the second direction in such a way that the number of the second row and the number of the second column are both less than the number of the first row. The second direction is different from the first direction and is a vertical direction.

Citation Information

Patent Citations

  • Image processing device and method

    CN104469376A

  • Signaling quantization matrices for video coding

    US20130114695A1