Encoding device, decoding device, encoding method, and decoding method

The encoding device addresses the inefficiency in the quadratic conversion process by applying a linear transformation and optimizing the quadratic transformation on the predicted residual signal, resulting in reduced processing and improved efficiency.

JP2025071223AActive Publication Date: 2025-05-02PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025025546
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-07-13
Filing Date
2025-02-20
Publication Date
2025-05-02
Estimated Expiration
2039-06-07

AI Technical Summary

Technical Problem

The processing amount increases in the quadratic conversion process where the encoding device applies a primary conversion process to the predicted residual signal, leading to inefficiencies.

Method used

An encoding device that applies a linear transformation to the predicted residual signal and performs a quadratic transformation on the first transform coefficient to generate a second transform coefficient, optimizing the block size and transformation basis selection to reduce processing.

Benefits of technology

The proposed solution reduces the processing amount in the quadratic transformation process, enhancing the encoding efficiency compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071223000001_ABST
    Figure 2025071223000001_ABST
Patent Text Reader

Abstract

To provide an encoding device capable of reducing the amount of processing.SOLUTION: An encoding device (100) includes a circuit and a memory. The circuit applies a primary transform to a prediction residual signal; performs a transform process in which a secondary transform is further applied to first transform coefficients resulting from the transform to generate second transform coefficients; quantizes the second transform coefficients, in which the secondary transform selects one transform base from a first candidate group in the case of a first block size and selects one transform base from a second candidate group in the case of a second block size; and applies the secondary transform to some of the first transform coefficients in the case of a block size larger than 4×4, where a size of a sub-block to which the secondary transform is applied in the current block having the first block size is the same as a size of a sub-block to which the secondary transform is applied in the current block having the second block size.SELECTED DRAWING: Figure 53
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an encoding device and the like that encodes a moving image including a plurality of pictures. [Background technology]

[0002] 2. Description of the Related Art Conventionally, H.265, also known as High Efficiency Video Coding (HEVC), exists as a standard for encoding moving images (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] H.265(ISO / IEC 23008-2 HEVC) / HEVC(High Efficiency Video Coding) Summary of the Invention [Problem to be solved by the invention]

[0004] However, there is a problem in that the amount of processing required increases in the secondary transform process that an encoding device or the like applies to transform coefficients resulting from the primary transform process applied to the prediction residual signal.

[0005] Therefore, the present disclosure provides an encoding device, etc. that can reduce the amount of processing compared to conventional methods in a secondary transformation process that the encoding device, etc. applies to transformation coefficients obtained by applying a primary transformation process to a prediction residual signal. [Means for solving the problem]

[0006] An encoding device according to one aspect of the present disclosure includes a circuit and a memory, and the circuit uses the memory to apply a primary transform to a prediction residual signal indicating a difference between a current block to be encoded and a predicted image of the current block, and performs a transform process in which a secondary transform is further applied to a first transform coefficient that is a transform result of the primary transform to generate a second transform coefficient of the current block, and quantizes the second transform coefficient. In the secondary transform, if a block size of the current block is a first block size, one transform base is selected from a first candidate group consisting of one or more transform base candidates, and if a block size of the current block is a second block size different from the first block size, one transform base is selected from a second candidate group different from the first candidate group. Whenever the block size is greater than 4×4, the secondary transform is applied to a portion of the first transform coefficient, and a size of a sub-block of the current block having the first block size to which the secondary transform is applied is the same as a size of a sub-block of the current block having the second block size to which the secondary transform is applied.

[0007] In addition, these comprehensive or specific aspects may be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. Effect of the Invention

[0008] An encoding device or the like according to an aspect of the present disclosure can reduce the amount of processing compared to conventional methods in a secondary transform that is further applied to transform coefficients obtained by applying a primary transform to a prediction residual signal. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of an encoding device according to an embodiment. [Diagram 2]FIG. 2 is a flowchart showing an example of the overall encoding process performed by the encoding device. [Diagram 3] FIG. 3 is a diagram showing an example of block division. [Figure 4A] FIG. 4A is a diagram showing an example of a slice configuration. [Figure 4B] FIG. 4B is a diagram showing an example of a tile configuration. [Figure 5A] FIG. 5A is a table showing the transform basis functions that correspond to each transform type. [Figure 5B] FIG. 5B is a diagram showing SVT (Spatially Varying Transform). [Figure 6A] FIG. 6A is a diagram showing an example of a filter shape used in an adaptive loop filter (ALF). [Figure 6B] FIG. 6B is a diagram showing another example of the shape of the filter used in the ALF. [Figure 6C] FIG. 6C is a diagram showing another example of the shape of the filter used in the ALF. [Figure 7] FIG. 7 is a block diagram showing an example of a detailed configuration of the loop filter unit functioning as the DBF. [Figure 8] FIG. 8 is a diagram showing an example of a deblocking filter having filter characteristics that are symmetric with respect to block boundaries. [Figure 9] FIG. 9 is a diagram for explaining block boundaries on which deblocking filter processing is performed. [Figure 10] FIG. 10 is a diagram showing an example of the Bs value. [Figure 11] FIG. 11 is a diagram illustrating an example of a process performed by the prediction processing unit of the encoding device. [Figure 12] FIG. 12 is a diagram illustrating another example of the process performed in the prediction processing unit of the encoding device. [Figure 13] FIG. 13 is a diagram illustrating another example of the process performed in the prediction processing unit of the encoding device. [Figure 14]FIG. 14 is a diagram showing an example of 67 intra prediction modes in intra prediction. [Figure 15] FIG. 15 is a flowchart showing the flow of basic inter prediction processing. [Figure 16] FIG. 16 is a flowchart showing an example of motion vector derivation. [Figure 17] FIG. 17 is a flowchart showing another example of motion vector derivation. [Figure 18] FIG. 18 is a flowchart showing another example of motion vector derivation. [Figure 19] FIG. 19 is a flowchart showing an example of inter prediction in the normal inter mode. [Figure 20] FIG. 20 is a flowchart showing an example of inter prediction in the merge mode. [Figure 21] FIG. 21 is a diagram for explaining an example of a motion vector derivation process in the merge mode. [Figure 22] FIG. 22 is a flowchart showing an example of frame rate up conversion (FRUC). [Figure 23] FIG. 23 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 24] FIG. 24 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 25A] FIG. 25A is a diagram for explaining an example of derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 25B] FIG. 25B is a diagram for explaining an example of derivation of a motion vector for each sub-block in the affine mode having three control points. [Figure 26A] FIG. 26A is a conceptual diagram for explaining the affine merge mode. [Figure 26B]FIG. 26B is a conceptual diagram for explaining an affine merge mode having two control points. [Figure 26C] FIG. 26C is a conceptual diagram for explaining an affine merge mode having three control points. [Figure 27] FIG. 27 is a flowchart showing an example of a process in the affine merge mode. [Figure 28A] FIG. 28A is a diagram for explaining an affine inter mode having two control points. [Figure 28B] FIG. 28B is a diagram for explaining an affine inter mode having three control points. [Figure 29] FIG. 29 is a flowchart showing an example of processing in the affine inter mode. [Figure 30A] FIG. 30A is a diagram illustrating an affine inter mode in which a current block has three control points and an adjacent block has two control points. [Figure 30B] FIG. 30B is a diagram illustrating an affine inter mode in which a current block has two control points and an adjacent block has three control points. [Figure 31A] FIG. 31A is a diagram showing the relationship between merge mode and DMVR (dynamic motion vector refreshing). [Figure 31B] FIG. 31B is a conceptual diagram for explaining an example of the DMVR process. [Diagram 32] FIG. 32 is a flowchart showing an example of generation of a predicted image. [Diagram 33] FIG. 33 is a flowchart showing another example of generation of a predicted image. [Diagram 34] FIG. 34 is a flowchart showing yet another example of generation of a predicted image. [Diagram 35] FIG. 35 is a flowchart illustrating an example of a predictive image correction process using overlapped block motion compensation (OBMC). [Diagram 36] FIG. 36 is a conceptual diagram for explaining an example of the predicted image correction process by the OBMC process. [Figure 37] FIG. 37 is a diagram for explaining generation of predicted images of two triangles. [Figure 38] FIG. 38 is a diagram for explaining a model assuming uniform linear motion. [Figure 39] FIG. 39 is a diagram for explaining an example of a predicted image generating method using luminance correction processing by LIC (local illumination compensation) processing. [Diagram 40] FIG. 40 is a block diagram showing an example of implementation of an encoding device. [Diagram 41] FIG. 41 is a block diagram showing a functional configuration of a decoding device according to an embodiment. As shown in FIG. [Diagram 42] FIG. 42 is a flowchart showing an example of the overall decoding process by the decoding device. [Diagram 43] FIG. 43 is a diagram illustrating an example of processing performed in the prediction processing unit of the decoding device. [Diagram 44] FIG. 44 is a diagram illustrating another example of the process performed in the prediction processing unit of the decoding device. [Diagram 45] FIG. 45 is a flowchart showing an example of inter prediction in the normal inter mode in the decoding device. [Figure 46] FIG. 46 is a block diagram showing an implementation example of a decoding device. [Figure 47] FIG. 47 is a diagram illustrating the secondary conversion process in the embodiment. [Figure 48] FIG. 48 is a flowchart showing a processing procedure in a conversion unit of the encoding device according to the embodiment. [Figure 49A] FIG. 49A is a table showing an example of the amount of processing required for primary conversion processing of the entire CTU in an embodiment. [Figure 49B] FIG. 49B is a table showing an example of the amount of processing required for secondary conversion processing of the entire CTU in an embodiment. [Figure 50] FIG. 50 is a table showing a first example in the embodiment. [Figure 51] FIG. 51 is a table showing a second example in the embodiment. [Figure 52] FIG. 52 is a table showing a third example in the embodiment. [Diagram 53] FIG. 53 is a table showing a fourth example in the embodiment. [Figure 54] FIG. 54 is a flowchart showing an example of the operation of the encoding device in the embodiment. [Figure 55] FIG. 55 is a flowchart showing an example of the operation of the decoding device according to the embodiment. [Figure 56] FIG. 56 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 57] FIG. 57 is a diagram showing an example of a coding structure for scalable coding. [Figure 58] FIG. 58 is a diagram showing an example of a coding structure in scalable coding. [Figure 59] FIG. 59 is a diagram showing an example of a display screen of a web page. [Figure 60] FIG. 60 is a diagram showing an example of a display screen of a web page. [Figure 61] FIG. 61 is a diagram showing an example of a smartphone. [Figure 62] FIG. 62 is a block diagram showing an example configuration of a smartphone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] (Findings on which this disclosure is based) For example, the encoding device or the like may perform a secondary transform, such as an orthogonal transform, on the transform coefficients obtained by applying a primary transform to the prediction residual signal. In this case, the encoding device or the like may apply a secondary transform of a plurality of block sizes to the transform coefficients obtained by applying the primary transform to the prediction residual signal.

[0011] Therefore, for example, an encoding device according to one aspect of the present disclosure includes a circuit and a memory, and the circuit uses the memory to perform a transformation process in which a transform coefficient obtained by applying a linear transform to a prediction residual signal in a target block among a plurality of blocks of a plurality of block sizes is further transformed by applying a secondary transform of a block size common to the plurality of blocks, and the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0012] As a result, when applying a secondary transform of a common block size to a processing target block, the encoding device can select a more appropriate candidate transform base than before and apply the selected candidate transform base to the processing target block. Therefore, the encoding device can reduce the amount of code in the secondary transform process more than before.

[0013] Also, for example, in the encoding device according to an embodiment of the present disclosure, the transform base of the quadratic transform of the common block size is a 4×4 square.

[0014] This allows the encoding device to select a transform base of the smallest size when applying a secondary transform of a common block size to a current block.

[0015] Also, for example, in the encoding device according to an embodiment of the present disclosure, the transform base of the quadratic transform of the common block size is an 8×8 square.

[0016] This allows the decoding device to select a transform base of an appropriate size when applying a secondary transform of a common block size to a current block.

[0017] Also, for example, an encoding device according to one aspect of the present disclosure assigns a common candidate for the transformation base to the candidate group in the secondary transformation for the processing target blocks of a portion of the multiple block sizes.

[0018] This allows the encoding device to reduce the amount of processing compared to the conventional method. For example, the encoding device can reduce the amount of processing by assigning a common base to a 16×16 block to be processed and a 32×32 block to be processed and performing secondary transformation.

[0019] Also, for example, an encoding device according to one aspect of the present disclosure determines not to apply the secondary transformation to the transform coefficients when the block size of the block to be processed is equal to or smaller than a predetermined block size, and determines to apply the secondary transformation to the transform coefficients when the block size of the block to be processed is larger than the predetermined block size.

[0020] As a result, the encoding device can reduce the amount of processing in the transform process compared to conventional methods by not performing secondary transform when the block to be processed has a block size that requires a large amount of processing in the secondary transform.

[0021] Also, for example, in the encoding device according to an embodiment of the present disclosure, the predetermined block size is a 4×4 square.

[0022] As a result, the encoding device can reduce the amount of processing in the transform process compared to conventional methods by not performing secondary transform when the block to be transformed has a block size of 4x4, which requires a large amount of processing in the secondary transform.

[0023] Also, for example, in the encoding device according to an embodiment of the present disclosure, the predetermined block size is a 4×8 or 8×4 rectangle.

[0024] As a result, the encoding device can reduce the amount of processing in the transform process compared to conventional methods by not performing secondary transform when the block to be transformed has a block size of 4x8 or 8x4, which requires a large amount of processing in the secondary transform.

[0025] Also, for example, in the encoding device according to one aspect of the present disclosure, the predetermined block size is equal to the smallest block size among one or more block sizes selectable in the secondary transform.

[0026] As a result, the encoding device can reduce the amount of processing in the conversion process more than before by not performing secondary conversion when the block to be processed on which the conversion process is performed is a block size that results in the greatest amount of processing in the secondary conversion among the sizes selectable by the encoding device.

[0027] Also, for example, a decoding device according to one aspect of the present disclosure includes a circuit and a memory, and the circuit uses the memory to perform an inverse transform process in which a linear transform is applied to transform coefficients obtained by applying a secondary transform of a block size common to a transform coefficient signal to a target block among a plurality of blocks of a plurality of block sizes, and the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0028] As a result, when applying a secondary transform of a common block size to a processing target block, the decoding device can select a more appropriate candidate transform base than in the past and apply the selected candidate transform base to the processing target block. Therefore, the decoding device can reduce the amount of code in the secondary transform process more than in the past.

[0029] Also, for example, in a decoding device according to an embodiment of the present disclosure, a transform base for the quadratic transform of the common block size is a 4×4 square.

[0030] This allows the decoding device to select a transform base of the smallest size when applying a secondary transform of a common block size to a current block.

[0031] Also, for example, in a decoding device according to an embodiment of the present disclosure, a transform base for the quadratic transform of the common block size is an 8×8 square.

[0032] This allows the decoding device to select a transform base of an appropriate size when applying a secondary transform of a common block size to a current block.

[0033] Also, for example, a decoding device according to one aspect of the present disclosure assigns a common candidate for the transformation base to the candidate group in the secondary transformation for the processing target blocks of a portion of the multiple block sizes.

[0034] This allows the decoding device to reduce the amount of processing compared to the conventional method. For example, the decoding device can reduce the amount of processing by assigning a common base to a 16×16 block to be processed and a 32×32 block to be processed and performing secondary transformation.

[0035] Also, for example, a decoding device according to one aspect of the present disclosure determines not to apply the secondary transformation to the transform coefficients when the block size of the block to be processed is equal to or smaller than a predetermined block size, and determines to apply the secondary transformation to the transform coefficients when the block size of the block to be processed is larger than the predetermined block size.

[0036] As a result, the decoding device is able to reduce the amount of processing in the conversion process compared to conventional methods by not performing secondary conversion when the block to be processed on has a block size that results in a large amount of processing in the secondary conversion.

[0037] Also, for example, in the decoding device according to an embodiment of the present disclosure, the predetermined block size is a 4×4 square.

[0038] As a result, the decoding device can reduce the amount of processing in the transform process compared to conventional methods by not performing secondary transform when the block to be transformed has a block size of 4x4, which requires a large amount of processing in the secondary transform.

[0039] Also, for example, in the decoding device according to an embodiment of the present disclosure, the predetermined block size is a 4×8 or 8×4 rectangle.

[0040] As a result, the decoding device can reduce the amount of processing in the transform process compared to conventional methods by not performing secondary transform when the block to be transformed has a block size of 4x8 or 8x4, which requires a large amount of processing in the secondary transform.

[0041] Also, for example, in a decoding device according to an aspect of the present disclosure, the predetermined block size is equal to a smallest block size among one or more block sizes selectable in the secondary transform.

[0042] As a result, the decoding device is able to reduce the amount of processing in the conversion process more than in the past by not performing secondary conversion when the block size is the one that requires the greatest amount of processing in the secondary conversion among the sizes selectable by the decoding device.

[0043] Furthermore, for example, an encoding method according to one aspect of the present disclosure performs a transformation process on a transform coefficient obtained by applying a linear transform to a prediction residual signal in a target block among multiple blocks of multiple block sizes, and further applies a secondary transform of a block size common to the multiple blocks, in which the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0044] As a result, the encoding method can achieve the same effects as the encoding device described above.

[0045] Furthermore, for example, a decoding method according to one aspect of the present disclosure performs an inverse transform process in which a primary transform is applied to transform coefficients obtained by applying a secondary transform of a block size common to a transform coefficient signal to a target block among a plurality of blocks of a plurality of block sizes, and the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0046] As a result, the decoding method can achieve the same effects as the above-mentioned decoding device.

[0047] Also, for example, an encoding device according to one aspect of the present disclosure may include a division unit, an intra prediction unit, an inter prediction unit, a loop filter unit, a transform unit, a quantization unit, and an entropy encoding unit.

[0048] The division unit may divide a picture into a plurality of blocks. The intra prediction unit may perform intra prediction on a block included in the plurality of blocks. The inter prediction unit may perform inter prediction on the block. The transformation unit may transform a prediction error between a predicted image obtained by the intra prediction or the inter prediction and an original image to generate a transformation coefficient. The quantization unit may quantize the transformation coefficient to generate a quantized coefficient. The entropy coding unit may code the quantized coefficient to generate a coded bit stream. The loop filter unit may apply a filter to a reconstructed image of the block.

[0049] Furthermore, for example, the encoding device may be an encoding device that encodes a moving image including a plurality of pictures.

[0050] The transform unit then performs a transform process on transform coefficients obtained by applying a linear transform to a prediction residual signal in a target block among multiple blocks of multiple block sizes, in which the transform unit further applies a secondary transform of a block size common to the multiple blocks, and the secondary transform of the common block size may be composed of one or more candidates for a transform base, and one of the transform bases may be selected from a group of candidates that differ depending on the block size of the target block.

[0051] Furthermore, for example, a decoding device according to one aspect of the present disclosure may include an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, and a loop filter unit.

[0052] The entropy decoding unit may decode quantized coefficients of a block in a picture from the encoded bitstream. The inverse quantization unit may inverse quantize the quantized coefficients to obtain transform coefficients. The inverse transform unit may inverse transform the transform coefficients to obtain prediction errors. The intra prediction unit may perform intra prediction on the block. The inter prediction unit may perform inter prediction on the block. The filter unit may apply a filter to a reconstructed image generated using a predicted image obtained by the intra prediction or the inter prediction and the prediction error.

[0053] Furthermore, for example, the decoding device may be a decoding device that decodes a moving image including a plurality of pictures.

[0054] The inverse transform unit then performs an inverse transform process in which a linear transform is applied to transform coefficients obtained by applying a secondary transform of a block size common to a transform coefficient signal to a target block among a plurality of blocks of a plurality of block sizes, and the secondary transform of the common block size may be composed of one or more candidates for a transform base, and one of the transform bases may be selected from a group of candidates that differ depending on the block size of the target block.

[0055] Furthermore, these comprehensive or specific aspects may be realized in a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized in any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0056] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the relationship and order of the steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0057] In the following, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, for example, any of the following may be implemented.

[0058] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0059] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, deleting, etc. For example, any function or process may be replaced or combined with another function or process described in any of the aspects of the present disclosure.

[0060] (3) In the method implemented by the encoding device or decoding device of the embodiment, some of the processes included in the method may be arbitrarily changed, such as added, replaced, deleted, etc. For example, any process in the method may be replaced or combined with another process described in any of the aspects of the present disclosure.

[0061] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in each aspect of the present disclosure.

[0062] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0063] (6) In a method implemented by an encoding device or a decoding device of an embodiment, any of the multiple processes included in the method may be replaced or combined with a process described in any of the aspects of the present disclosure or with any similar process.

[0064] (7) Some of the processes among the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0065] (8) The manner in which the processes and / or configurations described in each aspect of the present disclosure are implemented is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or configurations may be implemented in a device used for a purpose other than the video encoding or video decoding disclosed in the embodiment.

[0066] (Embodiment 1) [Encoding device] First, a coding device according to the present embodiment will be described. Fig. 1 is a block diagram showing a functional configuration of coding device 100 according to the present embodiment. Coding device 100 is a video coding device that codes a video on a block-by-block basis.

[0067] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0068] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The encoding device 100 may also be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0069] Below, the overall processing flow of the encoding device 100 will be described, and then each component included in the encoding device 100 will be described.

[0070] [Overall encoding process flow] FIG. 2 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0071] First, the division unit 102 of the encoding device 100 divides each picture included in an input image, which is a moving image, into a plurality of fixed-size blocks (128×128 pixels) (step Sa_1). Then, the division unit 102 selects a division pattern (also called a block shape) for the fixed-size blocks (step Sa_2). That is, the division unit 102 further divides the fixed-size block into a plurality of blocks constituting the selected division pattern. Then, the encoding device 100 performs the process of steps Sa_3 to Sa_9 for each of the plurality of blocks (i.e., the block to be encoded).

[0072] That is, a prediction processing unit consisting of all or part of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 generates a prediction signal (also called a prediction block) of the block to be coded (also called a current block) (step Sa_3).

[0073] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also called a difference block) (step Sa_4).

[0074] Next, the transform unit 106 and the quantization unit 108 perform transform and quantization on the difference block to generate a plurality of quantized coefficients (step Sa_5). Note that a block made up of a plurality of quantized coefficients is also called a coefficient block.

[0075] Next, the entropy coding unit 110 performs coding (specifically, entropy coding) on ​​the coefficient block and the prediction parameters related to the generation of the prediction signal to generate a coded signal (step Sa_6). The coded signal is also called a coded bit stream, a compressed bit stream, or a stream.

[0076] Next, the inverse quantization unit 112 and the inverse transformation unit 114 perform inverse quantization and inverse transformation on the coefficient block to reconstruct a plurality of prediction residuals (that is, difference blocks) (step Sa_7).

[0077] Next, the adder 116 reconstructs the current block into a reconstructed image (also called a reconstructed block or a decoded image block) by adding the predicted block to the restored difference block (step Sa_8). In this way, a reconstructed image is generated.

[0078] When this reconstructed image is generated, the loop filter unit 120 performs filtering on the reconstructed image as necessary (step Sa_9).

[0079] Then, the encoding device 100 determines whether or not encoding of the entire picture is completed (step Sa_10), and if it determines that encoding is not completed (No in step Sa_10), repeats the process from step Sa_2.

[0080] In the above example, the encoding device 100 selects one division pattern for fixed-size blocks and encodes each block according to the division pattern, but it may also encode each block according to each of a plurality of division patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of division patterns and select, for example, the encoded signal obtained by encoding according to the division pattern with the smallest cost as the encoded signal to be finally output.

[0081] Furthermore, the processes of steps Sa_1 to Sa_10 may be performed sequentially by encoding device 100, or some of the processes may be performed in parallel, or the order of the processes may be changed.

[0082] [Divided part] The division unit 102 divides each picture included in the input video into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides a picture into blocks of a fixed size (e.g., 128x128). The fixed-size blocks may be called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on, for example, recursive quadtree and / or binary tree block division. That is, the division unit 102 selects a division pattern. The variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in various implementation examples, CUs, PUs, and TUs do not need to be distinguished, and some or all of the blocks in a picture may be the processing units of CUs, PUs, and TUs.

[0083] Fig. 3 is a diagram showing an example of block division in this embodiment, in which solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0084] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0085] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0086] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14, 15 (binary tree block division).

[0087] The bottom left 64x64 block is divided into four square 32x32 blocks (quadtree block division). Of the four 32x32 blocks, the top left and bottom right blocks are further divided. The top left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block division). The bottom right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block division). As a result, the bottom left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0088] The bottom right 64x64 block 23 is not split.

[0089] 3, the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0090] In Fig. 3, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0091] [Picture Composition Slices / Tiles] In order to decode pictures in parallel, the pictures may be arranged in slice units or tile units. A picture arranged in slice units or tile units may be arranged by the division unit 102.

[0092] A slice is a basic coding unit constituting a picture. A picture is made up of, for example, one or more slices. Furthermore, a slice is made up of one or more consecutive coding tree units (CTUs).

[0093] FIG. 4A is a diagram showing an example of a slice configuration. For example, a picture includes 11×8 CTUs and is divided into four slices (slices 1-4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes 29 CTUs, and slice 4 includes 22 CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of the slice is obtained by dividing the picture in the horizontal direction. The boundary of the slice does not need to be the edge of the screen, and may be any boundary of the CTUs in the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, raster scan order. In addition, the slice includes header information and encoded data. The header information may describe the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.

[0094] A tile is a rectangular unit that makes up a picture. Each tile may be assigned a number called a TileId in raster scan order.

[0095] FIG. 4B is a diagram showing an example of a tile configuration. For example, a picture includes 11×8 CTUs and is divided into four rectangular tiles (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed in raster scan order. When tiles are used, at least one CTU is processed in raster scan order in each of multiple tiles. For example, as shown in FIG. 4B, the processing order of multiple CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0096] It should be noted that one tile may include one or more slices, and one slice may include one or more tiles.

[0097] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (a prediction sample input from a prediction control unit 128 described below) from an original signal (original sample) input from the division unit 102 for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also called a residual) of a block to be coded (hereinafter called a current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.

[0098] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0099] [Conversion section] The transform unit 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain, and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0100] The transform unit 106 may adaptively select a transform type from among a plurality of transform types, and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform may be called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0101] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transform basis functions corresponding to each transform type. In Figure 5A, N indicates the number of input pixels. The selection of the transform type from among the multiple transform types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0102] Such information indicating whether EMT or AMT is applied (e.g., called an EMT flag or an AMT flag) and information indicating the selected transformation type are usually signaled at a CU level, but the signaling of such information does not need to be limited to the CU level and may be at other levels (e.g., a bit sequence level, a picture level, a slice level, a tile level, or a CTU level).

[0103] Furthermore, the transform unit 106 may retransform the transform coefficients (transformation results). Such retransformation may be called an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transform unit 106 performs retransformation for each subblock (e.g., 4x4 subblock) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether or not to apply NSST and information regarding a transform matrix used in NSST are usually signaled at a CU level. Note that signaling of these pieces of information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0104] Separable transformation and non-separable transformation may be applied to the transformation unit 106. Separable transformation is a method of performing transformation multiple times by separating the input into directions for the number of dimensions, and non-separable transformation is a method of performing transformation collectively when the input is multidimensional, regarding two or more dimensions as one dimension.

[0105] For example, one example of a non-separable transformation is one in which, if the input is a 4x4 block, it is treated as a single array with 16 elements and a transformation process is performed on that array using a 16x16 transformation matrix.

[0106] As another example of a non-separable transformation, a 4x4 input block may be treated as a single array with 16 elements, and then a transformation (Hypercube Givens Transform) may be performed on the array by performing multiple Givens rotations.

[0107] In the transform in the transform unit 106, the type of basis for transforming into the frequency domain can be switched according to the area in the CU. One example is SVT (Spatially Varying Transform). In SVT, as shown in FIG. 5B, a CU is divided into two equal parts in the horizontal or vertical direction, and only one of the areas is transformed into the frequency domain. The type of transform basis can be set for each area, and for example, DST7 and DCT8 are used. In this example, only one of the two areas in the CU is transformed and the other is not transformed, but both areas may be transformed. In addition, the division method can be more flexible, such as not only dividing into two, but also dividing into four equal parts, or separately encoding information indicating the division and signaling it in the same way as the CU division. In addition, SVT is also called SBT (Sub-block Transform).

[0108] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter, referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0109] The predetermined scanning order is an order for quantizing / dequantizing transform coefficients, for example, the predetermined scanning order is defined as ascending (low to high) or descending (high to low) frequency order.

[0110] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0111] In addition, a quantization matrix may be used for quantization. For example, several kinds of quantization matrices may be used corresponding to frequency transform sizes such as 4x4 and 8x8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and chrominance. Note that quantization refers to digitizing values ​​sampled at predetermined intervals by associating them with predetermined levels, and in this technical field, expressions such as rounding, rounding, and scaling may also be used.

[0112] There are two methods of using a quantization matrix: one is to use a quantization matrix that is directly set on the encoding device side, and the other is to use a default quantization matrix (default matrix). By directly setting a quantization matrix on the encoding device side, it is possible to set a quantization matrix according to the characteristics of an image. However, in this case, there is a disadvantage that the amount of code increases due to the encoding of the quantization matrix.

[0113] On the other hand, there is a method that does not use a quantization matrix and quantizes the coefficients of high-frequency components and low-frequency components in the same way. Note that this method is equivalent to using a quantization matrix in which all coefficients have the same value (a flat matrix).

[0114] The quantization matrix may be specified, for example, in a Sequence Parameter Set (SPS) or a Picture Parameter Set (PPS). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and PPS are sometimes simply referred to as parameter sets.

[0115] [Entropy coding part] The entropy coding unit 110 generates a coded signal (coded bit stream) based on the quantized coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110, for example, binarizes the quantized coefficients, arithmetically codes the binary signal, and outputs a compressed bit stream or sequence.

[0116] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0117] [Inverse conversion section] The inverse transform unit 114 restores the prediction error (residual) by inverse transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0118] Note that the restored prediction error usually loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error usually contains a quantization error.

[0119] [Addition section] The adder 116 reconstructs a current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0120] [Block memory] The block memory 118 is a storage unit for storing, for example, blocks referenced in intra prediction and in a picture to be coded (referred to as a current picture). Specifically, the block memory 118 stores the reconstructed block output from the adder 116.

[0121] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0122] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF or DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0123] In ALF, a least squared error filter is applied to remove coding artifacts. For example, for each 2x2 sub-block in the current block, one filter is selected from among multiple filters based on local gradient direction and activity.

[0124] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes.

[0125] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived, for example, by adding gradients in multiple directions and quantizing the sum.

[0126] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0127] The shape of the filter used in the ALF is, for example, a circularly symmetric shape. FIGS. 6A to 6C are diagrams showing a number of examples of the shape of the filter used in the ALF. FIG. 6A shows a 5×5 diamond-shaped filter, FIG. 6B shows a 7×7 diamond-shaped filter, and FIG. 6C shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Note that the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, and may be at other levels (for example, the sequence level, slice level, tile level, CTU level, or CU level).

[0128] The on / off of ALF may be determined, for example, at the picture level or the CU level. For example, whether or not to apply ALF for luminance may be determined at the CU level, and whether or not to apply ALF for chrominance may be determined at the picture level. Information indicating whether or not to apply ALF is usually signaled at the picture level or the CU level. Note that the signaling of information indicating whether or not to apply ALF is not limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0129] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level, although the signaling of the coefficient sets need not be limited to the picture level, but may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or subblock level).

[0130] [Loop filter section > Deblocking filter] In the deblocking filter, the loop filter unit 120 reduces distortion at block boundaries of the reconstructed image by applying a filtering process to the block boundaries.

[0131] FIG. 7 is a block diagram showing an example of a detailed configuration of the loop filter unit 120 functioning as a deblocking filter.

[0132] The loop filter unit 120 includes a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 and 1206 .

[0133] The boundary determination unit 1201 determines whether or not a pixel to be deblocking-filtered (i.e., a target pixel) exists near a block boundary. Then, the boundary determination unit 1201 outputs the determination result to the switch 1202 and the process determination unit 1208.

[0134] When the boundary determination unit 1201 determines that the target pixel is located near the block boundary, the switch 1202 outputs the image before filtering to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel is not located near the block boundary, the switch 1202 outputs the image before filtering to the switch 1206.

[0135] The filter determination unit 1203 determines whether or not to perform deblocking filter processing on the target pixel based on the pixel value of at least one surrounding pixel around the target pixel. Then, the filter determination unit 1203 outputs the determination result to the switch 1204 and the processing determination unit 1208.

[0136] When the filter determination unit 1203 determines that the deblocking filter process is to be performed on the target pixel, the switch 1204 outputs the unfiltered image acquired via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that the deblocking filter process is not to be performed on the target pixel, the switch 1204 outputs the unfiltered image acquired via the switch 1202 to the switch 1206.

[0137] When the filtering unit 1205 acquires an unfiltered image via the switches 1202 and 1204, it executes deblocking filtering on the target pixel using the filter characteristics determined by the filter characteristics determination unit 1207. Then, the filtering unit 1205 outputs the filtered pixel to the switch 1206.

[0138] The switch 1206 selectively outputs pixels that have not been subjected to the deblocking filter process and pixels that have been subjected to the deblocking filter process by the filter processing unit 1205 under the control of the process determination unit 1208 .

[0139] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filter determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near a block boundary and the filter determination unit 1203 determines that the target pixel is to be subjected to deblocking filter processing, the processing determination unit 1208 causes the switch 1206 to output a pixel that has been subjected to deblocking filter processing. In addition, in cases other than the above, the processing determination unit 1208 causes the switch 1206 to output a pixel that has not been subjected to deblocking filter processing. By repeatedly outputting pixels in this manner, a filtered image is output from the switch 1206.

[0140] FIG. 8 is a diagram showing an example of a deblocking filter having filter characteristics that are symmetric with respect to block boundaries.

[0141] In the deblocking filter process, for example, one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter, is selected using pixel values ​​and a quantization parameter. In the strong filter, when pixels p0 to p2 and pixels q0 to q2 exist on either side of a block boundary as shown in Fig. 8, the pixel values ​​of pixels q0 to q2 are changed to pixel values ​​q'0 to q'2 by performing the calculation shown in the following formula.

[0142] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8 q'1=(p0+q0+q1+q2+2) / 4 q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0143] In the above equations, p0 to p2 and q0 to q2 are the pixel values ​​of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the opposite side of the block boundary. Also, on the right side of each equation, the coefficients by which the pixel values ​​of each pixel used in the deblocking filter process are multiplied are filter coefficients.

[0144] Furthermore, in the deblocking filter process, clipping may be performed so that the pixel value after the calculation does not change beyond a threshold. In this clipping process, the pixel value after the calculation according to the above formula is clipped to "pixel value before the calculation ±2 × threshold" using a threshold determined from the quantization parameter. This makes it possible to prevent excessive smoothing.

[0145] Fig. 9 is a diagram for explaining block boundaries on which deblocking filter processing is performed, and Fig. 10 is a diagram showing an example of a Bs value.

[0146] The block boundary where the deblocking filter process is performed is, for example, the boundary of a PU (Prediction Unit) or TU (Transform Unit) of an 8x8 pixel block as shown in Fig. 9. The deblocking filter process is performed in units of 4 rows or 4 columns. First, the Bs (Boundary Strength) value is determined for block P and block Q shown in Fig. 9 as shown in Fig. 10.

[0147] According to the Bs value in Fig. 10, it is determined whether or not to perform deblocking filter processing of different strengths even for block boundaries belonging to the same image. Deblocking filter processing for the color difference signal is performed when the Bs value is 2. Deblocking filter processing for the luminance signal is performed when the Bs value is 1 or more and a specific condition is satisfied. Note that the conditions for determining the Bs value are not limited to those shown in Fig. 10, and may be determined based on other parameters.

[0148] [Prediction processing unit (intra prediction unit, inter prediction unit, prediction control unit)] 11 is a diagram showing an example of processing performed in the prediction processing unit of the encoding device 100. Note that the prediction processing unit is made up of all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0149] The prediction processing unit generates a prediction image of the current block (step Sb_1). This prediction image is also called a prediction signal or a prediction block. The prediction signal includes, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image that has already been obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0150] The reconstructed image may be, for example, an image of a reference picture or an image of an encoded block in a current picture, which is a picture that includes the current block. The encoded block in the current picture may be, for example, a neighboring block of the current block.

[0151] FIG. 12 is a diagram showing another example of the process performed by the prediction processing unit of the encoding device 100. In FIG.

[0152] The prediction processing unit generates a predicted image by a first method (step Sc_1a), generates a predicted image by a second method (step Sc_1b), and generates a predicted image by a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In these prediction methods, the above-mentioned reconstructed image may be used.

[0153] Next, the prediction processing unit selects one of the multiple predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). This selection of the predicted image, that is, the selection of the method or mode for obtaining the final predicted image, may be performed by calculating a cost for each generated predicted image and based on the cost. Alternatively, the selection of the predicted image may be performed based on parameters used in the encoding process. The encoding device 100 may signal information for identifying the selected predicted image, method, or mode in an encoding signal (also called an encoded bit stream). The information may be, for example, a flag. This allows the decoding device to generate a predicted image according to the method or mode selected in the encoding device 100 based on the information. Note that in the example shown in FIG. 12, the prediction processing unit generates a predicted image in each method and then selects one of the predicted images. However, the prediction processing unit may select a method or mode based on parameters used in the encoding process described above before generating those predicted images, and generate a predicted image according to the method or mode.

[0154] For example, the first and second methods may be intra prediction and inter prediction, respectively, and the prediction processing unit may select a final predicted image for the current block from predicted images generated according to these prediction methods.

[0155] FIG. 13 is a diagram showing another example of the process performed by the prediction processing unit of the encoding device 100. In FIG.

[0156] First, the prediction processing unit generates a predicted image by intra prediction (step Sd_1a), and generates a predicted image by inter prediction (step Sd_1b). Note that the predicted image generated by intra prediction is also called an intra predicted image, and the predicted image generated by inter prediction is also called an inter predicted image.

[0157] Next, the prediction processing unit evaluates each of the intra-predicted image and the inter-predicted image (step Sd_2). This evaluation may use a cost. That is, the prediction processing unit calculates the cost C of each of the intra-predicted image and the inter-predicted image. This cost C is calculated by an RD optimization model formula, for example, C=D+λ×R. In this formula, D is the coding distortion of the predicted image, and is represented by, for example, the sum of absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. Furthermore, R is the generated code amount of the predicted image, and specifically, the code amount required for coding the motion information for generating the predicted image. Furthermore, λ is, for example, Lagrange's undetermined multiplier.

[0158] Then, the prediction processing unit selects the predicted image with the smallest calculated cost C from the intra-predicted image and the inter-predicted image as the final predicted image of the current block (step Sd_3). That is, a prediction method or mode for generating a predicted image of the current block is selected.

[0159] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called intra-screen prediction) of the current block with reference to a block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0160] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes typically includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0161] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.

[0162] The multiple directional prediction modes include, for example, 33 prediction modes defined in the H.265 / HEVC standard. The multiple directional prediction modes may include 32 prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). FIG. 14 is a diagram showing a total of 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions. (The two non-directional prediction modes are not shown in FIG. 14.)

[0163] In various implementation examples, a luminance block may be referenced in intra prediction of a chrominance block. That is, a chrominance component of a current block may be predicted based on a luminance component of the current block. Such intra prediction may be called a cross-component linear model (CCLM) prediction. An intra prediction mode of a chrominance block that references such a luminance block (e.g., called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0164] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on the gradient of reference pixels in the horizontal / vertical directions. Intra prediction with such correction may be called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is usually signaled at the CU level. Note that the signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0165] [Inter prediction section] The inter prediction unit 126 performs inter prediction (also called inter prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 block) in the current block. For example, the inter prediction unit 126 performs motion estimation in the reference picture for the current block or the current sub-block, and finds a reference block or sub-block that most closely matches the current block or the current sub-block. Then, the inter prediction unit 126 obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, and generates an inter prediction signal of the current block or sub-block. The inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0166] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms, for example, a motion vector may be signaled, or, as another example, a difference between a motion vector and a motion vector predictor may be signaled.

[0167] [Basic flow of inter prediction] FIG. 15 is a flowchart showing the basic flow of inter prediction.

[0168] The inter prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0169] Here, in generating a predicted image, the inter prediction unit 126 generates the predicted image by determining a motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). In determining an MV, the inter prediction unit 126 selects a candidate motion vector (candidate MV) (step Se_1) and derives an MV (step Se_2) to determine the MV. The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. In deriving an MV, the inter prediction unit 126 may further select at least one candidate MV from the at least one candidate MV, and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter prediction unit 126 may determine the MV of the current block by searching the area of ​​the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. Note that searching the area of ​​the reference picture may be referred to as motion estimation.

[0170] In the above example, steps Se_1 to Se_3 are performed by the inter prediction unit 126. However, the process of step Se_1 or step Se_2 may be performed by another component included in the encoding device 100.

[0171] [Motion vector derivation flow] FIG. 16 is a flowchart showing an example of motion vector derivation.

[0172] The inter prediction unit 126 derives the MV of the current block in a mode in which motion information (e.g., MV) is coded. In this case, for example, the motion information is coded as a prediction parameter and signaled. That is, the coded motion information is included in a coded signal (also called a coded bitstream).

[0173] Alternatively, the inter prediction unit 126 derives the MV in a mode in which motion information is not coded. In this case, the motion information is not included in the coded signal.

[0174] Here, the MV derivation modes include normal inter mode, merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the modes for encoding motion information include normal inter mode, merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Note that the motion information may include not only MV but also predicted motion vector selection information, which will be described later. Also, the modes for not encoding motion information include FRUC mode. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0175] FIG. 17 is a flowchart showing another example of motion vector derivation.

[0176] The inter prediction unit 126 derives the MV of the current block in a mode of encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0177] Alternatively, the inter prediction unit 126 derives the MV in a mode in which the differential MV is not coded. In this case, the coded differential MV is not included in the coded signal.

[0178] Here, as described above, the modes of deriving an MV include normal inter, merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the modes for encoding a differential MV include normal inter mode and affine mode (specifically, affine inter mode). Also, the modes for not encoding a differential MV include FRUC mode, merge mode, and affine mode (specifically, affine merge mode). The inter prediction unit 126 selects a mode for deriving an MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0179] [Motion vector derivation flow] FIG. 18 is a flowchart showing another example of motion vector derivation. There are a plurality of modes of MV derivation, that is, inter prediction modes, which are roughly divided into a mode in which a differential MV is coded and a mode in which a differential motion vector is not coded. The modes in which a differential MV is not coded include a merge mode, a FRUC mode, and an affine mode (specifically, an affine merge mode). Details of these modes will be described later, but simply, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from a surrounding coded block, and the FRUC mode is a mode in which the MV of the current block is derived by searching between coded regions. In addition, the affine mode is a mode in which the motion vector of each of a plurality of sub-blocks constituting the current block is derived as the MV of the current block, assuming an affine transformation.

[0180] Specifically, when the inter prediction mode information indicates 0 (Sf_1 is 0), the inter prediction unit 126 derives a motion vector using the merge mode (Sf_2). When the inter prediction mode information indicates 1 (Sf_1 is 1), the inter prediction unit 126 derives a motion vector using the FRUC mode (Sf_3). When the inter prediction mode information indicates 2 (Sf_1 is 2), the inter prediction unit 126 derives a motion vector using the affine mode (specifically, the affine merge mode) (Sf_4). When the inter prediction mode information indicates 3 (Sf_1 is 3), the inter prediction unit 126 derives a motion vector using a mode for encoding a differential MV (for example, normal inter mode) (Sf_5).

[0181] [MV Derivation > Normal Inter Mode] The normal inter mode is an inter prediction mode in which the MV of the current block is derived by finding a block similar to the image of the current block from the region of the reference picture indicated by the candidate MV. In addition, in this normal inter mode, the differential MV is coded.

[0182] FIG. 19 is a flowchart showing an example of inter prediction in the normal inter mode.

[0183] The inter prediction unit 126 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks around the current block in time or space (step Sg_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0184] Next, the inter prediction unit 126 extracts N candidate MVs (N is an integer equal to or greater than 2) from the multiple candidate MVs acquired in step Sg_1 as motion vector predictor candidates (also called predicted MV candidates) according to a predetermined priority order (step Sg_2). Note that the priority order is predetermined for each of the N candidate MVs.

[0185] Next, the inter prediction unit 126 selects one of the N motion vector predictor candidates as a motion vector predictor (also called a prediction MV) of the current block (step Sg_3). At this time, the inter prediction unit 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor into a stream. Note that the stream is the above-mentioned encoded signal or encoded bit stream.

[0186] Next, the inter prediction unit 126 derives the MV of the current block by referring to the coded reference picture (step Sg_4). At this time, the inter prediction unit 126 further encodes the difference value between the derived MV and the predicted motion vector as a differential MV into a stream. Note that the coded reference picture is a picture consisting of a plurality of blocks reconstructed after coding.

[0187] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Sg_5). Note that the predicted image is the above-mentioned inter prediction signal.

[0188] Furthermore, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the predicted image, which is included in the coded signal, is coded as, for example, a prediction parameter.

[0189] The candidate MV list may be used in common with lists used in other modes. Furthermore, processing related to the candidate MV list may be applied to processing related to lists used in other modes. Processing related to this candidate MV list may include, for example, extraction or selection of candidate MVs from the candidate MV list, sorting of the candidate MVs, or deletion of candidate MVs.

[0190] [MV Derivation > Merge Mode] Merge mode is an inter prediction mode that derives the MV of the current block by selecting a candidate MV from a candidate MV list as the MV for that block.

[0191] FIG. 20 is a flowchart showing an example of inter prediction in the merge mode.

[0192] The inter prediction unit 126 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks around the current block in time or space (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0193] Next, the inter prediction unit 126 derives the MV of the current block by selecting one candidate MV from the multiple candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream.

[0194] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Sh_3).

[0195] Furthermore, information indicating the inter prediction mode (merge mode in the above example) used to generate the predicted image, which is included in the coded signal, is coded as, for example, a prediction parameter.

[0196] FIG. 21 is a diagram illustrating an example of a motion vector derivation process for a current picture in the merge mode.

[0197] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include a spatially adjacent prediction MV, which is an MV held by a plurality of coded blocks located spatially around the target block, a temporally adjacent prediction MV, which is an MV held by a nearby block projected onto the position of the target block in a coded reference picture, a joint prediction MV, which is an MV generated by combining the MV values ​​of the spatially adjacent prediction MV and the temporally adjacent prediction MV, and a zero prediction MV, which is an MV with a value of zero.

[0198] Next, one predicted MV is selected from the multiple predicted MVs registered in the predicted MV list, and is determined as the MV for the target block.

[0199] Furthermore, the variable length coding unit writes merge_idx, which is a signal indicating which predicted MV has been selected, into the stream and codes it.

[0200] Note that the predicted MVs registered in the predicted MV list described in Figure 21 are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include additional predicted MVs other than the types of predicted MVs shown in the figure.

[0201] The final MV may be determined by performing dynamic motion vector refreshing (DMVR) processing, which will be described later, using the MV of the target block derived in the merge mode.

[0202] The prediction MV candidates are the above-mentioned candidate MVs, and the prediction MV list is the above-mentioned candidate MV list. The candidate MV list may also be called a candidate list. merge_idx is MV selection information.

[0203] [MV derivation > FRUC mode] The motion information may be derived on the decoding device side without being signaled from the encoding device side. As described above, the merge mode defined in the H.265 / HEVC standard may be used. For example, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation is performed on the decoding device side without using pixel values ​​of the current block.

[0204] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.

[0205] An example of the FRUC process is shown in FIG. 22. First, a list of multiple candidates (i.e., a candidate MV list, which may be common to the merge list) each having a predicted motion vector (MV) is generated by referring to the motion vector of the coded block spatially or temporally adjacent to the current block (step Si_1). Next, a best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate MV is selected based on the evaluation value. Then, a motion vector for the current block is derived based on the motion vector of the selected candidate (step Si_4). Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as it is as the motion vector for the current block. Also, for example, the motion vector for the current block may be derived by performing pattern matching in a peripheral area of ​​a position in a reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed on the area around the best candidate MV using pattern matching and evaluation values ​​in the reference picture, and if an MV with a better evaluation value is found, the best candidate MV is updated to that MV and used as the final MV for the current block. It is also possible to configure the system without performing the process of updating to an MV with a better evaluation value.

[0206] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Si_5).

[0207] The same processing may be performed when processing is performed in sub-block units.

[0208] The evaluation value may be calculated by various methods. For example, a reconstructed image of an area in a reference picture corresponding to the motion vector is compared with a reconstructed image of a predetermined area (which may be, for example, an area of ​​another reference picture or an area of ​​an adjacent block of the current picture, as shown below). Then, the difference between pixel values ​​of the two reconstructed images may be calculated and used as the evaluation value of the motion vector. Note that the evaluation value may be calculated using other information in addition to the difference value.

[0209] Next, pattern matching will be described in detail. First, one candidate MV included in a candidate MV list (for example, a merge list) is selected as a start point of a search by pattern matching. As the pattern matching, a first pattern matching or a second pattern matching is used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0210] [MV derivation > FRUC > Bilateral matching] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0211] FIG. 23 is a diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in FIG. 23, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of blocks that best match among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of a current block (Cur block). Specifically, for the current block, a difference is derived between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by a candidate MV and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV by a display time interval, and an evaluation value is calculated using the obtained difference value. It is preferable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.

[0212] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures in time and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives bidirectional motion vectors that are mirror-symmetric.

[0213] [MV derivation > FRUC > Template matching] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0214] Fig. 24 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in Fig. 24, in the second pattern matching, a motion vector of a current block is derived by searching in a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in a current picture (Cur Pic). Specifically, for a current block, a difference is derived between a reconstructed image of both or either of the left adjacent and upper adjacent coded areas and a reconstructed image at the same position in a coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and a candidate MV with the best evaluation value among a plurality of candidate MVs is selected as a best candidate MV.

[0215] Information indicating whether such a FRUC mode is applied (e.g., called a FRUC flag) may be signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating an applicable pattern matching method (first pattern matching or second pattern matching) may be signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0216] [MV derivation > Affine mode] Next, a description will be given of an affine mode in which a motion vector is derived for each sub-block based on the motion vectors of a plurality of adjacent blocks. This mode is sometimes called an affine motion compensation prediction mode.

[0217] FIG. 25A is a diagram for explaining an example of derivation of a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 25A, the current block includes 16 4x4 sub-blocks. Here, a motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and similarly, a motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, the two motion vectors v0 and v1 are projected by the following equation (1A) to obtain the motion vectors (v x ,v y ) is derived.

[0218]

number

[0219] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting coefficient.

[0220] Such information indicating the affine mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating the affine mode does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0221] In addition, such an affine mode may include several modes that differ in the method of deriving the motion vectors of the upper-left and upper-right corner control points. For example, the affine mode includes two modes: an affine inter (also called an affine normal inter) mode and an affine merge mode.

[0222] [MV derivation > Affine mode] FIG. 25B is a diagram for explaining an example of derivation of motion vectors on a sub-block basis in an affine mode having three control points. In FIG. 25B, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vector of an adjacent block, and similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vector of the adjacent block, and the motion vector v2 of the lower left corner control point of the current block is derived based on the motion vector of the adjacent block. Then, the three motion vectors v0, v1, and v2 are projected according to the following equation (1B) to derive the motion vectors (v x ,v y ) is derived.

[0223]

number

[0224] Here, x and y respectively indicate the horizontal and vertical positions of the subblock center, w indicates the width of the current block, and h indicates the height of the current block.

[0225] Affine modes with different numbers of control points (e.g., two and three) may be switched and signaled at the CU level. Note that information indicating the number of control points of the affine mode used at the CU level may also be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0226] In addition, such an affine mode having three control points may include several modes with different methods of deriving the motion vectors of the upper left, upper right, and lower left corner control points. For example, the affine mode includes two modes: affine inter (also called affine normal inter) mode and affine merge mode.

[0227] [MV Derivation > Affine Merge Mode] 26A, 26B, and 26C are conceptual diagrams for explaining the affine merge mode.

[0228] In the affine merge mode, as shown in FIG. 26A, for example, among the coded blocks A (left), B (top), C (top right), D (bottom left) and E (top left) adjacent to the current block, the predicted motion vectors of each of the control points of the current block are calculated based on a plurality of motion vectors corresponding to the blocks coded in affine mode. Specifically, these blocks are inspected in the order of coded blocks A (left), B (top), C (top right), D (bottom left) and E (top left), and the first valid block coded in affine mode is identified. Based on a plurality of motion vectors corresponding to this identified block, the predicted motion vector of the control point of the current block is calculated.

[0229] For example, as shown in Figure 26B, when the block A adjacent to the left of the current block is coded in affine mode with two control points, motion vectors v3 and v4 are derived by projecting the upper left corner and upper right corner positions of the coded block including block A. Then, from the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner are calculated.

[0230] For example, as shown in Figure 26C, when the block A adjacent to the left of the current block is coded in affine mode with three control points, motion vectors v3, v4 and v5 are derived by projecting to the upper left corner, upper right corner and lower left corner positions of the coded block including block A. Then, from the derived motion vectors v3, v4 and v5, the predicted motion vector v0 of the control point of the upper left corner of the current block, the predicted motion vector v1 of the control point of the upper right corner and the predicted motion vector v2 of the control point of the lower left corner are calculated.

[0231] Note that this predicted motion vector derivation method may be used to derive predicted motion vectors for each control point of the current block in step Sj_1 of FIG. 29, which will be described later.

[0232] FIG. 27 is a flow chart illustrating an example of the affine merge mode.

[0233] In the affine merge mode, first, the inter prediction unit 126 derives prediction MVs for each of the control points of the current block (step Sk_1). The control points are the upper left and upper right corners of the current block as shown in Fig. 25A, or the upper left, upper right and lower left corners of the current block as shown in Fig. 25B.

[0234] That is, the inter prediction unit 126 examines the coded blocks in the following order, as shown in FIG. 26A: coded block A (left), block B (top), block C (top right), block D (bottom left) and block E (top left), and identifies the first valid block coded in affine mode.

[0235] Then, when block A is identified and has two control points, as shown in FIG. 26B, the inter prediction unit 126 calculates the motion vector v0 of the control point of the upper left corner of the current block and the motion vector v1 of the control point of the upper right corner from the motion vectors v3 and v4 of the upper left corner and upper right corner of the coded block including block A. For example, the inter prediction unit 126 calculates the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner by projecting the motion vectors v3 and v4 of the upper left corner and upper right corner of the coded block onto the current block.

[0236] Alternatively, when block A is specified and block A has three control points, as shown in FIG. 26C, the inter prediction unit 126 calculates the motion vector v0 of the control point of the upper left corner of the current block, the motion vector v1 of the control point of the upper right corner, and the motion vector v2 of the control point of the lower left corner from the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, the inter prediction unit 126 calculates the predicted motion vector v0 of the control point of the upper left corner of the current block, the predicted motion vector v1 of the control point of the upper right corner, and the motion vector v2 of the control point of the lower left corner by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block.

[0237] Next, the inter prediction unit 126 performs motion compensation for each of the sub-blocks included in the current block. That is, for each of the sub-blocks, the inter prediction unit 126 calculates the motion vector of the sub-block as an affine MV using two predicted motion vectors v0 and v1 and the above formula (1A), or three predicted motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter prediction unit 126 performs motion compensation for the sub-block using the affine MVs and the coded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0238] [MV Derivation > Affine Intermode] FIG. 28A is a diagram for explaining an affine inter mode having two control points.

[0239] In this affine inter mode, as shown in Figure 28A, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as a predicted motion vector v0 of the control point of the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as a predicted motion vector v1 of the control point of the upper right corner of the current block.

[0240] FIG. 28B is a diagram for explaining an affine inter mode having three control points.

[0241] In this affine inter mode, as shown in FIG. 28B, a motion vector selected from the motion vectors of the coded blocks A, B and C adjacent to the current block is used as the predicted motion vector v0 of the control point of the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point of the upper right corner of the current block. Furthermore, a motion vector selected from the motion vectors of the coded blocks F and G adjacent to the current block is used as the predicted motion vector v2 of the control point of the lower left corner of the current block.

[0242] FIG. 29 is a flowchart showing an example of the affine inter mode.

[0243] In the affine inter mode, the inter prediction unit 126 first derives predicted MVs (v0, v1) or (v0, v1, v2) of two or three control points of the current block (step Sj_1). The control points are the upper left corner, upper right corner, or lower left corner of the current block, as shown in FIG. 25A or FIG. 25B.

[0244] That is, the inter prediction unit 126 derives the predicted motion vector (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the motion vector of any one of the coded blocks in the vicinity of each control point of the current block shown in Figure 28A or Figure 28B. At this time, the inter prediction unit 126 codes the predicted motion vector selection information for identifying the two selected motion vectors into a stream.

[0245] For example, the inter prediction unit 126 may use cost evaluation or the like to determine which motion vector of an encoded block adjacent to the current block to select as the predicted motion vector of the control point, and may write a flag indicating which predicted motion vector has been selected in the bitstream.

[0246] Next, the inter prediction unit 126 performs motion search (steps Sj_3 and Sj_4) while updating each of the predicted motion vectors selected or derived in step Sj_1 (step Sj_2). That is, the inter prediction unit 126 calculates the motion vector of each sub-block corresponding to the predicted motion vector to be updated as an affine MV using the above-mentioned formula (1A) or formula (1B) (step Sj_3). Then, the inter prediction unit 126 performs motion compensation for each sub-block using the affine MVs and the coded reference picture (step Sj_4). As a result, the inter prediction unit 126 determines, in the motion search loop, for example, the predicted motion vector that provides the smallest cost as the motion vector of the control point (step Sj_5). At this time, the inter prediction unit 126 further codes the difference value between the determined MV and the predicted motion vector as a differential MV into a stream.

[0247] Finally, the inter prediction unit 126 performs motion compensation on the current block using the determined MV and the encoded reference picture to generate a predicted image of the current block (step Sj_6).

[0248] [MV Derivation > Affine Intermode] When affine modes with different numbers of control points (for example, two and three) are switched and signaled at the CU level, the number of control points may differ between the coded block and the current block. Figures 30A and 30B are conceptual diagrams for explaining a method of deriving a predicted vector of a control point when the number of control points differs between the coded block and the current block.

[0249] For example, as shown in FIG. 30A, when the current block has three control points, which are the upper left corner, the upper right corner and the lower left corner, and the block A adjacent to the left of the current block is coded in affine mode with two control points, motion vectors v3 and v4 are derived that are projected to the upper left corner and the upper right corner of the coded block including block A.Then, from the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner are calculated.Furthermore, from the derived motion vectors v0 and v1, the predicted motion vector v2 of the control point of the lower left corner is calculated.

[0250] For example, as shown in Figure 30B, if the current block has two control points at the upper left corner and the upper right corner, and the block A adjacent to the left of the current block is coded in affine mode with three control points, motion vectors v3, v4 and v5 are derived by projecting the upper left corner, upper right corner and lower left corner positions of the coded block including block A. Then, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated from the derived motion vectors v3, v4 and v5.

[0251] This prediction motion vector derivation method may be used to derive the prediction motion vector for each control point of the current block in step Sj_1 of FIG.

[0252] [MV Derivation > DMVR] FIG. 31A is a diagram showing the relationship between the merge mode and the DMVR.

[0253] The inter prediction unit 126 derives a motion vector of the current block in merge mode (step Sl_1). Next, the inter prediction unit 126 determines whether or not to search for a motion vector, that is, to perform motion search (step Sl_2). Here, when the inter prediction unit 126 determines not to perform motion search (No in step Sl_2), it determines the motion vector derived in step Sl_1 as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is determined in merge mode.

[0254] On the other hand, when it is determined in step Sl_1 that motion search is to be performed (Yes in step Sl_2), the inter prediction unit 126 derives a final motion vector for the current block by searching the surrounding area of ​​the reference picture indicated by the motion vector derived in step Sl_1 (step Sl_3). That is, in this case, the motion vector of the current block is determined by DMVR.

[0255] FIG. 31B is a conceptual diagram for explaining an example of DMVR processing for determining an MV.

[0256] First, the optimal MVP set for the current block (e.g., in merge mode) is set as the candidate MV. Then, according to the candidate MV(L0), reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 direction. Similarly, according to the candidate MV(L1), reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 direction. A template is generated by averaging these reference pixels.

[0257] Next, the template is used to search the surrounding areas of the candidate MVs in the first reference picture (L0) and the second reference picture (L1), and the MV with the smallest cost is determined as the final MV. Note that the cost value may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search area, the candidate MV value, etc.

[0258] The encoding device and a decoding device (to be described later) basically have the same processing configuration and operations as those described here.

[0259] Any process may be used other than the process described here as long as it is capable of searching the vicinity of the candidate MV and deriving the final MV.

[0260] [Motion compensation > BIO / OBMC] In motion compensation, there is a mode in which a predicted image is generated and the predicted image is corrected, such as BIO and OBMC, which will be described later.

[0261] FIG. 32 is a flowchart showing an example of generation of a predicted image.

[0262] The inter prediction unit 126 generates a predicted image (step Sm_1), and corrects the predicted image using any of the above-mentioned modes (step Sm_2).

[0263] FIG. 33 is a flowchart showing another example of generation of a predicted image.

[0264] The inter prediction unit 126 determines a motion vector of the current block (step Sn_1). Next, the inter prediction unit 126 generates a predicted image (step Sn_2) and determines whether or not to perform correction processing (step Sn_3). Here, if the inter prediction unit 126 determines that correction processing is to be performed (Yes in step Sn_3), it generates a final predicted image by correcting the predicted image (step Sn_4). On the other hand, if the inter prediction unit 126 determines that correction processing is not to be performed (No in step Sn_3), it outputs the predicted image as the final predicted image without correcting it (step Sn_5).

[0265] Furthermore, motion compensation has a mode in which luminance is corrected when generating a predicted image, such as LIC, which will be described later.

[0266] FIG. 34 is a flowchart showing yet another example of generation of a predicted image.

[0267] The inter prediction unit 126 derives a motion vector of the current block (step So_1). Next, the inter prediction unit 126 determines whether or not to perform luminance correction processing (step So_2). Here, if the inter prediction unit 126 determines to perform luminance correction processing (Yes in step So_2), it generates a predicted image while performing luminance correction (step So_3). That is, the predicted image is generated by LIC. On the other hand, if the inter prediction unit 126 determines not to perform luminance correction processing (No in step So_2), it generates a predicted image by normal motion compensation without performing luminance correction (step So_4).

[0268] [Motion Compensation > OBMC] An inter prediction signal may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks. Specifically, a prediction signal based on the motion information obtained by motion search (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture) may be weighted and added to generate an inter prediction signal for each sub-block in the current block. Such inter prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0269] In the OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is applied (e.g., called OBMC flag) may be signaled at the CU level. Note that the signaling level of these pieces of information does not need to be limited to the sequence level and the CU level, and may be other levels (e.g., the picture level, slice level, tile level, CTU level, or sub-block level).

[0270] The OBMC mode will now be described in more detail. Figures 35 and 36 are a flowchart and a conceptual diagram for explaining an overview of the predictive image correction process in the OBMC process.

[0271] First, as shown in Fig. 36, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to a processing target (current) block. In Fig. 36, the arrow "MV" indicates a reference picture, and indicates what the current block of the current picture refers to in order to obtain a predicted image.

[0272] Next, the motion vector (MV_L) already derived for the coded left adjacent block is applied (reused) to the block to be coded to obtain a predicted image (Pred_L). The motion vector (MV_L) is indicated by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by superimposing the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between the adjacent blocks.

[0273] Similarly, a motion vector (MV_U) already derived for the coded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. The predicted image Pred_U is then superimposed on the predicted images (e.g., Pred and Pred_L) that have been corrected the first time, thereby performing a second correction of the predicted image. This has the effect of blending the boundaries between the adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block, with the boundaries with the adjacent blocks blended (smoothed).

[0274] Note that the above example is a two-pass correction method using the left adjacent and above adjacent blocks, but the correction method may also be a three-pass or more pass correction method using the right adjacent and / or below adjacent blocks.

[0275] The area in which overlapping is performed does not have to be the entire pixel area of ​​the block, but may be only a part of the area near the block boundary.

[0276] Here, the OBMC predicted image correction process for obtaining one predicted image Pred by superimposing additional predicted images Pred_L and Pred_U from one reference picture has been described. However, when the predicted image is corrected based on multiple reference pictures, the same process may be applied to each of the multiple reference pictures. In such a case, the OBMC image correction based on multiple reference pictures is performed to obtain a corrected predicted image from each reference picture, and then the obtained multiple corrected predicted images are further superimposed to obtain a final predicted image.

[0277] In addition, in OBMC, the unit of the target block may be a prediction block unit, or a sub-block unit obtained by further dividing the prediction block.

[0278] As a method of determining whether or not to apply OBMC processing, for example, there is a method of using obmc_flag, which is a signal indicating whether or not to apply OBMC processing. As a specific example, the encoding device may determine whether or not the target block belongs to an area with complex motion. If the target block belongs to an area with complex motion, the encoding device sets a value of 1 as obmc_flag and applies OBMC processing to perform encoding, and if the target block does not belong to an area with complex motion, the encoding device sets a value of 0 as obmc_flag and performs encoding of the block without applying OBMC processing. On the other hand, the decoding device decodes obmc_flag described in a stream (e.g., a compressed sequence), and switches whether or not to apply OBMC processing depending on the value to perform decoding.

[0279] In the above example, the inter prediction unit 126 generates one rectangular predicted image for the rectangular current block. However, the inter prediction unit 126 may generate multiple predicted images of shapes other than a rectangle for the rectangular current block, and combine the multiple predicted images to generate a final rectangular predicted image. The shape other than a rectangle may be, for example, a triangle.

[0280] FIG. 37 is a diagram for explaining generation of predicted images of two triangles.

[0281] The inter prediction unit 126 generates a predicted image of a triangle by performing motion compensation on a first partition of a triangle in the current block using a first MV of the first partition. Similarly, the inter prediction unit 126 generates a predicted image of a triangle by performing motion compensation on a second partition of a triangle in the current block using a second MV of the second partition. Then, the inter prediction unit 126 generates a predicted image of the same rectangle as the current block by combining these predicted images.

[0282] In the example shown in Fig. 37, the first partition and the second partition are each triangular, but they may be trapezoidal or may have different shapes. Furthermore, in the example shown in Fig. 37, the current block is composed of two partitions, but it may be composed of three or more partitions.

[0283] Also, the first partition and the second partition may overlap each other, i.e., the first partition and the second partition may include the same pixel area. In this case, the predicted image of the current block may be generated using the predicted image of the first partition and the predicted image of the second partition.

[0284] Furthermore, in this example, a predicted image is generated by inter prediction for both of the two partitions, but a predicted image may be generated by intra prediction for at least one partition.

[0285] [Motion compensation > BIO] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called a BIO (bi-directional optical flow) mode.

[0286] Fig. 38 is a diagram for explaining a model assuming uniform linear motion. In Fig. 38, (vx, vy) indicates a velocity vector, and τ0 and τ1 indicate the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) indicates a motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates a motion vector corresponding to the reference picture Ref1.

[0287] In this case, under the assumption of uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (2) holds.

[0288]

number

[0289] Here, I(k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like may be corrected pixel by pixel.

[0290] Note that the decoding device may derive a motion vector using a method other than the method based on a model assuming uniform linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0291] [Motion Compensation > LIC] Next, an example of a mode in which a predicted image (prediction) is generated using LIC (local illumination compensation) processing will be described.

[0292] FIG. 39 is a diagram for explaining an example of a predicted image generating method using a luminance correction process by LIC processing.

[0293] First, the MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.

[0294] Next, extract information indicating how the luminance value of the current block has changed between the reference picture and the current picture. This extraction is performed based on the luminance pixel values ​​of the coded left adjacent reference area (peripheral reference area) and the coded upper adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the derived MV. Then, calculate a luminance correction parameter using the information indicating how the luminance value has changed.

[0295] A luminance correction process is performed by applying the luminance correction parameters to a reference image in a reference picture specified by the MV, thereby generating a predicted image for the current block.

[0296] It should be noted that the shape of the surrounding reference region in FIG. 39 is just an example, and other shapes may be used.

[0297] Although the process of generating a predicted image from one reference picture has been described here, the process is similar when generating a predicted image from multiple reference pictures, and a luminance correction process may be performed on the reference images obtained from each reference picture in a manner similar to that described above before generating a predicted image.

[0298] As a method of determining whether or not to apply LIC processing, for example, there is a method of using lic_flag, which is a signal indicating whether or not to apply LIC processing. As a specific example, in an encoding device, it is determined whether or not the current block belongs to an area where a luminance change occurs, and if it belongs to an area where a luminance change occurs, a value of 1 is set as lic_flag and LIC processing is applied and encoding is performed, and if it does not belong to an area where a luminance change occurs, a value of 0 is set as lic_flag and encoding is performed without applying LIC processing. On the other hand, a decoding device may decode lic_flag described in a stream, and switch whether or not to apply LIC processing depending on the value and perform decoding.

[0299] Another method of determining whether to apply LIC processing is, for example, a method of determining according to whether LIC processing has been applied to surrounding blocks.As a specific example, when the current block is in merge mode, it is determined whether the surrounding coded blocks selected when deriving MV in merge mode processing have been coded by applying LIC processing.Depending on the result, it is switched to whether to apply LIC processing and performs coding.In this example, the same processing is also applied to the processing on the decoding device side.

[0300] The LIC process (luminance correction process) has been described with reference to FIG. 39, and will be described in detail below.

[0301] First, the inter prediction unit 126 derives a motion vector for obtaining a reference image corresponding to the current block to be coded from a reference picture that is a coded picture.

[0302] Next, the inter prediction unit 126 uses the luminance pixel values ​​of the coded surrounding reference areas adjacent to the left and above the coding target block and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the motion vector to extract information indicating how the luminance values ​​have changed between the reference picture and the coding target picture, and calculates a luminance correction parameter. For example, the luminance pixel value of a pixel in the surrounding reference area in the coding target picture is set to p0, and the luminance pixel value of a pixel in the surrounding reference area in the reference picture at the equivalent position to the pixel is set to p1. The inter prediction unit 126 calculates coefficients A and B that optimize A×p1+B=p0 for multiple pixels in the surrounding reference areas as the luminance correction parameter.

[0303] Next, the inter prediction unit 126 performs luminance correction processing on a reference image in a reference picture specified by the motion vector using the luminance correction parameter to generate a predicted image for the block to be coded. For example, the luminance pixel value in the reference image is p2, and the luminance pixel value of the predicted image after the luminance correction processing is p3. The inter prediction unit 126 calculates A×p2+B=p3 for each pixel in the reference image to generate a predicted image after the luminance correction processing.

[0304] Note that the shape of the surrounding reference area in FIG. 39 is an example, and other shapes may be used. Also, a part of the surrounding reference area shown in FIG. 39 may be used. For example, an area including a predetermined number of pixels thinned out from each of the upper adjacent pixel and the left adjacent pixel may be used as the surrounding reference area. Also, the surrounding reference area is not limited to an area adjacent to the encoding target block, and may be an area not adjacent to the encoding target block. Also, in the example shown in FIG. 39, the surrounding reference area in the reference picture is an area specified by the motion vector of the encoding target picture from the surrounding reference area in the encoding target picture, but may be an area specified by another motion vector. For example, the other motion vector may be the motion vector of the surrounding reference area in the encoding target picture.

[0305] Although the operation of encoding device 100 has been described above, the operation of decoding device 200 is also similar.

[0306] The LIC process may be applied to color difference instead of luminance. In this case, correction parameters may be derived separately for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.

[0307] Alternatively, the LIC process may be applied on a subblock basis. For example, the correction parameters may be derived using a surrounding reference region of the current subblock and a surrounding reference region of a reference subblock in a reference picture specified by the MV of the current subblock.

[0308] [Predictive control unit] The prediction control unit 128 selects either an intra-prediction signal (a signal output from the intra-prediction unit 124) or an inter-prediction signal (a signal output from the inter-prediction unit 126), and outputs the selected signal to the subtraction unit 104 and the addition unit 116 as a prediction signal.

[0309] As shown in FIG. 1, in various implementations, the prediction control unit 128 may output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bitstream and perform the same prediction process as that performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used in the intra predictor 124 or the inter predictor 126), or any index, flag, or value based on or indicating the prediction process performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128.

[0310] [Example of an encoding device implementation] 40 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, several components of the encoding device 100 shown in FIG. 1 are implemented by the processor a1 and the memory a2 shown in FIG.

[0311] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. The processor a1 may be a processor such as a CPU. The processor a1 may also be a collection of multiple electronic circuits. For example, the processor a1 may play the roles of multiple components of the encoding device 100 shown in FIG. 1 and the like, excluding components for storing information.

[0312] The memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode a moving image is stored. The memory a2 may be an electronic circuit and may be connected to the processor a1. The memory a2 may be included in the processor a1. The memory a2 may be a collection of multiple electronic circuits. The memory a2 may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium, etc. The memory a2 may be a non-volatile memory or a volatile memory.

[0313] For example, the memory a2 may store a video to be encoded, or a bit string corresponding to the encoded video, or may store a program for the processor a1 to encode the video.

[0314] Also, for example, the memory a2 may play the role of a component for storing information among the multiple components of the encoding device 100 shown in Fig. 1 etc. Specifically, the memory a2 may play the role of the block memory 118 and the frame memory 122 shown in Fig. 1. More specifically, the memory a2 may store reconstructed blocks, reconstructed pictures, etc.

[0315] It should be noted that not all of the components shown in Fig. 1 and the like may be implemented, and not all of the processes described above may be performed, in the encoding device 100. Some of the components shown in Fig. 1 and the like may be included in another device, and some of the processes described above may be executed by another device.

[0316] [Decryption device] Next, a description will be given of a decoding device capable of decoding the coded signal (coded bit stream) output from the above coding device 100. Fig. 41 is a block diagram showing a functional configuration of a decoding device 200 according to this embodiment. The decoding device 200 is a video decoding device that decodes a video on a block-by-block basis.

[0317] As shown in FIG. 41, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0318] The decoding device 200 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The decoding device 200 may also be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0319] Below, the overall processing flow of the decoding device 200 will be described, and then each component included in the decoding device 200 will be described.

[0320] [Overall flow of decryption process] FIG. 42 is a flowchart showing an example of the overall decoding process by the decoding device 200.

[0321] First, the entropy decoding unit 202 of the decoding device 200 identifies a division pattern of fixed-size blocks (128×128 pixels) (step Sp_1). This division pattern is the division pattern selected by the encoding device 100. Then, the decoding device 200 performs the processes of steps Sp_2 to Sp_6 on each of the multiple blocks constituting the division pattern.

[0322] That is, the entropy decoding unit 202 decodes (specifically, entropy decodes) the coded quantized coefficients and prediction parameters of the block to be decoded (also called the current block) (step Sp_2).

[0323] Next, the inverse quantization unit 204 and the inverse transform unit 206 perform inverse quantization and inverse transform on the multiple quantized coefficients to reconstruct multiple prediction residuals (that is, difference blocks) (step Sp_3).

[0324] Next, a prediction processing unit consisting of all or a part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction signal (also called a prediction block) of the current block (step Sp_4).

[0325] Next, the adder 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted block to the difference block (step Sp_5).

[0326] Then, when this reconstructed image is generated, the loop filter unit 212 performs filtering on the reconstructed image (step Sp_6).

[0327] Then, the decoding device 200 determines whether or not the decoding of the entire picture is completed (step Sp_7), and if it determines that the decoding is not completed (No in step Sp_7), it repeats the process from step Sp_1.

[0328] The processes of steps Sp_1 to Sp_7 may be performed sequentially by the decoding device 200, or some of the processes may be performed in parallel, or the order of the processes may be changed.

[0329] [Entropy Decoding Part] The entropy decoding unit 202 entropy decodes the coded bitstream. Specifically, for example, the entropy decoding unit 202 arithmetically decodes the coded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. The entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 on a block-by-block basis. The entropy decoding unit 202 may output prediction parameters included in the coded bitstream (see FIG. 1) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can execute the same prediction process as the process executed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the coding device side.

[0330] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of a block to be decoded (hereinafter, referred to as a current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0331] [Inverse conversion section] The inverse transform unit 206 restores the prediction error by inverse transforming the transform coefficients input from the inverse quantization unit 204 .

[0332] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0333] Also for example, if the information interpreted from the coded bitstream indicates to apply NSST, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0334] [Addition section] The adder 208 reconstructs the current block by adding the prediction error, which is an input from the inverse transformer 206, and the prediction sample, which is an input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0335] [Block memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter, referred to as a current picture). Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0336] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0337] If the information indicating ALF on / off read from the encoded bitstream indicates ALF on, one filter is selected from among multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0338] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0339] [Prediction processing unit (intra prediction unit, inter prediction unit, prediction control unit)] 43 is a diagram showing an example of processing performed in the prediction processing unit of the decoding device 200. Note that the prediction processing unit is made up of all or some of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0340] The prediction processing unit generates a prediction image of the current block (step Sq_1). This prediction image is also called a prediction signal or a prediction block. The prediction signal includes, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image that has already been obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0341] The reconstructed image may be, for example, an image of a reference picture, or an image of a decoded block in a current picture, which is a picture that includes the current block. The decoded block in the current picture may be, for example, a neighboring block of the current block.

[0342] FIG. 44 is a diagram showing another example of the process performed by the prediction processing unit of the decoding device 200. In FIG.

[0343] The prediction processing unit determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode may be determined based on prediction parameters, for example.

[0344] When the prediction processing unit determines the first method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the first method (step Sr_2a). When the prediction processing unit determines the second method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the second method (step Sr_2b). When the prediction processing unit determines the third method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the third method (step Sr_2c).

[0345] The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter-prediction method, an intra-prediction method, and other prediction methods. These prediction methods may use the above-mentioned reconstructed image.

[0346] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on an intra prediction mode interpreted from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0347] Note that, when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0348] Furthermore, when information interpreted from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects pixel values ​​after intra prediction based on the gradients of reference pixels in the horizontal / vertical directions.

[0349] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) in the current block. For example, the inter prediction unit 218 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the encoded bit stream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.

[0350] If the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks.

[0351] Also, if the information interpreted from the encoded bitstream indicates that the FRUC mode is applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the encoded bitstream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0352] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when information interpreted from the encoded bitstream indicates that an affine motion compensation prediction mode is applied, the inter prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0353] [MV Derivation > Normal Inter Mode] If the information interpreted from the encoded bitstream indicates that normal inter mode is to be applied, the inter prediction unit 218 derives an MV based on the information interpreted from the encoded bitstream, and performs motion compensation (prediction) using the MV.

[0354] FIG. 45 is a flowchart showing an example of inter prediction in the normal inter mode in the decoding device 200.

[0355] The inter prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded blocks around the current block in time or space (step Ss_1). That is, the inter prediction unit 218 creates a candidate MV list.

[0356] Next, the inter prediction unit 218 extracts N candidate MVs (N is an integer equal to or greater than 2) from the multiple candidate MVs acquired in step Ss_1 as motion vector predictor candidates (also called prediction MV candidates) according to a predetermined priority order (step Ss_2). Note that the priority order is predetermined for each of the N prediction MV candidates.

[0357] Next, the inter prediction unit 218 decodes the predicted motion vector selection information from the input stream (i.e., the encoded bit stream), and uses the decoded predicted motion vector selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted motion vector (also called predicted MV) of the current block (step Ss_3).

[0358] Next, the inter prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the differential value, which is the decoded differential MV, to the selected predicted motion vector (step Ss_4).

[0359] Finally, the inter prediction unit 218 performs motion compensation on the current block using the derived MV and the decoded reference picture to generate a predicted image of the current block (step Ss_5).

[0360] [Predictive control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as a prediction signal to the adder unit 208. Overall, the configurations, functions, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side may correspond to the configurations, functions, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.

[0361] [Example of implementation of a decryption device] Fig. 46 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, several components of the decoding device 200 shown in Fig. 41 are implemented by the processor b1 and the memory b2 shown in Fig. 46.

[0362] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes encoded video (i.e., an encoded bitstream). The processor b1 may be a processor such as a CPU. The processor b1 may also be a collection of multiple electronic circuits. For example, the processor b1 may play the role of multiple components of the decoding device 200 shown in FIG. 41 and the like, excluding the components for storing information.

[0363] The memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode the encoded bit stream is stored. The memory b2 may be an electronic circuit and may be connected to the processor b1. The memory b2 may be included in the processor b1. The memory b2 may be a collection of multiple electronic circuits. The memory b2 may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium, etc. The memory b2 may be a non-volatile memory or a volatile memory.

[0364] For example, the memory b2 may store a video image or a coded bitstream, and may store a program for the processor b1 to decode the coded bitstream.

[0365] Also, for example, the memory b2 may play the role of a component for storing information among the multiple components of the decoding device 200 shown in Fig. 41 etc. Specifically, the memory b2 may play the role of the block memory 210 and the frame memory 214 shown in Fig. 41. More specifically, the memory b2 may store reconstructed blocks, reconstructed pictures, etc.

[0366] Note that not all of the components shown in Fig. 41 and the like may be implemented, and not all of the above-described processes may be performed, in the decoding device 200. Some of the components shown in Fig. 41 and the like may be included in another device, and some of the above-described processes may be executed by another device.

[0367] [Definition of each term] As an example, each term may be defined as follows:

[0368] A picture is an array of luma samples in monochrome format, or two corresponding arrays of luma samples and chroma samples in 4:2:0, 4:2:2 and 4:4:4 color formats. A picture may be a frame or a field.

[0369] A frame is a composition of a top field from which a number of sample rows occur: 0, 2, 4, . . . and a bottom field from which a number of sample rows occur: 1, 3, 5, . . .

[0370] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit.

[0371] A tile is a rectangular region of multiple coding tree blocks in a particular tile column and a particular tile row in a picture. A tile may also be a rectangular region of a frame that is intended to be decoded and coded independently, although loop filters across tile edges may still be applied.

[0372] A block is an MxN (N rows and M columns) array of samples, or an MxN array of transform coefficients. A block may be a square or rectangular region of pixels consisting of one luma and two chroma matrices.

[0373] A CTU (coding tree unit) may be a coding tree block of luma samples for a picture with three sample arrangements, or two corresponding coding tree blocks of chroma samples, or a coding tree block of samples for either monochrome pictures or pictures coded with three separated color planes and a syntax structure used for coding the samples.

[0374] A superblock may comprise one or two mode information blocks, or may be a square block of 64x64 pixels that can be recursively divided into four 32x32 blocks and further divided.

[0375] [Explanation of secondary conversion process] FIG. 47 is a diagram for explaining the secondary transform process in the embodiment. The secondary transform process is a transform process that the encoding device 100 or the decoding device 200 performs on the prediction residual signal after performing the primary transform on the prediction residual signal. In the secondary transform process, an orthogonal transform or the like is performed as a transform process. The implementation area of ​​the secondary transform process may be different from the implementation area of ​​the primary transform process. For example, even if the primary transform process is performed on the entire target block, the secondary transform process may be performed on a part of the target block as shown in FIG. 47. Here, the part of the target block may be, for example, a sub-block on the low frequency side.

[0376] Furthermore, the size of the sub-block on which the secondary transform process is performed does not have to be a fixed size. For example, the encoding device 100 may change the size of the sub-block on which the secondary transform process is performed depending on the block size of the current block.

[0377] Furthermore, the primary conversion process and the secondary conversion process may be separable or non-separable.

[0378] There may be a plurality of candidates for the basis used in the secondary transform process. For example, the encoding device 100 may hold a total of six basis candidates, including 4×4 basis A, 4×4 basis B, 4×4 basis C, 8×8 basis D, 8×8 basis E, and 8×8 basis F. The encoding device 100 may select a candidate to be used in the secondary transform process from among the plurality of candidates, and write information about the selected candidate into the bitstream.

[0379] When selecting a basis candidate to be used in the secondary conversion process from a plurality of basis candidates, the number of basis candidates to be used may be limited based on any parameter. For example, when selecting a basis candidate to be used in the secondary conversion process from a plurality of basis candidates, if the length of the short side of the processing target block is 8 or more, a basis of 8×8 may be used. Also, when selecting a basis candidate to be used in the secondary conversion process from a plurality of basis candidates, if the length of the short side of the processing target block is 4, a basis of 4×4 may be used.

[0380] [Internal configuration of the conversion unit of the encoding device] FIG. 48 is a flowchart showing a processing procedure in a conversion unit of the encoding device according to the embodiment.

[0381] First, the encoding device 100 determines whether the block to be processed is equal to or smaller than a predetermined block size (step S1000). Here, for example, the predetermined block size may be a 4×4 square block size. Alternatively, the predetermined block size may be a 4×8 or 8×4 rectangular block size. Alternatively, the predetermined block size may be the smallest block size among the block sizes of candidates for bases used in the secondary transform process that the encoding device 100 can select.

[0382] When the encoding device 100 determines that the block to be processed is equal to or smaller than the predetermined block size (Yes in step S1000), the encoding device 100 ends the operation without performing the secondary conversion process on the block to be processed. At this time, the encoding device 100 does not need to write a signal related to the secondary conversion process into the bit stream. In other words, the encoding device 100 does not need to code a signal related to the secondary conversion process in the bit stream.

[0383] When the encoding device 100 determines that the current block is larger than the predetermined block size (No in step S1000), the encoding device 100 determines whether or not to apply secondary transformation processing to the current block (step S1001).

[0384] When the encoding device 100 determines that the secondary transform process is applied to the block to be processed (Yes in step S1001), the encoding device 100 selects one base candidate from one or more base candidates in the secondary transform process (step S1002). Here, the determination in step S1001 and the selection in step S1002 may be made according to information such as the encoding mode of the block to be processed. Furthermore, the determination in step S1001 and the selection in step S1002 may be made by evaluating the cost by performing a tentative transform process or the like using each of the base candidates in the one or more secondary transform processes in step S1002. Furthermore, a signal indicating the results of the determination and selection made in steps S1001 and S1002 may be written into the bitstream by the encoding device 100. That is, the signal indicating the results of the determination and selection made in steps S1001 and S1002 may be coded in the bitstream by the encoding device 100.

[0385] In step S1002, one or more candidate bases for the secondary transform process may be changed depending on the size of the target block. For example, when the length of the short side of the target block is smaller than 16, the encoding device 100 may select a 4×4 square base as a candidate base for use in the secondary transform process. When the length of the short side of the target block is 16 or more, the encoding device 100 may select an 8×8 square base as a candidate base for use in the secondary transform process.

[0386] Next, the encoding device 100 performs a secondary transform process using the basis candidates selected by the encoding device 100 in step S1002 (step S1003), and then the encoding device 100 ends its operation.

[0387] If the encoding device 100 determines that the secondary transformation process is not to be applied to the current block (No in step S1001), the encoding device 100 ends the operation.

[0388] Note that the processing flow described in FIG. 48 is just an example, and the order of the processing described in FIG. 48 may be changed, some of the processing described may be removed, or processing not described may be added.

[0389] 48 are similarly performed in the inverse transform unit of the decoding device 200. In the inverse transform unit of the decoding device 200, the operation of encoding a signal in a bit stream, which is performed in the transform unit of the encoding device 100, is changed to the operation of decoding a signal from the bit stream.

[0390] Note that the processing flow of the decoding device 200 described above is merely an example, and the order of the processing described may be changed, some of the processing described may be removed, or processing not described may be added.

[0391] Fig. 49A is a table showing an example of the amount of processing required for primary conversion processing per block in the embodiment. Fig. 49B is a table showing an example of the amount of processing required for secondary conversion processing per block in the embodiment. According to the configuration of the embodiment, the encoding device 100 or the decoding device 200 may be able to reduce the amount of processing required for conversion processing.

[0392] 49A and 49B, the amount of processing required for the primary conversion process and the secondary conversion process for each block will be described with a specific example. The amount of processing required for the primary conversion process and the secondary conversion process for the entire CTU (Coding Tree Unit) is, for example, (Processing amount required for primary conversion processing and secondary conversion processing of the entire CTU) = {(Processing amount required for primary conversion processing) + (Processing amount required for secondary conversion processing) × (Number of blocks to be spread in the CTU)} It can be calculated using the following formula.

[0393] In the primary transformation process, the block size of the processing target block is set to a square block size using values ​​that are powers of 2, such as 4 x 4, 8 x 8, 16 x 16, and 32 x 32. Figures 49A and 49B show the numerical values ​​assumed as the number of processes required for the primary transformation process and the secondary transformation process for each of the above block sizes.

[0394] Here, the amount of processing required for the primary conversion process and the secondary conversion process may be interpreted as the number of multiplications, the number of additions, and the sum of the number of multiplications and the number of additions.

[0395] Also, here, it is assumed that the size of the sub-block on which the secondary transformation process is performed, that is, the size of the base used in the secondary transformation process, is a 4×4 square or an 8×8 square.

[0396] Fig. 50 is a table showing a first example in the embodiment. In Fig. 50, a first example in which the basis candidates used in the secondary conversion process are only basis candidates of a 4 x 4 square size will be described.

[0397] Assume that the shape of the CTU in the first example is a square of 128 x 128. For example, when a 4 x 4 square block is a processing target block, the amount of processing required for the primary conversion process and the secondary conversion process of the entire CTU is calculated by the following formula.

[0398] (48+256)×{(128 / 4)^2}=311296(times)

[0399] The amount of processing required for the primary transform process and the secondary transform process of the entire CTU for each block size of the block to be processed, calculated by the same calculation as above, is shown in Fig. 50. In the first example shown in Fig. 50, the encoding device 100 or the decoding device 200 uses a base of a 4 × 4 square size in the secondary transform process for all sizes of blocks to be processed that are subjected to the primary transform process.

[0400] In the first example shown in FIG. 50, the largest processing amount is required when the block size of the processing target block is a square block size of 32×32. In contrast, the second largest processing amount is required when the block size of the processing target block is a square block size of 4×4, which is the largest number of blocks in the CTU. However, for example, when the block size of the processing target block is 8×8 or more, if the encoding device 100 uses a base of a square size of 8×8 for the secondary transform processing, the processing amount increases more significantly than the processing amount shown in FIG. 50. That is, in the first example shown in FIG. 50, the size of the sub-block on which the secondary transform processing is performed is reduced to suppress the processing amount required for the primary transform processing and the secondary transform processing. The first example is a preferable example of a candidate base used for the secondary transform processing selected for the block size of the processing target block on which the primary transform processing is performed.

[0401] However, it is assumed that the conversion process of the CTU performed by the conversion unit of the encoding device 100 or the inverse conversion unit of the decoding device 200 involves processes other than the primary conversion process and the secondary conversion process shown in FIG. 50. Therefore, depending on the amount of processing other than the primary conversion process and the secondary conversion process, in the first example, the amount of processing required when the block size of the processing target block is 4×4 may be significantly larger than in other cases. Here, the other processing is processing required for each processing target block. For example, it is pre-processing or post-processing for performing the conversion process. Specifically, the pre-processing is a process of determining the memory storage method to be used, copying data to the memory, converting the copied data, scanning the converted data in units of blocks, and transmitting it. Therefore, the number of blocks in the CTU on which the primary conversion process is performed and the number of sub-blocks on which the secondary conversion process is performed are the largest in the case of 4×4, and therefore, when the processing other than the primary conversion process and the secondary conversion process is taken into consideration, the amount of processing may be the largest.

[0402] Therefore, an example will be shown for reducing the amount of processing performed by the encoding device 100, taking into consideration processes other than those required for the primary transform process and the secondary transform process. In the following example, examples of candidate bases used in the secondary transform process and selected for each block size of a target block on which the primary transform process is performed will be described.

[0403] Fig. 51 is a table showing a second example in the embodiment. In the second example shown in Fig. 51, the encoding device 100 or the decoding device 200 does not perform a secondary transform process when the block size of the processing target block to be subjected to the primary transform process is 4x4, and performs a secondary transform process using a candidate base of a 4x4 square size when the block size of the processing target block is other than 4x4. Note that instead of the encoding device 100 not performing the secondary transform process, the encoding device 100 may be configured to perform the secondary transform process using a base having a transform characteristic such that coefficient values ​​before and after the transform are equal. Fig. 51 shows the amount of processing required for the primary transform process and the secondary transform process in the entire CTU for each block size of the processing target block in the second example calculated by the calculation formula used in Fig. 50.

[0404] As shown in Fig. 51, in the second example, the amount of processing required for the primary transformation process and the secondary transformation process when the block size of the processing target block is 4x4, in which the amount of processing other than the primary transformation process and the secondary transformation process that occurs for each processing target block is the largest, is reduced compared to the first example. Therefore, even when the amount of processing other than the primary transformation process and the secondary transformation process that occurs for each processing target block is large, it is possible to suppress the maximum amount of processing that may occur in the transformation process of the entire CTU. Therefore, the encoding device 100 can promote reduction in the circuit scale in a device implemented to perform the transformation process.

[0405] In the second example described in FIG. 51, the encoding device 100 or the decoding device 200 does not perform the secondary transform process when the block size of the processing target block to be subjected to the primary transform process is 4×4. However, the encoding device 100 or the decoding device 200 may be configured not to perform the secondary transform process when the block size of the processing target block to be subjected to the primary transform process is other than 4×4. For example, the encoding device 100 or the decoding device 200 may not perform the secondary transform process when the block size of the processing target block to be subjected to the primary transform process is 8×8. Also, for example, the encoding device 100 or the decoding device 200 may not perform the secondary transform process when the block size of the processing target block to be subjected to the primary transform process is 4×8 or 8×4. Also, for example, the encoding device 100 or the decoding device 200 may not perform the secondary transform process when the block size of the processing target block to be subjected to the primary transform process is other than 8×8, 4×8, and 8×4. In other words, the encoding device 100 may be configured not to perform the secondary transform process when the size of the block to be processed is equal to or smaller than the smallest block size among one or more block sizes selectable in the secondary transform process. In this case, the encoding device 100 may be configured to be able to apply the secondary transform process when the size of the block to be processed is larger than the smallest block size among one or more block sizes selectable in the secondary transform process.

[0406] Alternatively, instead of the encoding device 100 performing the secondary transform processing, the encoding device 100 may be configured to perform the secondary transform processing using a basis having transform characteristics such that the coefficient values ​​before and after the transform are equal.

[0407] With the above configuration, the encoding device 100 or the decoding device 200 can set the block size of the target block on which the primary transformation process is performed to mean that there is no candidate for the secondary transformation base for the block size that may require a large amount of processing other than the primary transformation process and the secondary transformation process required for each target block. In other words, the encoding device 100 or the decoding device 200 can set the block size of the target block on which the primary transformation process is performed to mean that the secondary transformation process is not performed for the block size that may require a large amount of processing other than the primary transformation process and the secondary transformation process required for each target block.

[0408] For example, the encoding device 100 or the decoding device 200 can set the candidate secondary transform base to mean that there is no candidate when the block size of the processing target block to be subjected to the primary transform process is 8×8. Also, for example, the encoding device 100 or the decoding device 200 can set the candidate secondary transform base to mean that there is no candidate when the block size of the processing target block to be subjected to the primary transform process is 4×8 or 8×4. Also, for example, the encoding device 100 or the decoding device 200 can set the candidate secondary transform base to mean that there is no candidate when the block size of the processing target block to be subjected to the primary transform process is other than 8×8, 4×8, and 8×4. Also, for example, the encoding device 100 or the decoding device 200 can set the candidate secondary transform base to mean that there is no candidate when the block size of the processing target block to be subjected to the primary transform process is other than 4×4.

[0409] This allows the encoding device 100 or the decoding device 200 to improve the possibility of suppressing the maximum amount of processing that may occur in the conversion process of a CTU. Therefore, the encoding device 100 or the decoding device 200 can promote reduction in the circuit scale in a device implemented for performing the conversion process.

[0410] When encoding device 100 performs secondary transform processing using a 4×4 square base, part of the 4×4 square base may be set to 0. In other words, the 4×4 square base may have transform characteristics such that some transform coefficient values ​​of the target block after secondary transform processing are forcibly set to 0.

[0411] FIG. 52 is a table showing a third example in the embodiment. In the third example shown in FIG. 52, when the block size of the block to be processed is 4×4, the secondary transform process is not performed. When the block size of the block to be processed is 8×8, the secondary transform process is performed using a base of a square size of 4×4. When the block size of the block to be processed is 16×16 or 32×32, the secondary transform process is performed using a base of a square size of 8×8. Note that instead of the encoding device 100 not performing the secondary transform process, the encoding device 100 may be configured to perform the secondary transform process using a base having a transform characteristic such that coefficient values ​​before and after the transform are equal. FIG. 52 shows the amount of processing required for the primary transform process and the secondary transform process in the entire CTU for each block size of the block to be processed in the third example calculated by the calculation formula used in FIG. 50.

[0412] As shown in FIG. 52, in the third example, the amount of processing required for the primary transformation and secondary transformation increases when 16×16 and 32×32 processing target blocks are used, compared with the first and second examples. However, the amount and rate of increase are not large. On the other hand, in the third example, the amount of processing required for the primary transformation and secondary transformation when the block size of the processing target block is 4×4, in which the amount of processing other than the primary transformation and secondary transformation processing that occurs for each processing target block is the largest, is reduced compared with the first example. Therefore, even when the amount of processing other than the primary transformation and secondary transformation processing that occurs for each processing target block is large, it is possible to suppress the maximum amount of processing that may occur in the transformation processing of the entire CTU. In addition, since the size of the base used in the secondary transformation is partially larger than that of the first and second examples, it is possible to perform more efficient transformation processing and improve the encoding efficiency. Therefore, the encoding device 100 can promote reduction in the circuit scale in a device implemented for performing the transformation processing.

[0413] When performing secondary transformation processing using a base of 8×8 square size, a part of the base of 8×8 square size may be set to 0. In other words, the base of 8×8 square size may have transformation characteristics such that some transformation coefficient values ​​of the processing target block after secondary transformation processing are forcibly set to 0.

[0414] Note that the processes described in the second example illustrated in FIG. 51 and the third example illustrated in FIG. 52 are not necessarily applied when the amount of processing required for each processing block other than the primary conversion processing and the secondary conversion processing is large. The processes described in the second example illustrated in FIG. 51 and the third example illustrated in FIG. 16 may be applied when the amount of processing required for each processing block other than the primary conversion processing and the secondary conversion processing is small. In this case, the maximum amount of processing that may occur in the conversion processing of the CTU is smaller than when the processes described in the second example illustrated in FIG. 51 and the third example illustrated in FIG. 52 are applied when the amount of processing required for each processing block other than the primary conversion processing and the secondary conversion processing is large. Therefore, the encoding device 100 or the decoding device 200 can promote reduction in the circuit scale in a device implemented for performing the conversion processing.

[0415] FIG. 53 is a table showing a fourth example in the embodiment. In the first example shown in FIG. 50, among the processing target blocks to be subjected to the primary transformation process, the base of 4×4 square size used for the secondary transformation is commonly used for all sizes of processing target blocks, thereby making it possible to suppress the maximum amount of processing required for the primary transformation process and the secondary transformation process. However, the above method is difficult to deal with the fact that the tendency of coefficient values ​​differs depending on the size of the processing target block during the primary transformation process. For example, there is a high possibility that the tendency of coefficient values ​​is significantly different between the coefficient values ​​after the primary transformation process in a processing target block of 4×4 square size and the coefficient values ​​after the primary transformation process in a 4×4 square region corresponding to the low frequency side of a processing target block of 16×16 square size. In this case, if a base candidate used in a common secondary transformation process is used for processing target blocks of different sizes, there is a high possibility that the optimal base candidate cannot be used for the secondary transformation process.

[0416] Therefore, in the fourth example, as shown in FIG. 53, even if the bases used in the secondary transform process are the same size, the encoding device 100 assigns a candidate group with different bases used as the secondary transform bases for each size of the processing target block in the primary transform process to the sub-block in which the secondary transform process is performed. The encoding device 100 selects a base to be actually applied in the secondary transform process from the candidate group assigned to the sub-block in which the secondary transform process is performed. Here, the candidate group may have multiple candidates for the base used in the secondary transform process, or may have one candidate for the base used in the secondary transform process. The candidates included in the candidate group may be multiple candidates that differ depending on the direction of intra prediction.

[0417] As a result, the optimal transformation base is defined according to the tendency of coefficient values ​​after the primary transformation process in the area in the processing target block where the secondary transformation process is performed as a group of candidate bases used in the secondary transformation process assigned to the processing target block in the primary transformation process according to the block size. Therefore, the encoding device 100 can select more appropriate candidate bases used in the secondary transformation process than in the first example.

[0418] In the fourth example described in Fig. 53, the shape of the base used in the secondary transform process is a 4 x 4 square, but it may be a shape other than a 4 x 4 square. In the fourth example, a base of a different size may be used in the secondary transform process depending on the size of the block to be processed. Furthermore, the encoding device 100 may not perform the secondary transform process for some sizes of blocks to be processed.

[0419] In addition, in the fourth example described in Figure 53, the encoding device 100 uses a different set of candidates for the bases used in the secondary transformation process for each size of the processing target block on which the primary transformation process is performed, but the encoding device 100 may also use a common set of candidates for the bases used in the secondary transformation process for blocks of different sizes on which the primary transformation process is performed.

[0420] In the example described in Figures 50 to 53, a configuration is shown in which one base candidate is used when a secondary conversion process is performed for each block size of a processing target block to which a primary conversion process is performed. However, in the example described in Figures 50 to 53, a plurality of base candidates may be used when a secondary conversion process is performed for each block size of a processing target block to which a primary conversion process is performed. Also, in the example described in Figures 50 to 53, among the block sizes of a processing target block to which a primary conversion process is performed, there may be a block size that has a plurality of base candidates to be used when a secondary conversion process is performed.

[0421] For example, when the block size of the processing target block on which the primary transformation process is performed is 32×32, the encoding device 100 or the decoding device 200 may be configured to select a basis from a base of 4×4 size and a base of 8×8 size as a basis to be used in the secondary transformation process. When the block size of the processing target block on which the primary transformation process is performed is 16×16, the encoding device 100 or the decoding device 200 may be configured to select a basis from a base of 4×4 size and a base of 8×8 size as a basis to be used in the secondary transformation process. Furthermore, when the secondary transformation process is performed on a plurality of processing target blocks having a plurality of block sizes on which the primary transformation process is performed, there may be a plurality of candidates for the basis to be used in the secondary transformation process.

[0422] In the fourth example described in FIG. 53, the encoding device 100 selects a basis used in the secondary transform process from a different candidate group of bases for each size of the processing target block on which the primary transform process is performed, but the present invention is not limited to this example. For example, the candidate group may be configured to have different candidates that have a common base, but differ in whether or not the secondary transform process is performed by replacing some coefficients of the base with 0. That is, in a predetermined candidate group, the candidates may have a transform characteristic that some transform coefficient values ​​of the processing target block after the secondary transform process are forcibly set to 0. In other words, a candidate group of different bases is not only a candidate group consisting of candidates that simply include different bases, but may also be considered to be a candidate group of different bases when the candidate group includes candidates that include the same coefficients but have different coefficients of the block after the secondary transform process using the base.

[0423] The encoding device 100 or the decoding device 200 may perform different secondary transform processes depending on the block size of the processing target block on which the primary transform process is performed.

[0424] The processing contents of the encoding device shown in Figures 50 to 53 are similarly performed in the decoding device.

[0425] [Variations] When the block division structure of the block to be processed differs between the color difference signal and the luminance signal, the encoding device 100 or the decoding device 200 may apply the encoding method or decoding method, etc. in the embodiments of the present disclosure to only the luminance signal or only the color difference signal.

[0426] Furthermore, the encoding device 100 or the decoding device 200 may determine whether or not to apply the encoding method or the decoding method according to the embodiments of the present disclosure to the current block on a slice-by-slice or tile-by-tile basis.

[0427] In addition, the encoding device 100 or the decoding device 200 may determine whether to apply the encoding method and the decoding method, etc. in the embodiment of the present disclosure depending on the slice type (I-slice, P-slice, B-slice) in the block to be processed.

[0428] In addition, the encoding device 100 or the decoding device 200 may write a flag indicating that the encoding method or decoding method, etc. in an embodiment of the present disclosure has been applied to the block to be processed, into the syntax of the sequence layer, picture layer, slice layer, etc.

[0429] In addition, when applying the encoding method or the decoding method in the embodiment of the present disclosure to the processing target block, the encoding device 100 or the decoding device 200 may use a determination method different from the determination method of the base candidate used in the secondary transformation process in the embodiment of the present disclosure. In addition, the encoding device 100 or the decoding device 200 may use a determination method different from the determination method of the base candidate used in the secondary transformation process in the embodiment of the present disclosure in combination with the determination method of the base candidate used in the secondary transformation process in the embodiment of the present disclosure. For example, the encoding device 100 or the decoding device 200 may use a combination of the determination method of the base candidate used in the secondary transformation process using the intra prediction mode and the determination method of the base candidate used in the secondary transformation process in the embodiment of the present disclosure.

[0430] In the encoding method, the decoding method, etc. according to the embodiment of the present disclosure, the block to be processed is a square, but the block to be processed does not have to be a square. In the encoding method, the decoding method, etc. according to the embodiment of the present disclosure, the block to be processed may be a rectangle, for example.

[0431] In the encoding method or decoding method, etc., according to the embodiment of the present disclosure, the shape of the base used in the secondary transformation process is a square, but the shape of the base used in the secondary transformation process does not have to be a square. In the encoding method or decoding method, etc., according to the embodiment of the present disclosure, the shape of the base used in the secondary transformation process may be, for example, a rectangle.

[0432] [Representative example] Fig. 54 is a flowchart showing an example of the operation of the encoding device in the embodiment. For example, the encoding device 100 shown in Fig. 40 performs the operation shown in Fig. 54 when performing a transform process of applying a secondary transform process to a prediction residual signal that has been subjected to a primary transform. Specifically, the processor a1 performs the following operation using the memory a2.

[0433] First, the encoding device 100 selects one transform basis from a candidate group that is made up of one or more transform basis candidates and that differs depending on the block size of a block to be processed (Step S2001).

[0434] Next, the encoding device 100 further applies a secondary transform of a common block size to the transform coefficients obtained by applying the primary transform to the prediction residual signal (Step S2002).

[0435] Furthermore, in the encoding device 100, the transform base of the secondary transform may be a 4×4 square.

[0436] Furthermore, in the encoding device 100, the transform base of the secondary transform may be an 8×8 square.

[0437] Furthermore, in the encoding device 100, common transform base candidates may be assigned in the secondary transform to processing target blocks of some sizes among a plurality of block sizes.

[0438] In addition, the encoding device 100 may determine not to apply secondary transformation to the transform coefficients when the block size of the block to be processed is equal to or smaller than a predetermined block size, and may determine to apply secondary transformation to the transform coefficients when the block size of the block to be processed is larger than the predetermined block size.

[0439] Furthermore, the predetermined block size of the current block when encoding apparatus 100 determines not to apply secondary transform to the current block may be a 4×4 square.

[0440] Furthermore, the predetermined block size of the current block when encoding apparatus 100 determines not to apply secondary transform to the current block may be a 4×8 or 8×4 rectangle.

[0441] In addition, when encoding device 100 determines not to apply secondary transformation to a target block, the specified block size of the target block may be equal to the smallest block size of the target block that can be selected for the secondary transformation among the block sizes of the target block that can be selected for the secondary transformation.

[0442] Fig. 55 is a flowchart showing an example of the operation of the decoding device in the embodiment. For example, the decoding device 200 shown in Fig. 46 performs the operation shown in Fig. 55 when performing inverse transform processing in which a primary transform is further applied to transform coefficients to which a secondary transform has been applied. Specifically, the processor b1 performs the following operation using the memory b2.

[0443] First, the decoding device 200 selects one transformation basis from a candidate group that is made up of one or more transformation basis candidates and that differs depending on the block size of a processing target block (Step S3001).

[0444] Next, the decoding device 200 performs an inverse transform process of applying a linear transform to the transform coefficients obtained by applying a secondary transform of a common block size to the transform coefficient signals (Step S3002).

[0445] Furthermore, in the decoding device 200, the transform base of the secondary transform may be a 4×4 square.

[0446] Furthermore, in the decoding device 200, the transform base of the secondary transform may be an 8×8 square.

[0447] Furthermore, in the decoding device 200, a common candidate transform base may be assigned in the secondary transform to processing target blocks of some sizes among a plurality of block sizes.

[0448] In addition, the decoding device 200 may determine not to apply secondary transformation to the transform coefficients when the block size of the block to be processed is equal to or smaller than a predetermined block size, and may determine to apply secondary transformation to the transform coefficients when the block size of the block to be processed is larger than the predetermined block size.

[0449] Furthermore, the predetermined block size of the current block when the decoding device 200 determines not to apply secondary transform to the current block may be a 4×4 square.

[0450] Furthermore, when the decoding device 200 determines not to apply secondary transform to the current block, the predetermined block size of the current block may be a 4×8 or 8×4 rectangle.

[0451] In addition, when the decoding device 200 determines not to apply secondary transformation to the target block, the specified block size of the target block may be equal to the smallest block size of the target block that can be selected in the secondary transformation among the block sizes of the target block that can be selected in the secondary transformation.

[0452] [supplement] The encoding device 100 and the decoding device 200 in this embodiment may be used as an image encoding device and an image decoding device, respectively, or may be used as a video encoding device and a video decoding device.

[0453] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0454] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuitry and a storage device electrically connected to and accessible from the processing circuitry. For example, the processing circuitry corresponds to the processor a1 or b1, and the storage device corresponds to the memory a2 or b2.

[0455] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a storage device. In addition, when the processing circuit includes a program execution unit, the storage device stores a software program to be executed by the program execution unit.

[0456] Here, the software for realizing the encoding device 100 or the decoding device 200 according to the present embodiment is a program as follows.

[0457] In other words, this program may cause the computer to perform a transformation process in which a transform coefficient obtained by applying a linear transform to a prediction residual signal in a target block among multiple blocks of multiple block sizes is further transformed by applying a secondary transform of a block size common to the multiple blocks, and the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0458] Alternatively, the program may cause the computer to execute a decoding method in which, in a target block among a plurality of blocks of a plurality of block sizes, an inverse transform process is performed in which a linear transform is further applied to transform coefficients obtained by applying a secondary transform of a block size common to the plurality of blocks to a transform coefficient signal, and the secondary transform of the common block size is composed of one or more candidates for a transform base, and one of the transform bases is selected from a group of candidates that differ depending on the block size of the target block.

[0459] Also, each component may be a circuit, as described above. These circuits may form one circuit as a whole, or each may be a separate circuit. Also, each component may be realized by a general-purpose processor, or may be realized by a dedicated processor.

[0460] Furthermore, a process executed by a specific component may be executed by another component. The order in which the processes are executed may be changed, or multiple processes may be executed in parallel. Furthermore, the encoding / decoding device may include the encoding device 100 and the decoding device 200.

[0461] In addition, the ordinal numbers such as first and second used in the description may be changed as appropriate. Furthermore, new ordinal numbers may be given to components, or ordinal numbers may be removed.

[0462] Although the aspects of the encoding device 100 and the decoding device 200 have been described above based on the embodiment, the aspects of the encoding device 100 and the decoding device 200 are not limited to this embodiment. As long as it does not deviate from the spirit of this disclosure, various modifications conceived by a person skilled in the art to this embodiment and forms constructed by combining components in different embodiments may also be included within the scope of the aspects of the encoding device 100 and the decoding device 200.

[0463] This aspect may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some of the processes, some of the configurations of the device, and some of the syntax described in the flowcharts of this aspect may be implemented in combination with other aspects.

[0464] (Embodiment 2) [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can usually be realized by an MPU (micro processing unit) and a memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memories. It is also possible to realize each functional block by hardware (dedicated circuitry).

[0465] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Also, the processor that executes the above program may be single or multiple. That is, centralized processing or distributed processing may be performed.

[0466] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, which are also included within the scope of the aspects of the present disclosure.

[0467] Further, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems implementing the application examples will be described. Such a system may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device including both. Other configurations of such a system can be appropriately changed depending on the case.

[0468] [Usage example] 56 is a diagram showing the overall configuration of an appropriate content supply system ex100 for realizing a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0469] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may be configured to connect any of the above devices in combination. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network or short-distance wireless communication, etc., without the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 may be connected to a terminal in a hotspot in an airplane ex117, etc., via a satellite ex116.

[0470] Instead of the base stations ex106 to ex110, wireless access points or hot spots may be used. The streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0471] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handyphone System) that supports the mobile communication system, such as 2G, 3G, 3.9G, 4G, and in the future, 5G.

[0472] The home appliance ex114 is a refrigerator, or an appliance included in a home fuel cell cogeneration system.

[0473] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live distribution or the like. In live distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content photographed by a user using the terminal, may multiplex the video data obtained by encoding with sound data obtained by encoding sound corresponding to the video, and may transmit the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0474] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal in an airplane ex117, or the like, capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0475] [Distributed processing] The streaming server ex103 may be a plurality of servers or computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content distribution may be realized by a network that connects a large number of edge servers distributed around the world. In a CDN, an edge server that is physically close to the client is dynamically assigned according to the client. The content is cached and distributed to the edge server, thereby reducing delays. In addition, when some types of errors occur or communication conditions change due to an increase in traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the part of the network where a failure has occurred, thereby realizing high-speed and stable distribution.

[0476] In addition to the distributed processing of the distribution itself, the encoding processing of the captured data may be performed by each terminal, may be performed by the server side, or may be shared among the terminals. As an example, in the encoding processing, a processing loop is generally performed twice. In the first loop, the complexity of the image or the amount of code is detected for each frame or scene. In the second loop, processing is performed to maintain the image quality and improve the encoding efficiency. For example, the terminal performs the first encoding processing, and the server side that receives the content performs the second encoding processing, thereby improving the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode almost in real time, the data encoded once by the terminal can be received and played back by other terminals, making it possible to perform more flexible real-time distribution.

[0477] As another example, the camera ex113 etc. extracts features from an image, compresses data related to the features as metadata, and transmits the data to the server. The server performs compression according to the meaning of the image (or the importance of the content), for example by determining the importance of an object from the features and switching the quantization precision. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server re-compresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a high processing load such as CABAC (context-adaptive binary arithmetic coding).

[0478] As another example, in a stadium, a shopping mall, a factory, etc., there may be a plurality of video data in which almost the same scene has been shot by a plurality of terminals. In this case, using the plurality of terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, coding processing is assigned to each of them, for example, in units of GOPs (Group of Pictures), in units of pictures, or in units of tiles obtained by dividing a picture, for distributed processing. This reduces delays and realizes better real-time performance.

[0479] Since the multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referenced. The server may also receive the encoded data from each terminal and change the reference relationship between the multiple data, or correct or replace the pictures themselves and re-encode them. This makes it possible to generate a stream with improved quality and efficiency for each piece of data.

[0480] Furthermore, the server may distribute the video data after performing transcoding to change the encoding method of the video data. For example, the server may convert an MPEG-based encoding method to a VP-based encoding method (e.g., VP9), or convert H.264 to H.265.

[0481] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, descriptions such as "server" or "terminal" are used to indicate the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0482] [3D, multi-angle] Images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are almost synchronized with each other are increasingly being integrated and used. The videos taken by each device are integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points included in the videos match.

[0483] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. If the server can obtain the relative positional relationship between the shooting terminals, the server may generate a 3D shape of the scene based on not only 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point cloud or the like, or may generate images to be transmitted to the receiving terminal by selecting or reconstructing images from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0484] In this way, the user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting terminal, or can enjoy content in which a video of a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0485] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and the left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0486] In the case of an AR image, the server superimposes virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may obtain or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect them to create superimposed data. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to the request for virtual object information. The server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and deliver it to the decoding device. Note that the superimposed data has an α value indicating the transparency in addition to RGB, and the server may set the α value of the part other than the object created from the three-dimensional data to 0, etc., and encode the data in a state in which the part is transparent. Alternatively, the server may generate data in which a predetermined value of RGB value is set to the background like a chromakey, and the part other than the object is the background color.

[0487] Similarly, the decoding process of the distributed data may be performed by each client terminal, or may be performed by the server side, or may be shared among them. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request may be received by other terminals, decoded, and the decoded signal may be transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-enabled terminals themselves, data with good image quality can be reproduced. In another example, while large-sized image data is received by a TV or the like, a part of the area, such as tiles into which the picture is divided, may be decoded and displayed on the viewer's personal terminal. This allows the viewer to share the overall picture while checking his / her own area of ​​responsibility or the area he / she wants to check in more detail at hand.

[0488] It may be possible to seamlessly receive content using delivery system standards such as MPEG-DASH in situations where multiple short-range, medium-range, or long-range wireless communication is available indoors and outdoors. A user may freely select and switch in real time between a decoding device or a display device, such as a user's terminal or a display device placed indoors or outdoors. In addition, decoding can be performed while switching between a decoding device and a display device using the user's location information, etc. This makes it possible to map and display information on a part of the wall or ground of a neighboring building where a displayable device is embedded while the user is moving to a destination. It is also possible to switch the bit rate of the received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be accessed from the receiving terminal in a short time, or copied to an edge server in a content delivery service.

[0489] [Scalable Coding] The switching of contents will be described using a scalable stream compressed and coded by applying the video coding method shown in each of the above embodiments, as shown in FIG. 57. The server may have multiple streams with the same content but different qualities as individual streams, but may be configured to switch contents by taking advantage of the characteristics of a temporal / spatial scalable stream realized by coding in layers as shown in the figure. In other words, the decoding side can freely switch and decode low-resolution content and high-resolution content by determining which layer to decode according to an internal factor such as performance and an external factor such as the state of the communication band. For example, if a user wants to continue watching a video that he or she was watching on the smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream up to a different layer, thereby reducing the burden on the server side.

[0490] Furthermore, as described above, pictures are coded for each layer, and in addition to the configuration in which scalability is realized in an enhancement layer above the base layer, the enhancement layer may include meta-information based on image statistics and the like. The decoding side may generate high-quality content by super-resolving pictures in the base layer based on the meta-information. The super-resolution may improve the signal-to-noise ratio while maintaining and / or increasing the resolution. The meta-information includes information for specifying linear or non-linear filter coefficients to be used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least squares calculation to be used in the super-resolution process.

[0491] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of an object in an image. The decoding side selects tiles to be decoded to decode only a part of the area. Furthermore, by storing attributes of objects (such as a person, a car, a ball, etc.) and positions in a video (such as coordinate positions in the same image) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 58, the meta information may be stored using a data storage structure different from pixel data, such as a supplemental enhancement information (SEI) message in HEVC. This meta information indicates, for example, the position, size, or color of a main object.

[0492] Meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in a video, and by combining the picture-by-picture information with the time information, it can identify the picture in which an object exists and determine the position of the object within the picture.

[0493] [Web page optimization] FIG. 59 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. FIG. 60 is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As shown in FIG. 59 and FIG. 60, a web page may include multiple link images that are links to image content, and the appearance of the web page differs depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture that each content has as a link image, or may display an image such as a gif animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the image until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image enters the screen.

[0494] When a link image is selected by a user, the display device performs decoding while giving top priority to the base layer. If the HTML constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, in order to ensure real-time performance, before selection or when the communication bandwidth is very tight, the display device decodes and displays only forward-reference pictures (I-pictures, P-pictures, and B-pictures with forward reference only), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of decoding the content to the start of display). Furthermore, the display device may intentionally ignore the reference relationship of pictures, roughly decode all B-pictures and P-pictures with forward reference, and perform normal decoding as the number of received pictures increases over time.

[0495] [Automatic driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0496] In this case, since a car, drone, or airplane including a receiving terminal moves, the receiving terminal can realize seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting location information of the receiving terminal. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information according to a user's selection, a user's situation, and / or a communication band state.

[0497] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0498] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distributors, but also low-quality, short-duration content from individuals via unicast or multicast distribution. Such personal content is expected to continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, by using the following configuration.

[0499] During shooting, in real time or after accumulating, the server performs recognition processing such as shooting errors, scene search, semantic analysis, and object detection from the original image data or the encoded data. Then, based on the recognition results, the server manually or automatically performs editing such as correcting focus deviation or camera shake, deleting less important scenes such as scenes that are less bright than other pictures or out of focus, emphasizing object edges, and changing color. The server encodes the edited data based on the editing results. It is also known that if the shooting time is too long, the viewer rating will decrease, and the server may automatically clip not only scenes with less importance as described above but also scenes with little movement based on the image processing results so that the content will be within a specific time range depending on the shooting time. Alternatively, the server may generate a digest based on the result of the semantic analysis of the scene and encode it.

[0500] In some cases, personal content may contain content that infringes copyright, moral rights, or portrait rights, and the range of sharing may exceed the intended range, which may be inconvenient for individuals. Therefore, for example, the server may change the image to an unfocused image of a person's face on the periphery of the screen, or the inside of a house, and encode it. Furthermore, the server may recognize whether the image to be encoded contains a face of a person other than a person registered in advance, and if so, may perform processing such as applying a mosaic to the face. Alternatively, as pre-processing or post-processing of encoding, the user may specify a person or background area that he or she wishes to process in the image from the viewpoint of copyright, etc. The server may replace the specified area with another image, or may perform processing such as blurring the focus. If it is a person, the person can be tracked in the video and the image of the person's face can be replaced.

[0501] Since viewing of personal content with a small amount of data requires real-time performance, the decoding device first receives the base layer as a top priority and performs decoding and playback, although this depends on the bandwidth. The decoding device may receive an enhancement layer during this time, and when playback is looped or otherwise played two or more times, play high-quality video including the enhancement layer. With a stream that has been scalably encoded in this way, it is possible to provide an experience in which the video is rough when not selected or when viewing begins, but the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played the first time and a second stream that is encoded with reference to the first video are configured as a single stream.

[0502] [Other application examples] Moreover, these encoding or decoding processes are generally processed in an LSIex500 possessed by each terminal. The LSI (large scale integration circuitry)ex500 (see FIG. 56) may be a one-chip or a multi-chip configuration. Note that software for encoding or decoding moving images may be incorporated into some kind of recording medium (such as a CD-ROM, a flexible disk, or a hard disk) that can be read by the computer ex111 or the like, and the encoding or decoding process may be performed using the software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera may be transmitted. The video data at this time is data that has been encoded and processed by the LSIex500 possessed by the smartphone ex115.

[0503] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0504] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is carried and transmitted over broadcasting radio waves using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to the encoding process and decoding process.

[0505] [Hardware configuration] FIG. 61 is a diagram showing further details of the smartphone ex115 shown in FIG. 56. FIG. 62 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying the video captured by the camera unit ex465 and the decoded data of the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded data such as captured video or still images, recorded audio, received video or still images, and e-mail, or decoded data, and a slot unit ex464 which is an interface unit with a SIMex468 for identifying a user and authenticating access to various data including a network. In addition, an external memory may be used instead of the memory unit ex467.

[0506] A main control unit ex460, which comprehensively controls the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0507] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state and supplies power to each unit from the battery pack.

[0508] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, a ROM, and a RAM. During a telephone call, a voice signal collected by a voice input unit ex456 is converted into a digital voice signal by a voice signal processing unit ex454, and then subjected to spectrum spreading processing by a modulation / demodulation unit ex452, and then subjected to digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit ex451, and the resulting signal is transmitted via an antenna ex450. In addition, the received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, and then subjected to spectrum inverse spreading processing by a modulation / demodulation unit ex452, and then converted into an analog voice signal by a voice signal processing unit ex454, and then output from a voice output unit ex457. During a data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 based on the operation of an operation unit ex466 of the main unit. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / separation unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing the video or still image, and sends the encoded audio data to the multiplexing / separation unit ex453. The multiplexing / separation unit ex453 multiplexes the encoded video data and the encoded audio data by a predetermined method, and performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450.

[0509] In the case of receiving a video attached to an e-mail or a chat, or a video linked to a web page, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / separation unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and the video or still image included in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. As real-time streaming becomes more and more popular, audio playback may not be socially appropriate depending on the user's situation. Therefore, as an initial setting, it is preferable to have a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0510] Also, although the smartphone ex115 has been described as an example here, three other implementation formats are conceivable for the terminal: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed onto video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Also, video data itself may be received or transmitted instead of the multiplexed data.

[0511] Although the main control unit ex460 including the CPU controls the encoding or decoding process, various terminals are often equipped with a GPU. Therefore, a configuration may be used in which a wide area is processed collectively by utilizing the performance of the GPU using a memory shared by the CPU and GPU, or a memory whose addresses are managed so that they can be used in common. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processing of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transformation and quantization collectively in units such as pictures by the GPU, rather than by the CPU. [Industrial Applicability]

[0512] The present disclosure is applicable to, for example, television receivers, digital video recorders, car navigation systems, mobile phones, digital cameras, digital video cameras, video conference systems, electronic mirrors, and the like. [Explanation of symbols]

[0513] 100 Encoding device 102 Division 104 Subtraction section 106 Conversion unit 108 Quantization section 110 Entropy coding unit 112, 204 Inverse quantization section 114, 206 Inverse conversion unit 116, 208 Addition section 118, 210 Block Memory 120, 212 Loop filter section 122, 214 frame memory 124, 216 Intra prediction section 126, 218 Inter prediction section 128, 220 Predictive control unit 200 Decryption device 202 Entropy Decoding Unit 1201 Boundary determination section 1202, 1204, 1206 Switches 1203 Filter Judgment Unit 1205 Filter processing section 1207 Filter characteristic determination section 1208 Processing Judgment Unit a1, b1 processor a2, b2 memory

Claims

1. The circuit, A memory, The circuit uses the memory to: applying a primary transform to a prediction residual signal indicating a difference between a current block to be coded and a predicted image of the current block, and performing a transform process of further applying a secondary transform to first transform coefficients which are a transform result of the primary transform to generate second transform coefficients of the current block; quantizing the second transform coefficients; In the secondary transformation, when the block size of the current block is a first block size, one transformation base is selected from a first candidate group consisting of one or more transformation base candidates, and when the block size of the current block is a second block size different from the first block size, one transformation base is selected from a second candidate group different from the first candidate group; applying the secondary transform to a portion of the first transform coefficients whenever the block size has a size greater than 4×4; a size of a sub-block to which the secondary transformation is applied among the current block having the first block size is the same as a size of a sub-block to which the secondary transformation is applied among the current block having the second block size; Encoding device.

2. The circuit, A memory, The circuit uses the memory to: applying a secondary transform to a first transform coefficient obtained by inverse quantizing a current block to be decoded, and performing an inverse transform process of further applying a primary transform to a transform result of the secondary transform, and generating an image based on a prediction residual signal obtained by the inverse transform process; In the secondary transformation, when the block size of the current block is a first block size, one transformation base is selected from a first candidate group consisting of one or more transformation base candidates, and when the block size of the current block is a second block size different from the first block size, one transformation base is selected from a second candidate group different from the first candidate group; applying the secondary transform to a portion of the first transform coefficients whenever the block size has a size greater than 4×4; a size of a sub-block to which the secondary transformation is applied among the current block having the first block size is the same as a size of a sub-block to which the secondary transformation is applied among the current block having the second block size; Decryption device.

3. applying a primary transform to a prediction residual signal indicating a difference between a current block to be coded and a predicted image of the current block, and performing a transform process of further applying a secondary transform to first transform coefficients which are a transform result of the primary transform to generate second transform coefficients of the current block; quantizing the second transform coefficients; In the secondary transformation, when the block size of the current block is a first block size, one transformation base is selected from a first candidate group consisting of one or more transformation base candidates, and when the block size of the current block is a second block size different from the first block size, one transformation base is selected from a second candidate group different from the first candidate group; applying the secondary transform to a portion of the first transform coefficients whenever the block size has a size greater than 4×4; a size of a sub-block to which the secondary transformation is applied among the current block having the first block size is the same as a size of a sub-block to which the secondary transformation is applied among the current block having the second block size; Encoding method.

4. applying a secondary transform to a first transform coefficient obtained by inverse quantizing a current block to be decoded, and performing an inverse transform process of further applying a primary transform to a transform result of the secondary transform, and generating an image based on a prediction residual signal obtained by the inverse transform process; In the secondary transformation, when the block size of the current block is a first block size, one transformation base is selected from a first candidate group consisting of one or more transformation base candidates, and when the block size of the current block is a second block size different from the first block size, one transformation base is selected from a second candidate group different from the first candidate group; applying the secondary transform to a portion of the first transform coefficients whenever the block size has a size greater than 4×4; a size of a sub-block to which the secondary transformation is applied among the current block having the first block size is the same as a size of a sub-block to which the secondary transformation is applied among the current block having the second block size; Decryption method.