Symbolizing device and decoding device

By determining asymmetric filter characteristics based on pixel values and error likelihood, the method addresses the challenge of error distribution in video encoding, enhancing encoding efficiency and reducing errors.

JP7699704B2Active Publication Date: 2025-06-27PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024161157
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-05-30
Filing Date
2024-09-18
Publication Date
2025-06-27
Estimated Expiration
2038-04-04

AI Technical Summary

Technical Problem

Existing video encoding technologies, such as HEVC, face challenges in further improving encoding and decoding efficiency, particularly in managing error distribution across block boundaries.

Method used

An encoding and decoding method that determines asymmetric filter characteristics across block boundaries based on pixel values, error likelihood, and quantization parameters, allowing for tailored deblocking filter processing to reduce errors.

Benefits of technology

The proposed method enhances error reduction by increasing the influence of filter processing on pixels with large errors and reducing its influence on pixels with small errors, thereby improving encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699704000003
    Figure 0007699704000003
  • Figure 0007699704000004
    Figure 0007699704000004
  • Figure 0007699704000005
    Figure 0007699704000005
Patent Text Reader

Abstract

To provide an encoding device capable of further improvement.SOLUTION: The encoding device performs a series of luminance filter processing on a first boundary between a first luminance block and a second luminance block and a series of chrominance filter processing on a second boundary between a first color difference block and a second color difference block. In the luminance filter processing, values of multiple luminance samples are changed by using a series of clip processing in a manner that the change in the values of multiple luminance samples included in the first luminance block and the second luminance block which are aligned in a direction perpendicular to the first boundary, does not exceed each of the first thresholds. Each first threshold value is determined based on the size of the first luma block, the size of the second luma block, and the quantization parameter, and the first boundary is set to be symmetric or asymmetric.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method.

Background Art

[0002] A video encoding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In such encoding and decoding technologies, further improvement is required.

[0005] Therefore, an object of the present disclosure is to provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement.

Means for Solving the Problems

[0006] An encoding device according to an aspect of the present disclosure includes a processing circuit and a memory connected to the processing circuit. The processing circuit performs a luminance filter process on a first boundary between a first luminance block and a second luminance block using the memory, and performs a color difference filter process on a second boundary between a first color difference block and a second color difference block. In the luminance filter process, a clipping process is used such that a change amount of values of a plurality of luminance samples included in the first luminance block and the second luminance block arranged in a direction orthogonal to the first boundary does not exceed each first threshold value, and the values of the plurality of luminance samples are changed. Each first threshold value is determined based on the sizes of the first luminance block and the second luminance block and a quantization parameter, and is set symmetrically or asymmetrically with respect to the first boundary. In the color difference filter process, a clipping process is used such that a change amount of values of a plurality of color difference samples included in the first color difference block and the second color difference block arranged in a direction orthogonal to the second boundary does not exceed each second threshold value, and the values of the plurality of color difference samples are changed.

[0007] Note that these general or specific aspects may be implemented in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

Effect of the Invention

[0008] The present disclosure can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvement.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

DETAILED DESCRIPTION OF THE INVENTION

[0010] In the deblocking filter process of H.265 / HEVC, which is one of the image encoding methods, a filter having symmetric characteristics is applied across a block boundary. As a result, for example, when the error distribution is discontinuous, such as when the error of pixels located on one side across the block boundary is small and the error of pixels located on the other side across the block boundary is large, the error reduction efficiency may decrease due to symmetric filter processing. Here, the error is the difference in pixel values between the original image and the reconstructed image.

[0011] An encoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine non-symmetric filter characteristics across a block boundary and performs deblocking filter processing having the determined filter characteristics.

[0012] According to this, the encoding apparatus may be able to reduce errors by performing filter processing having non-symmetric filter characteristics across a block boundary.

[0013] For example, in the determination of the filter characteristics, the non-symmetric filter characteristics may be determined such that the influence of the deblocking filter processing is greater for pixels with a higher likelihood of having a large error from the original image.

[0014] According to this, since the encoding device can increase the influence of the filter processing on pixels with large errors, there is a possibility that the errors of the pixels can be further reduced. Also, since the encoding device can reduce the influence of the filter processing on pixels with small errors, there is a possibility that an increase in the errors of the pixels can be suppressed.

[0015] For example, in determining the filter characteristics, the asymmetric filter characteristics may be determined by changing the filter coefficients of the reference filter asymmetrically across the block boundary.

[0016] For example, in determining the filter characteristics, an asymmetric weight is determined across the block boundary, and in the deblocking filter processing, a filter operation using the filter coefficients is performed, and the amount of change in the pixel value before and after the filter operation may be weighted by the determined asymmetric weight.

[0017] For example, in determining the filter characteristics, an asymmetric offset value is determined across the block boundary, and in the deblocking filter processing, a filter operation using the filter coefficients is performed, and the determined asymmetric offset value may be added to the pixel value after the filter operation.

[0018] For example, in determining the filter characteristics, an asymmetric reference value is determined across the block boundary, and in the deblocking filter processing, a filter operation using the filter coefficients is performed, and when the amount of change in the pixel value before and after the filter operation exceeds the reference value, the amount of change may be clipped to the reference value.

[0019] For example, in determining the filter characteristics, the condition for determining whether to perform the deblocking filter processing may be set asymmetrically across the block boundary.

[0020] The decoding device according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across block boundaries and performs deblocking filter processing having the determined filter characteristics.

[0021] According to this, the decoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across block boundaries.

[0022] For example, in determining the filter characteristics, the asymmetric filter characteristics may be determined such that the greater the likelihood that a pixel has a large error with respect to the original image, the greater the influence of the deblocking filter processing on that pixel.

[0023] According to this, the decoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of those pixels can be further reduced. Also, the decoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of those pixels can be suppressed.

[0024] For example, in determining the filter characteristics, the asymmetric filter characteristics may be determined by changing the filter coefficients of the reference filter asymmetrically across the block boundary.

[0025] For example, in determining the filter characteristics, asymmetric weights are determined across the block boundary. In the deblocking filter processing, a filter operation using the filter coefficients is performed, and the amount of change in the pixel value before and after the filter operation is weighted by the determined asymmetric weights.

[0026] For example, in determining the filter characteristics, asymmetric offset values are determined across the block boundary. In the deblocking filter processing, a filter operation using the filter coefficients is performed, and the determined asymmetric offset values are added to the pixel value after the filter operation.

[0027] For example, in determining the filter characteristics, an asymmetric reference value is determined across the block boundary, and in the deblocking filter process, a filter operation using filter coefficients is performed. When the change amount of the pixel value before and after the filter operation exceeds the reference value, the change amount may be clipped to the reference value.

[0028] For example, in determining the filter characteristics, the condition for determining whether to perform the deblocking filter process may be set asymmetrically across the block boundary.

[0029] An encoding method according to one aspect of the present disclosure determines asymmetric filter characteristics across a block boundary and performs a deblocking filter process having the determined filter characteristics.

[0030] According to this, there is a possibility that the encoding method can reduce errors by performing a filter process having asymmetric filter characteristics across a block boundary.

[0031] A decoding method according to one aspect of the present disclosure determines asymmetric filter characteristics across a block boundary and performs a deblocking filter process having the determined filter characteristics.

[0032] According to this, there is a possibility that the decoding method can reduce errors by performing a filter process having asymmetric filter characteristics across a block boundary.

[0033] An encoding apparatus according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine asymmetric filter characteristics across a block boundary based on pixel values across the block boundary and performs a deblocking filter process having the determined filter characteristics.

[0034] According to this, the encoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across a block boundary. Further, the encoding device can determine appropriate filter characteristics based on pixel values across the block boundary.

[0035] For example, in determining the filter characteristics, the filter characteristics may be determined based on the difference between the pixel values.

[0036] For example, in determining the filter characteristics, as the difference between the pixel values is larger, the difference in the filter characteristics across the block boundary may be made larger.

[0037] According to this, the encoding device may be able to suppress unnecessary smoothing, for example, when the block boundary coincides with the edge of an object in the image.

[0038] For example, in determining the filter characteristics, the difference between the pixel values is compared with a threshold value based on a quantization parameter, and when the difference between the pixel values is larger than the threshold value, the difference in the filter characteristics across the block boundary may be made larger than when the difference between the pixel values is smaller than the threshold value.

[0039] According to this, the encoding device can determine filter characteristics taking into account the influence on the quantization parameter error.

[0040] For example, in determining the filter characteristics, as the difference between the pixel values is larger, the difference in the filter characteristics across the block boundary may be made smaller.

[0041] According to this, the decoding device may be able to suppress the smoothing from being weakened due to asymmetry, for example, when the block boundary is likely to be subjectively conspicuous, and may be able to suppress deterioration of subjective performance.

[0042] For example, in determining the filter characteristics, the filter characteristics may be determined based on the variance of the pixel values.

[0043] A decoding device according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on pixel values across the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0044] According to this, the decoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across a block boundary. Further, the decoding device may be able to determine appropriate filter characteristics based on the difference in pixel values across the block boundary.

[0045] For example, in determining the filter characteristics, the filter characteristics may be determined based on the difference in the pixel values.

[0046] For example, in determining the filter characteristics, as the difference in the pixel values is larger, the difference in the filter characteristics across the block boundary may be made larger.

[0047] According to this, the decoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Further, the decoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0048] For example, in determining the filter characteristics, the difference in the pixel values is compared with a threshold value based on a quantization parameter, and when the difference in the pixel values is larger than the threshold value, the difference in the filter characteristics across the block boundary is made larger than when the difference in the pixel values is smaller than the threshold value.

[0049] According to this, the decoding device can determine filter characteristics taking into account the influence on the errors of the quantization parameter.

[0050] For example, in determining the filter characteristics, the greater the difference in the pixel values, the smaller the difference in the filter characteristics across the block boundary may be made.

[0051] According to this, the decoding apparatus may be able to suppress unnecessary smoothing, for example, when the block boundary coincides with the edge of an object in the image.

[0052] For example, in determining the filter characteristics, the filter characteristics may be determined based on the variance of the pixel values.

[0053] An encoding method according to an aspect of the present disclosure determines non-symmetric filter characteristics across a block boundary based on pixel values across the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0054] According to this, the encoding method may be able to reduce errors by performing filter processing having non-symmetric filter characteristics across a block boundary. Also, the encoding method can determine appropriate filter characteristics based on the difference in pixel values across the block boundary.

[0055] A decoding method according to an aspect of the present disclosure determines non-symmetric filter characteristics across a block boundary based on pixel values across the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0056] According to this, the decoding method may be able to reduce errors by performing filter processing having non-symmetric filter characteristics across a block boundary. Also, the decoding method can determine appropriate filter characteristics based on the difference in pixel values across the block boundary.

[0057] An encoding device according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on an angle between a prediction direction of intra prediction and the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0058] According to this, the encoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across a block boundary. Also, the encoding device can determine appropriate filter characteristics based on an angle between a prediction direction of intra prediction and the block boundary.

[0059] For example, in determining the filter characteristics, the difference in the filter characteristics across the block boundary may be increased as the angle is closer to vertical.

[0060] According to this, the encoding device can increase the influence of filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the encoding device can reduce the influence of filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0061] For example, in determining the filter characteristics, the difference in the filter characteristics across the block boundary may be decreased as the angle is closer to horizontal.

[0062] According to this, the encoding device can increase the influence of filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the encoding device can reduce the influence of filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0063] A decoding device according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric with respect to a block boundary based on an angle between a prediction direction of intra prediction and the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0064] According to this, the decoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric with respect to a block boundary. Further, the decoding device may be able to determine appropriate filter characteristics based on an angle between a prediction direction of intra prediction and the block boundary.

[0065] For example, in determining the filter characteristics, the difference in the filter characteristics across the block boundary may be increased as the angle is closer to perpendicular.

[0066] According to this, the decoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the decoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0067] For example, in determining the filter characteristics, the difference in the filter characteristics across the block boundary may be decreased as the angle is closer to horizontal.

[0068] According to this, the decoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the decoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0069] An encoding method according to one aspect of the present disclosure determines filter characteristics that are asymmetric with respect to a block boundary based on an angle between a prediction direction of intra prediction and the block boundary, and performs deblocking filter processing having the determined filter characteristics.

[0070] According to this, the encoding method may be able to reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Further, the encoding method can determine appropriate filter characteristics based on the angle between the prediction direction of intra prediction and the block boundary.

[0071] A decoding method according to one aspect of the present disclosure determines filter characteristics that are asymmetric across the block boundary based on the angle between the prediction direction of intra prediction and the block boundary, and performs a deblocking filter process having the determined filter characteristics.

[0072] According to this, the decoding method may be able to reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Further, the decoding method can determine appropriate filter characteristics based on the angle between the prediction direction of intra prediction and the block boundary.

[0073] An encoding apparatus according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on the position within a block of a target pixel, and performs a deblocking filter process having the determined filter characteristics on the target pixel.

[0074] According to this, the encoding apparatus may be able to reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Further, the encoding apparatus can determine appropriate filter characteristics based on the position within a block of a target pixel.

[0075] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that pixels farther from the reference pixels of intra prediction are more greatly affected by the filtering process.

[0076] According to this, the encoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the encoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0077] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that the influence of the deblocking filter processing on the pixel in the lower right is greater than the influence of the deblocking filter processing on the pixel in the upper left.

[0078] According to this, the encoding device can increase the influence of the filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Also, the encoding device can reduce the influence of the filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0079] The decoding device according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on the position within a block of a target pixel, and performs deblocking filter processing having the determined filter characteristics on the target pixel.

[0080] According to this, the decoding device may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across a block boundary. Also, the decoding device can determine appropriate filter characteristics based on the position within a block of the target pixel.

[0081] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that the influence of the filter processing is greater for pixels farther from the reference pixels of the intra prediction.

[0082] According to this, since the decoding device can increase the influence of the filter process on pixels with large errors, there is a possibility that the errors of the pixels can be further reduced. Further, since the decoding device can reduce the influence of the filter process on pixels with small errors, there is a possibility that an increase in the errors of the pixels can be suppressed.

[0083] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that the influence of the deblocking filter process on the pixel in the lower right is greater than the influence of the deblocking filter process on the pixel in the upper left.

[0084] According to this, since the decoding device can increase the influence of the filter process on pixels with large errors, there is a possibility that the errors of the pixels can be further reduced. Further, since the decoding device can reduce the influence of the filter process on pixels with small errors, there is a possibility that an increase in the errors of the pixels can be suppressed.

[0085] The encoding method according to one aspect of the present disclosure determines filter characteristics that are asymmetric across a block boundary based on the position within a block of a target pixel, and performs a deblocking filter process having the determined filter characteristics on the target pixel.

[0086] According to this, the encoding method may be able to reduce errors by performing a filter process having filter characteristics that are asymmetric across a block boundary. Further, the encoding method can determine appropriate filter characteristics based on the position within a block of a target pixel.

[0087] The decoding method according to one aspect of the present disclosure determines filter characteristics that are asymmetric across a block boundary based on the position within a block of a target pixel, and performs a deblocking filter process having the determined filter characteristics on the target pixel.

[0088] According to this, there is a possibility that the decoding method can reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Also, the decoding method can determine appropriate filter characteristics based on the position of a target pixel within a block.

[0089] An encoding device according to one aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on quantization parameters, and performs a deblocking filter process having the determined filter characteristics.

[0090] According to this, there is a possibility that the encoding device can reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Also, the encoding device can determine appropriate filter characteristics based on quantization parameters.

[0091] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that the greater the quantization parameter, the greater the influence of the deblocking filter process.

[0092] According to this, the encoding device can increase the influence of the filtering process on pixels with large errors, so there is a possibility that the errors of such pixels can be further reduced. Also, the encoding device can reduce the influence of the filtering process on pixels with small errors, so there is a possibility that an increase in the errors of such pixels can be suppressed.

[0093] For example, in the determination of the filter characteristics, the filter characteristics may be determined such that the change in the influence accompanying a change in the quantization parameter of the upper left pixel is greater than the change in the influence accompanying a change in the quantization parameter of the lower right pixel.

[0094] According to this, the encoding device can increase the influence of the filtering process on pixels with large errors, so there is a possibility that the errors of such pixels can be further reduced.

[0095] A decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to determine filter characteristics that are asymmetric across a block boundary based on quantization parameters, and performs deblocking filter processing having the determined filter characteristics.

[0096] According to this, the decoding apparatus may be able to reduce errors by performing filter processing having filter characteristics that are asymmetric across a block boundary. Further, the decoding apparatus can determine appropriate filter characteristics based on quantization parameters.

[0097] For example, in determining the filter characteristics, the filter characteristics may be determined such that the greater the quantization parameter, the greater the influence of the deblocking filter processing.

[0098] According to this, the decoding apparatus can increase the influence of filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced. Further, the decoding apparatus can reduce the influence of filter processing on pixels with small errors, so there is a possibility that an increase in the errors of the pixels can be suppressed.

[0099] For example, in determining the filter characteristics, the filter characteristics may be determined such that a change in the influence accompanying a change in the quantization parameter of the upper left pixel is greater than a change in the influence accompanying a change in the quantization parameter of the lower right pixel.

[0100] According to this, the decoding apparatus can increase the influence of filter processing on pixels with large errors, so there is a possibility that the errors of the pixels can be further reduced.

[0101] An encoding method according to an aspect of the present disclosure may determine filter characteristics that are asymmetric across a block boundary based on quantization parameters, and perform deblocking filter processing having the determined filter characteristics.

[0102] According to this, the encoding method may be able to reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Also, the encoding method can determine appropriate filter characteristics based on quantization parameters.

[0103] A decoding method according to one aspect of the present disclosure determines filter characteristics that are asymmetric across a block boundary based on quantization parameters, and performs a deblocking filter process having the determined filter characteristics.

[0104] According to this, the decoding method may be able to reduce errors by performing a filtering process having filter characteristics that are asymmetric across a block boundary. Also, the decoding method can determine appropriate filter characteristics based on quantization parameters.

[0105] Note that these general or specific aspects may be implemented in a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, method, integrated circuit, computer program, and recording medium.

[0106] Hereinafter, embodiments will be specifically described with reference to the drawings.

[0107] Note that all of the embodiments described below show general or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Also, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.

[0108] (Embodiment 1) First, as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure to be described later are applicable, an outline of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.

[0109] When applying the processes and / or configurations described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.

[0110] (1) With respect to the encoding device or the decoding device of Embodiment 1, among the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure.

[0111] (2) With respect to the encoding device or the decoding device of Embodiment 1, after performing any changes such as addition, replacement, or deletion of functions or processes to be performed on some of the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure.

[0112] (3) With respect to the method performed by the encoding device or the decoding device of Embodiment 1, after performing addition of processes and / or any changes such as replacement or deletion of some of the plurality of processes included in the method, the process corresponding to the process described in each aspect of the present disclosure is replaced with the process described in each aspect of the present disclosure.

[0113] (4) Combine some of the components that make up the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components that include some of the functions of the components described in each aspect of the present disclosure, or components that perform some of the processes performed by the components described in each aspect of the present disclosure.

[0114] (5) Combine a component that includes some of the functions of some of the components that make up the encoding device or decoding device of Embodiment 1, or a component that performs some of the processes performed by some of the components that make up the encoding device or decoding device of Embodiment 1, with the components described in each aspect of the present disclosure, components that include some of the functions of the components described in each aspect of the present disclosure, or components that perform some of the processes performed by the components described in each aspect of the present disclosure.

[0115] (6) Replace, in the method performed by the encoding device or decoding device of Embodiment 1, the process corresponding to the process described in each aspect of the present disclosure among the plurality of processes included in the method with the process described in each aspect of the present disclosure.

[0116] (7) Combine some of the processes included in the method performed by the encoding device or decoding device of Embodiment 1 with the processes described in each aspect of the present disclosure.

[0117] Note that the ways of implementing the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, they may be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Also, the processes and / or configurations described in different aspects may be implemented in combination.

[0118] [Overview of Encoding Device] First, the outline of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.

[0119] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0120] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0121] Hereinafter, each component included in the encoding device 100 will be described.

[0122] [Division Unit] The splitting unit 102 splits each picture included in the input moving picture into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (for example, 128x128). Such blocks of a fixed size are sometimes called Coding Tree Units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of a variable size (for example, 64x64 or less) based on recursive quadtree and / or binary tree block splitting. Such blocks of a variable size are sometimes called Coding Units (CUs), Prediction Units (PUs), or Transformation Units (TUs). Note that in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be processing units of CUs, PUs, and TUs.

[0123] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, solid lines represent block boundaries by quadtree block splitting, and broken lines represent block boundaries by binary tree block splitting.

[0124] Here, block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).

[0125] The upper left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper left 64x64 block is split into two 16x64 blocks 11, 12 and a 32x64 block 13.

[0126] The upper right 64x64 block is horizontally split into two rectangular 64x32 blocks 14, 15 (binary tree block splitting).

[0127] The 64x64 block in the lower left is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the 64x64 block in the lower left is divided into 16 16x32 blocks, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.

[0128] The 64x64 block 23 in the lower right is not divided.

[0129] As described above, in FIG. 2, the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.

[0130] In FIG. 2, one block is divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary-tree block division). A division including such a ternary-tree block division is sometimes called MBT (multi type tree) division.

[0131] [Subtraction unit] The subtraction unit 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0132] The original signal is the input signal of the encoding device 100, and is a signal representing the image of each picture constituting the moving image (for example, a luminance signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0133] [Conversion Unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain, and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0134] Note that the conversion unit 106 may adaptively select a conversion type from a plurality of conversion types, and convert the prediction error into conversion coefficients using a transform basis function corresponding to the selected conversion type. Such a conversion may be called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).

[0135] The plurality of conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the conversion basis functions corresponding to each conversion type. In FIG. 3, N indicates the number of input pixels. The selection of the conversion type from among these plurality of conversion types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.

[0136] Information indicating whether to apply such EMT or AMT (for example, called an AMT flag) and information indicating the selected conversion type are signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0137] Further, the conversion unit 106 may re-convert the conversion coefficient (conversion result). Such re-conversion may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (e.g., 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of these pieces of information is not necessarily limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0138] Here, the separable transform is a method of performing multiple conversions by separating for each direction by the number of dimensions of the input, and the non-separable transform is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing the conversion collectively.

[0139] For example, as an example of the non-separable transform, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.

[0140] Similarly, after regarding a 4×4 input block as an array having 16 elements, a method of performing a plurality of Givens rotations on the array (Hypercube Givens Transform) is also an example of the non-separable transform.

[0141] [Quantization Unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order, and quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0142] The predetermined order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0143] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0144] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable-length encoding on the quantization coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.

[0145] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantization coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized conversion coefficients of the current block to the inverse conversion unit 114.

[0146] [Inverse Conversion Unit] The inverse transform unit 114 restores the prediction error by inversely transforming the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients, thereby restoring the prediction error of the current block. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0147] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.

[0148] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.

[0149] [Block memory] The block memory 118 is a storage unit for storing blocks within the coded target picture (hereinafter referred to as the current picture), which are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.

[0150] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0151] In ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 sub-block within a current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.

[0152] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (e.g., C = 5D + A) is calculated using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (e.g., 15 or 25 classes).

[0153] The gradient direction value D is derived, for example, by comparing gradients in a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding gradients in a plurality of directions and quantizing the addition result.

[0154] Based on the results of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0155] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Note that the signaling of information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0156] The on / off of ALF is determined, for example, at the picture level or CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance difference, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or CU level. Note that the signaling of the information indicating the on / off of ALF does not have to be limited to the picture level or CU level, and it may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0157] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets does not have to be limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0158] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0159] [Intra prediction unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to the blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0160] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0161] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).

[0162] The plurality of directional prediction modes includes, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0163] In the intra prediction of a chrominance block, a luminance block may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0164] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of reference pixels in the horizontal / vertical direction. Intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of PDPC application (for example, called a PDPC flag) is signaled at, for example, the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0165] [Inter prediction unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. Inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0166] The motion information used for motion compensation is signaled. A motion vector predictor may be used for signaling the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.

[0167] In addition to the motion information of the current block obtained by motion search, the motion information of adjacent blocks may also be used to generate an inter prediction signal. Specifically, an inter prediction signal may be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).

[0168] In such an OBMC mode, information indicating the size of sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).

[0169] The OBMC mode will be described more specifically. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process by OBMC processing.

[0170] First, a prediction image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.

[0171] Next, the motion vector (MV_L) of the encoded left adjacent block is applied to the block to be coded to obtain a prediction image (Pred_L), and the first correction of the prediction image is performed by weighting and superimposing the prediction image and Pred_L.

[0172] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_U). The predicted image after the first correction and Pred_U are weighted and superimposed to perform the second correction of the predicted image, which is used as the final predicted image.

[0173] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to configure to perform more corrections than two stages using the right adjacent block or the lower adjacent block.

[0174] Note that the region for superimposition may be only a partial region near the block boundary, rather than the pixel region of the entire block.

[0175] Here, the predicted image correction process from a single reference picture has been described. However, the same applies when correcting the predicted image from multiple reference pictures. After obtaining the predicted images corrected from each reference picture, the obtained predicted images are further superimposed to obtain the final predicted image.

[0176] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.

[0177] As a method for determining whether to apply OBMC processing, for example, there is a method using an obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as the obmc_flag and encoding is performed by applying OBMC processing. If it does not belong to a region with complex motion, the value 0 is set as the obmc_flag and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC processing according to the value.

[0178] Note that the motion information may be derived on the decoder side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also for example, the motion information may be derived by performing motion search on the decoder side. In this case, the motion search is performed without using the pixel values of the current block.

[0179] Here, a mode in which motion search is performed on the decoder side will be described. This mode in which motion search is performed on the decoder side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.

[0180] An example of the FRUC process is shown in FIG. 5D. First, by referring to the motion vectors of the encoded blocks adjacent to the current block spatially or temporally, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0181] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, a motion vector for the current block may be derived. That is, search is performed in the same way for the region around the best candidate MV, and if there is an MV for which the evaluation value is a good value, the best candidate MV may be updated to the MV, and that may be used as the final MV of the current block. Note that it is also possible to adopt a configuration in which the said process is not performed.

[0182] The same processing may be performed in the case of processing in sub-block units.

[0183] The evaluation value is calculated by obtaining the difference value of the reconstructed image by pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that the evaluation value may be calculated using information other than the difference value.

[0184] As the pattern matching, first pattern matching or second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.

[0185] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the candidate described above, a region in another reference picture along the motion trajectory of the current block is used.

[0186] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is advisable to select the candidate MV with the best evaluation value among the plurality of candidate MVs as the final MV.

[0187] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0188] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidate.

[0189] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that most closely matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or either of the left adjacent and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.

[0190] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the pattern matching method (first pattern matching or second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0191] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called the BIO (bi-directional optical flow) mode.

[0192] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v y( ) indicates the velocity vector, and τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.

[0193] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0194]

Equation

[0195] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from the merge list, etc., are corrected in pixel units.

[0196] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0197] Here, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.

[0198] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, a current block includes 16 4×4 sub-blocks. Here, based on the motion vectors of adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (v x , v y ) of each sub-block within the current block is derived by the following equation (2).

[0199] [Equation]

[0200] Here, x and y respectively indicate the horizontal position and the vertical position of the sub-block, and w indicates a predetermined weight coefficient.

[0201] Such an affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right control points are different. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode is not necessarily limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0202] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0203] Here, an example of deriving the motion vector of the picture to be encoded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.

[0204] First, a prediction MV list in which candidates for the prediction MV are registered is generated. As candidates for the prediction MV, there are a spatial adjacent prediction MV which is the MV of a plurality of encoded blocks located spatially adjacent to the block to be encoded, a temporal adjacent prediction MV which is the MV of a nearby block obtained by projecting the position of the block to be encoded in the encoded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV which is an MV with a value of zero, and the like.

[0205] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be encoded.

[0206] Furthermore, in the variable length encoding unit, a merge_idx which is a signal indicating which prediction MV has been selected is described in the stream and encoded.

[0207] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are merely examples, and the number may be different from that in the figure, the configuration may not include some of the types of prediction MVs in the figure, or the configuration may include prediction MVs other than the types of prediction MVs in the figure.

[0208] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be encoded derived in the merge mode.

[0209] Here, an example of determining the MV using the DMVR process will be described.

[0210] FIG. 9C is a conceptual diagram for explaining the outline of DMVR processing.

[0211] First, using the optimal MVP set for the processing target block as a candidate MV, reference pixels are respectively obtained from a first reference picture which is a processed picture in the L0 direction and a second reference picture which is a processed picture in the L1 direction according to the candidate MV, and a template is generated by taking the average of each reference pixel.

[0212] Next, using the template, the peripheral regions of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search region, the MV value, and the like.

[0213] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.

[0214] Note that even if it is not the processing itself described here, other processing may be used as long as it is a processing capable of searching the periphery of the candidate MV to derive the final MV.

[0215] Here, a mode of generating a predicted image using the LIC processing will be described.

[0216] FIG. 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction processing by the LIC processing.

[0217] First, an MV for obtaining a reference image corresponding to the encoding target block is derived from a reference picture which is an encoded picture.

[0218] Next, for the block to be encoded, information indicating how the luminance values change between the reference picture and the picture to be encoded is extracted using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, and the luminance correction parameter is calculated.

[0219] By performing luminance correction processing on the reference image in the reference picture specified by the MV using the luminance correction parameter, a predicted image for the block to be encoded is generated.

[0220] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0221] Also, although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures. After performing luminance correction processing on the reference images obtained from each reference picture in the same manner, a predicted image is generated.

[0222] As a method for determining whether to apply the LIC process, for example, there is a method using lic_flag, which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply the LIC process according to the value.

[0223] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the peripheral encoded blocks selected during the derivation of the MV in the merge mode process are encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. Note that in the case of this example, the processing in decoding is exactly the same.

[0224] [Outline of Decoder] Next, an outline of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing the functional configuration of a decoder 200 according to Embodiment 1. The decoder 200 is a moving image / image decoder that decodes moving images / images in units of blocks.

[0225] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0226] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0227] Each component included in the decoder 200 will be described below.

[0228] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bit stream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in block units.

[0229] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0230] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.

[0231] For example, when the information decoded from the encoded bit stream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.

[0232] Also, for example, when the information decoded from the encoded bit stream indicates that NSST is to be applied, the inverse transform unit 206 applies an inverse reverse transform to the transform coefficients.

[0233] [Addition Unit] The adder 208 reconstructs the current block by adding the prediction error which is the input from the inverse transform unit 206 and the prediction sample which is the input from the predictive control unit 220. Then, the adder 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0234] [Block Memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture), which are blocks referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0235] [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.

[0236] When the information indicating the on / off of the ALF read from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0237] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0238] [Intra Prediction Unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bitstream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0239] In addition, when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance difference block is selected, the intra prediction unit 216 may predict the chrominance difference component of the current block based on the luminance component of the current block.

[0240] Also, when the information decoded from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction.

[0241] [Inter prediction unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bitstream, and outputs the inter prediction signal to the prediction control unit 220.

[0242] In addition, when the information decoded from the encoded bitstream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.

[0243] Also, when the information decoded from the encoded bitstream indicates that the FRUC mode is applicable, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0244] In addition, when the BIO mode is applicable, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is applicable, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0245] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208.

[0246] [Deblocking filter processing] Next, the deblocking filter processing performed in the encoding device 100 and the decoding device 200 configured as described above will be specifically described with reference to the drawings. In the following, the operation of the loop filter unit 120 provided in the encoding device 100 will be mainly described, but the operation of the loop filter unit 212 provided in the decoding device 200 is the same.

[0247] As described above, when encoding an image, the encoding device 100 calculates a prediction error by subtracting a prediction signal generated by the intra prediction unit 124 or the inter prediction unit 126 from the original signal. The encoding device 100 generates quantization coefficients by performing orthogonal transformation processing, quantization processing, etc. on the prediction error. Further, the encoding device 100 restores the prediction error by inverse quantization and inverse orthogonal transformation of the obtained quantization coefficients. Here, since the quantization process is an irreversible process, the restored prediction error has an error (quantization error) with respect to the prediction error before the transformation.

[0248] The deblocking filter process performed by the loop filter unit 120 is a type of filter process performed for purposes such as reducing this quantization error. The deblocking filter process is applied to the block boundary to remove block noise. Hereinafter, this deblocking filter process will also be simply referred to as a filter process.

[0249] FIG. 11 is a flowchart showing an example of the deblocking filter process performed by the loop filter unit 120. For example, the process shown in FIG. 11 is performed for each block boundary.

[0250] First, the loop filter unit 120 calculates a block boundary strength (Bs) in order to determine the behavior of the deblocking filter process (S101). Specifically, the loop filter unit 120 determines Bs using the prediction mode of the block to be filtered, the nature of the motion vector, etc. For example, if at least one of the blocks sandwiching the boundary is an intra prediction block, Bs = 2 is set. Also, if at least one of the following conditions (1) to (3) is satisfied: (1) at least one of the blocks sandwiching the block boundary includes dominant orthogonal transformation coefficients, (2) the difference between the motion vectors of the two blocks sandwiching the block boundary is greater than or equal to a threshold value, and (3) the number of motion vectors or the reference images of the two blocks sandwiching the block boundary are different, then Bs = 1 is set. If none of the conditions (1) to (3) apply, Bs = 0 is set.

[0251] Next, the loop filter unit 120 determines whether the set Bs is greater than the first threshold (S102). If Bs is less than or equal to the first threshold (No in S102), the loop filter unit 120 does not perform the filter process (S107).

[0252] On the other hand, if the set Bs is greater than the first threshold (Yes in S102), the loop filter unit 120 calculates the pixel variation d in the boundary region using the pixel values in the blocks on both sides of the block boundary (S103). This process will be described with reference to FIG. 12. When the pixel values at the block boundary are defined as shown in FIG. 12, the loop filter unit 120 calculates, for example, d = |p30 - 2×p20 + p10| + |p33 - 2×p23 + p13| + |q30 - 2×q20 + q10| + |q33 - 2×q23 + q13|.

[0253] Next, the loop filter unit 120 determines whether the calculated d is greater than the second threshold (S104). If d is less than or equal to the second threshold (No in S104), the loop filter unit 120 does not perform the filter process (S107). Note that the first threshold and the second threshold are different.

[0254] If the calculated d is greater than the second threshold (Yes in S104), the loop filter unit 120 determines the filter characteristics (S105) and performs the filter process with the determined filter characteristics (S106). For example, a 5-tap filter such as (1, 2, 2, 2, 1) / 8 is used. That is, for p10 shown in FIG. 12, an operation of (1×p30 + 2×p20 + 2×p10 + 2×q10 + 1×q20) / 8 is performed. Here, during the filter process, a clip process is performed so that the displacement is within a certain range so as not to cause excessive smoothing. The clip process mentioned here means, for example, that when the threshold of the clip process is tc and the pixel value before filtering is q, the pixel value after filtering can only take values in the range of q ± tc.

[0255] Next, an example of applying an asymmetric filter across a block boundary in the deblocking filter process by the loop filter unit 120 according to this embodiment will be described.

[0256] FIG. 13 is a flowchart showing an example of the deblocking filter process according to this embodiment. Note that the process shown in FIG. 13 may be performed for each block boundary or for each unit pixel including one or more pixels.

[0257] First, the loop filter unit 120 acquires encoding parameters, and determines asymmetric filter characteristics across a block boundary using the acquired encoding parameters (S111). In the present disclosure, it is assumed that the acquired encoding parameters characterize, for example, an error distribution.

[0258] Here, the filter characteristics are filter coefficients and parameters used for controlling the filter process, etc. Also, the encoding parameters may be any parameters that can be used to determine the filter characteristics. The encoding parameters may be information indicating the error itself, or information or parameters related to the error (e.g., affecting the magnitude relationship of the error).

[0259] Also, hereinafter, a pixel determined to have a large or small error based on the encoding parameters, that is, a pixel highly likely to have a large or small error, will simply be referred to as a pixel with a large or small error.

[0260] Here, it is not necessary to perform the determination process each time, and the process may be performed according to a rule that associates the encoding parameters and the filter characteristics determined in advance.

[0261] Note that even a pixel that is statistically likely to have a small error may have a larger error than the error of a pixel that is likely to have a large error when viewed pixel by pixel.

[0262] Next, the loop filter unit 120 executes filter processing having the determined filter characteristics (S112).

[0263] Here, the filter characteristics determined in step S111 do not necessarily have to be asymmetric, and a symmetric design is also possible. In the following, a filter having asymmetric filter characteristics across a block boundary is also referred to as an asymmetric filter, and a filter having symmetric filter characteristics across a block boundary is also referred to as a symmetric filter.

[0264] Specifically, the filter characteristics are determined in consideration of two points: pixels determined to have a small error are less likely to be affected by pixels with a large error in the surroundings, and pixels determined to have a large error are more likely to be affected by pixels with a small error in the surroundings. That is, the filter characteristics are determined such that the greater the error of a pixel, the greater the influence of the filter processing. For example, the filter characteristics are determined such that the greater the error of a pixel, the greater the amount of change in the pixel values before and after the filter processing. Thereby, for pixels that are likely to have a small error, it is possible to prevent the values from deviating from the true value by large fluctuations. Conversely, for pixels that are likely to have a large error, the error can be reduced by strongly being affected by pixels with a small error and varying the value.

[0265] In the following, an element that changes the displacement by the filter is defined as the weight of the filter. In other words, the weight indicates the degree of influence of the filter processing on symmetric pixels. Increasing the weight means that the influence of the filter processing on the pixel increases. In other words, it means making the pixel value after the filter processing more likely to be affected by other pixels. Specifically, increasing the weight means determining the filter characteristics such that the amount of change in the pixel values before and after the filter processing increases, or such that the filter processing is more likely to be performed.

[0266] That is, the loop filter unit 120 increases the weight for pixels with larger errors. Note that increasing the weight for pixels with larger errors does not necessarily mean continuously changing the weight based on the error, but also includes cases where the weight is changed stepwise. That is, the weight of the first pixel only needs to be smaller than the weight of the second pixel with a larger error than the first pixel. The same expression will be used hereinafter.

[0267] Note that in the finally determined filter characteristics, it is not necessary that pixels with larger errors have larger weights. That is, the loop filter unit 120 may, for example, correct the reference filter characteristics determined by a conventional method so that pixels with larger errors tend to have larger weights.

[0268] Hereinafter, a plurality of specific methods for asymmetrically changing the weight will be described. Note that any one of the methods shown below may be used, or a method combining a plurality of methods may be used.

[0269] As a first method, the loop filter unit 120 decreases the filter coefficient for pixels with larger errors. For example, the loop filter unit 120 decreases the filter coefficient for pixels with larger errors and increases the filter coefficient for pixels with smaller errors.

[0270] An example of deblocking filter processing performed on pixel p1 shown in FIG. 12 will be described. Hereinafter, without applying this method, for example, a filter determined by a conventional method is called a reference filter. The reference filter is a 5-tap filter perpendicular to the block boundary, and is assumed to be a filter extending over (p3, p2, p1, q1, q2). Also, assume that the filter coefficients are (1, 2, 2, 2, 1) / 8. Also, assume that the error of block P is likely to be large and the error of block Q is likely to be small. In this case, the filter coefficients are set so that block P with a large error is likely to be affected by block Q with a small error. Specifically, the filter coefficients used for pixels with small errors are set large, and the filter coefficients used for pixels with large errors are set small. For example, (0.5, 1.0, 1.0, 2.0, 1.5) / 6 is used as the filter coefficients.

[0271] As another example, 0 may be used for the filter coefficients of pixels with small errors. For example, (0, 0, 1, 2, 2) / 5 may be used as the filter coefficients. That is, the filter taps may be changed. Conversely, the current filter coefficient of 0 may be set to a value other than 0. For example, (1, 2, 2, 2, 1, 1) / 9 etc. may be used as the filter coefficients. That is, the loop filter unit 120 may extend the filter taps to the side with a small error.

[0272] Note that the reference filter does not have to be a filter symmetric about the target pixel like (1, 2, 2, 2, 1) / 8 above. In such a case, the loop filter unit 120 further adjusts the filter. For example, the filter coefficients of the reference filter used for the leftmost pixel of block Q are (1, 2, 3, 4, 5) / 15, and the filter coefficients of the reference filter used for the rightmost pixel of block P are (5, 4, 3, 2, 1) / 15. That is, in this case, filter coefficients that are horizontally reversed are used between pixels sandwiching the block boundary. Such a filter characteristic that is symmetric with respect to the block boundary can also be called "a filter characteristic that is symmetric with respect to the block boundary". That is, a filter characteristic that is asymmetric with respect to the block boundary is a filter characteristic that is not symmetric with respect to the block boundary.

[0273] Also, in the same manner as described above, when the error of block P is large and the error of block Q is small, the loop filter unit 120 changes, for example, the filter coefficient (5, 4, 3, 2, 1) / 15, which is the filter coefficient of the reference filter used for the pixel at the right end of block P, to (2.5, 2.0, 1.5, 2.0, 1.0) / 9.

[0274] In this way, in the deblocking filter process, a filter whose filter coefficients change asymmetrically across the block boundary is used. For example, the loop filter unit 120 determines a reference filter having symmetric filter characteristics across the block boundary according to a predetermined standard. The loop filter unit 120 changes the reference filter so as to have asymmetric filter characteristics across the block boundary. Specifically, the loop filter unit 120 performs at least one of increasing the filter coefficient of at least one pixel with a small error and decreasing the filter coefficient of at least one pixel with a large error among the filter coefficients of the reference filter.

[0275] Next, a second method of changing the weights asymmetrically will be described. First, the loop filter unit 120 performs a filter operation using the reference filter. Next, the loop filter unit 120 performs asymmetric weighting across the block boundary on the reference change amount Δ0, which is the amount of change in the pixel values before and after the filter operation using the reference filter. Hereinafter, for the sake of distinction, the process using the reference filter is called a filter operation, and a series of processes including the filter operation and the subsequent correction process (for example, asymmetric weighting) is called a filter process (deblocking filter process).

[0276] For example, for pixels with small errors, the loop filter unit 120 calculates the corrected change amount Δ1 by multiplying the reference change amount Δ0 by a coefficient smaller than 1. Also, for pixels with large errors, the loop filter unit 120 calculates the corrected change amount Δ1 by multiplying the reference change amount Δ0 by a coefficient larger than 1. Next, the loop filter unit 120 generates the pixel value after the filter process by adding the corrected change amount Δ1 to the pixel value before the filter operation. Note that the loop filter unit 120 may perform only one of the process for pixels with small errors and the process for pixels with large errors.

[0277] For example, similar to the above, assume that the error of block P is large and the error of block Q is small. In this case, for the pixels included in the block Q with small errors, the loop filter unit 120 calculates the corrected change amount Δ1 by, for example, multiplying the reference change amount Δ0 by 0.8. Also, for the pixels included in the block P with large errors, the loop filter unit 120 calculates the corrected change amount Δ1 by, for example, multiplying the reference change amount Δ0 by 1.2. By doing so, the variation of the values of pixels with small errors can be reduced. Also, the variation of the values of pixels with large errors can be increased.

[0278] Note that as the ratio between the coefficient multiplied by the reference change amount Δ0 of pixels with small errors and the coefficient multiplied by the reference change amount Δ0 of pixels with large errors, 1:1 may be selected. In this case, the filter characteristics are symmetric across the block boundary.

[0279] Further, the loop filter unit 120 may calculate a coefficient to be multiplied by the reference change amount Δ0 by multiplying a reference coefficient by a constant. In this case, the loop filter unit 120 uses a larger constant for pixels with larger errors than for pixels with smaller errors. As a result, the amount of change in the pixel value for pixels with larger errors increases, and the amount of change in the pixel value for pixels with a high possibility of having smaller errors decreases. For example, the loop filter unit 120 uses 1.2 or 0.8 as the constant for pixels adjacent to the block boundary, and uses 1.1 or 0.9 as the constant for pixels one pixel away from the pixels adjacent to the block boundary. Further, the reference coefficient is obtained, for example, by (A×(q1 - p1)-B×(q2 - p2)+C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Further, p1, p2, q1, and q2 are the pixel values of pixels having the positional relationship shown in FIG. 12 across the block boundary.

[0280] Next, a third method of changing weights asymmetrically will be described. The loop filter unit 120 performs a filter operation using the filter coefficients of the reference filter in the same manner as in the second method. Next, the loop filter unit 120 adds an offset value that is asymmetric across the block boundary to the pixel value after the filter operation. Specifically, the loop filter unit 120 adds a positive offset value to the pixel value of a pixel with a large error so that the value of the pixel with a large error approaches the value of a pixel with a high possibility of having a small error and the displacement of the pixel with a large error increases. Also, the loop filter unit 120 adds a negative offset to the pixel value of a pixel with a small error so that the value of the pixel with a small error does not approach the value of a pixel with a large error and the displacement of the pixel with a small error decreases. As a result, the amount of change in the pixel value for pixels with larger errors increases, and the amount of change in the pixel value for pixels with smaller errors decreases. Note that the loop filter unit 120 may perform only one of the processing for pixels with small errors and the processing for pixels with large errors.

[0281] For example, for pixels included in a block with a large error, the loop filter unit 120 calculates a corrected change amount Δ1 by adding a positive offset value (e.g., 1) to the absolute value of the reference change amount Δ0. Also, for pixels included in a block with a small error, the loop filter unit 120 calculates a corrected change amount Δ1 by adding a negative offset value (e.g., -1) to the absolute value of the reference change amount Δ0. Next, the loop filter unit 120 generates a pixel value after filter processing by adding the corrected change amount Δ1 to the pixel value before filter operation. Note that the loop filter unit 120 may add the offset value not to the change amount but to the pixel value after filter operation. Also, the offset value may not be symmetric across the block boundary.

[0282] Also, when the filter taps extend over a plurality of pixels from the block boundary, the loop filter unit 120 may change only the weight for a certain specific pixel or may change the weights for all pixels. Also, the loop filter unit 120 may change the weights according to the distance from the block boundary to the target pixel. For example, the loop filter unit 120 may make the filter coefficients for up to two pixels from the block boundary asymmetric and make the filter coefficients for the subsequent pixels symmetric. Also, the weights of the filter may be common to a plurality of pixels or may be set for each pixel.

[0283] Next, a fourth method of changing the weights asymmetrically will be described. The loop filter unit 120 performs a filter operation using the filter coefficients of the reference filter. Next, when the change amount Δ between the pixel values before and after the filter operation exceeds a clip width that is a reference value, the loop filter unit 120 clips the change amount Δ to the clip width. The loop filter unit 120 sets the clip width asymmetrically across the block boundary.

[0284] Specifically, the loop filter unit 120 makes the clip width for pixels with a large error larger than the clip width for pixels with a small error. For example, the loop filter unit 120 makes the clip width for pixels with a large error a constant multiple of the clip width for pixels with a small error. As a result of changing the clip width, the values of pixels with a small error cannot change significantly. Also, the values of pixels with a large error can change significantly.

[0285] Note that the loop filter unit 120 may adjust the absolute value of the clip width instead of specifying the ratio of the clip widths. For example, the loop filter unit 120 fixes the clip width for pixels with a large error to a multiple of a predetermined reference clip width. The loop filter unit 120 sets the ratio of the clip width for pixels with a small error to the clip width for pixels with a small error to 1.2:0.8. Specifically, for example, assume that the reference clip width is 10 and the change amount Δ before and after the filter operation is 12. In this case, when the reference clip width is used as it is, the change amount Δ is corrected to 10 by threshold processing. On the other hand, when the target pixel is a pixel with a large error, for example, the reference clip width is multiplied by 1.5. As a result, since the clip width becomes 15, threshold processing is not performed and the change amount Δ becomes 12.

[0286] Next, a fifth method of changing weights asymmetrically will be described. The loop filter unit 102 sets the condition for determining whether to perform filter processing asymmetrically across the block boundary. Here, the condition for determining whether to perform filter processing is, for example, the first threshold value or the second threshold value shown in FIG. 11.

[0287] Specifically, the loop filter unit 120 sets the condition so that filter processing is likely to be performed for pixels with a large error, and sets the condition so that filter processing is unlikely to be performed for pixels with a small error. For example, the loop filter unit 120 sets the threshold value for pixels with a small error higher than the threshold value for pixels with a large error. For example, the loop filter unit 120 sets the threshold value for pixels with a small error to a constant multiple of the threshold value for pixels with a large error.

[0288] Further, the loop filter unit 120 may not only specify the ratio of the threshold values, but also adjust the absolute value of the threshold values. For example, the loop filter unit 120 may fix the threshold value for pixels with small errors to a multiple of a predetermined reference threshold value, and set the ratio between the threshold value for pixels with small errors and the threshold value for pixels with large errors to 1.2:0.8.

[0289] Specifically, assume that the reference threshold value of the second threshold value in step 104 is 10 and d calculated from the pixel values in the block is 12. When the reference threshold value is used as the second threshold value as it is, it is determined that filter processing is to be performed. On the other hand, when the target pixel is a pixel with a small error, for example, a value obtained by multiplying the reference threshold value by 1.5 is used as the second threshold value. In this case, the second threshold value becomes 15, which is larger than d. Thereby, it is determined that filter processing is not to be performed.

[0290] Also, constants or the like indicating weights based on the errors used in the above first to fifth methods may be predetermined values in the encoding device 100 and the decoding device 200, or may be variable. Specifically, this constant is a coefficient multiplied by the filter coefficient in the first method or the filter coefficient of the reference filter, a coefficient multiplied by the reference change amount Δ0 in the second method or a constant multiplied by the reference coefficient, an offset value in the third method, a constant multiplied by the clip width or the reference clip width in the fourth method, and a constant multiplied by the threshold value or the reference threshold value in the fifth method, and the like.

[0291] When the constant is variable, information indicating the constant may be included in the bit stream as a parameter in, for example, sequence or slice units, and transmitted from the encoding device 100 to the decoding device 200. Note that the information indicating the constant may be information indicating the constant itself, or may be information indicating the ratio or difference from the reference value.

[0292] In addition, as a method of changing a coefficient or a constant according to an error, for example, there are a method of changing linearly, a method of changing in a quadratic function manner, a method of changing in an exponential function manner, or a method of using a look-up table showing the relationship between the error and the constant.

[0293] Also, when the error is above a reference or when the error is below the reference, a fixed value may be used as a constant. For example, when the error of the loop filter unit 120 is within a predetermined range, the variable is set to a first value, when the error is above the predetermined range, the variable is set to a second value, and when the error is within the predetermined range, the variable may be continuously changed from the first value to the second value according to the error.

[0294] Also, when the error of the loop filter unit 120 exceeds a predetermined reference, a symmetric filter (reference filter) may be used instead of using an asymmetric filter.

[0295] Also, when using a look-up table or the like, the loop filter unit 120 may hold both tables for large and small errors, or may hold only one table and calculate the other constant according to a predetermined rule from the content of the table.

[0296] As described above, the encoding device 100 and the decoding device 200 according to the present embodiment can reduce the error of the reconstructed image by using an asymmetric filter, so that the encoding efficiency can be improved.

[0297] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.

[0298] (Embodiment 2) In Embodiments 2 to 6, specific examples of the encoding parameters characterizing the above-described error distribution will be described. In the present embodiment, the loop filter unit 120 determines filter characteristics according to the position within the block of the target pixel.

[0299] FIG. 14 is a flowchart showing an example of the deblocking filter process according to the present embodiment. First, the loop filter unit 120 acquires information indicating the position within the block of the target pixel as an encoding parameter characterizing the error distribution. Based on the position, the loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary (S121).

[0300] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S122).

[0301] Here, pixels farther from the reference pixels for intra prediction are more likely to have a larger error than pixels closer to the reference pixels for intra prediction. Therefore, the loop filter unit 120 determines the filter characteristics so that the amount of change in the pixel values before and after the filter process becomes larger for pixels farther from the reference pixels for intra prediction.

[0302] For example, in the case of H.265 / HEVC or JEM, as shown in FIG. 15, pixels closer to the reference pixels are pixels existing in the upper left within the block, and pixels farther from the reference pixels are pixels existing in the lower right within the block. Therefore, the loop filter unit 120 determines the filter characteristics so that the weight of the pixel in the lower right within the block becomes larger than the weight of the pixel in the upper left.

[0303] Specifically, for pixels far from the reference pixels of intra prediction, the loop filter unit 120 determines filter characteristics so that the influence of the filter process becomes large as described in Embodiment 1. That is, the loop filter unit 120 increases the weight of pixels far from the reference pixels of intra prediction. Here, increasing the weight means, as described above, (1) reducing the filter coefficient, (2) increasing the filter coefficient of pixels sandwiching a boundary (i.e., pixels close to the reference pixels of intra prediction), (3) increasing the coefficient multiplied by the amount of change, (4) increasing the offset value of the amount of change, (5) increasing the clip width, and (6) modifying the threshold value so that the filter process is likely to be executed, by performing at least one of these. On the other hand, for pixels close to the reference pixels of intra prediction, the loop filter unit 120 determines filter characteristics so that the influence of the filter process becomes small. That is, the loop filter unit 120 reduces the weight of pixels close to the reference pixels of intra prediction. Here, reducing the weight means, as described above, (1) increasing the filter coefficient, (2) reducing the filter coefficient of pixels sandwiching a boundary (i.e., pixels close to the reference pixels of intra prediction), (3) reducing the coefficient multiplied by the amount of change, (4) reducing the offset value of the amount of change, (5) reducing the clip width, and (6) modifying the threshold value so that the filter process is unlikely to be executed, by performing at least one of these.

[0304] Note that when intra prediction is used, the above processing is performed, and it is not necessary to perform the above processing on blocks using inter prediction. However, since the nature of the intra prediction block may also be affected by inter prediction, the above processing may also be performed on the inter prediction block.

[0305] Further, the loop filter unit 120 may change the weight by arbitrarily specifying a position within a specific block. For example, as described above, the loop filter unit 120 may increase the weight of the pixel in the lower right of the block and decrease the weight of the pixel in the upper left of the block. Note that the loop filter unit 120 may change the weight by arbitrarily specifying a position within the block, not limited to the upper left and the lower right.

[0306] Also, as shown in FIG. 15, at the horizontal adjacent block boundary, the error of the left block increases and the error of the right block increases. Therefore, the loop filter unit 120 may increase the weight of the left block and decrease the weight of the right block with respect to the horizontal adjacent block boundary.

[0307] Similarly, at the vertical adjacent block boundary, the error of the upper block increases and the error of the lower block decreases. Therefore, the loop filter unit 120 may increase the weight of the upper block and decrease the weight of the lower block with respect to the vertical adjacent block boundary.

[0308] Further, the loop filter unit 120 may change the weight according to the distance from the reference pixel of the intra prediction. Also, the loop filter unit 120 may determine the weight in units of block boundaries or in units of pixels. The farther away from the reference pixel, the more likely the error is to increase. Therefore, the loop filter unit 120 determines the filter characteristics so that the weight gradient becomes steeper as the distance from the reference pixel increases. Also, the loop filter unit 120 determines the filter characteristics so that the weight gradient on the upper side of the right side of the block is gentler than the weight gradient on the lower side.

[0309] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0310] (Embodiment 3) In this embodiment, the loop filter unit 120 determines filter characteristics according to the orthogonal transform basis.

[0311] FIG. 16 is a flowchart showing an example of deblocking filter processing according to this embodiment. First, the loop filter unit 120 acquires information indicating the orthogonal transform basis used for the target block as an encoding parameter characterizing the error distribution. The loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary based on the orthogonal transform basis (S131).

[0312] Next, the loop filter unit 120 performs filter processing having the determined filter characteristics (S132).

[0313] The encoding device 100 selects one orthogonal transform basis, which is the transform basis when performing orthogonal transform, from among a plurality of candidates. The plurality of candidates includes, for example, a 0th-order transform basis such as DCT-II with a flat basis and a 0th-order transform basis such as DST-VII with a non-flat basis. FIG. 17 is a diagram showing the transform basis of DCT-II. FIG. 18 is a diagram showing the transform basis of DCT-VII.

[0314] The 0th-order basis of DCT-II is constant regardless of the position within the block. That is, when DCT-II is used, the error within the block is constant. Therefore, when both blocks sandwiching the block boundary are transformed by DCT-II, the loop filter unit 120 performs filter processing using a symmetric filter without using an asymmetric filter.

[0315] On the other hand, the value of the 0th basis of DST-VII increases as the distance from the left or upper block boundary increases. That is, the error is likely to increase as the distance from the left or upper block boundary increases. Therefore, when at least one of the two blocks sandwiching the block boundary is transformed by DST-VII, the loop filter unit 120 uses an asymmetric filter. Specifically, the loop filter unit 120 determines the filter characteristics so that the influence of the filter process becomes smaller for pixels with smaller values within the block of the lower-order (e.g., 0th order) basis.

[0316] Specifically, when both blocks sandwiching the block boundary are transformed by DST-VII, the loop filter unit 120 determines the filter characteristics so that the influence of the filter process becomes larger for the lower-right pixel within the block by the above-described method. Also, the loop filter unit 120 determines the filter characteristics so that the influence of the filter process becomes larger for the upper-left pixel within the block.

[0317] Also, when DST-VII and DCT-II are adjacent vertically, the loop filter unit 120 determines the filter characteristics so that the filter weight for the pixels in the lower part of the upper block where DST-VII is used, which is adjacent to the block boundary, becomes larger than the filter weight for the pixels in the upper part of the lower block where DCT-II is used. However, the difference in the amplitude of the lower-order basis in this case is smaller than the difference in the amplitude of the lower-order basis when DST-VIIs are adjacent. Therefore, the loop filter unit 120 sets the filter characteristics so that the slope of the weight in this case is smaller than the slope of the weight when DST-VIIs are adjacent. The loop filter unit 120 sets, for example, the weight when DCT-II and DCT-II are adjacent to 1:1 (symmetric filter), the weight when DST-VII and DST-VII are adjacent to 1.3:0.7, and the weight when DST-VII and DCT-II are adjacent to 1.2:0.8.

[0318] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, some processes described in the flowchart of this aspect, some configurations of the apparatus, some syntax, etc. may be implemented in combination with other aspects.

[0319] (Embodiment 4) In this embodiment, the loop filter unit 120 determines filter characteristics according to pixel values sandwiching a block boundary.

[0320] FIG. 19 is a flowchart showing an example of deblocking filter processing according to this embodiment. First, the loop filter unit 120 acquires information indicating pixel values in a block sandwiching a block boundary as an encoding parameter characterizing error distribution. Based on the pixel values, the loop filter unit 120 determines asymmetric filter characteristics across the block boundary (S141).

[0321] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S142).

[0322] For example, the loop filter unit 120 increases the difference in filter characteristics across the block boundary as the difference d0 in pixel values increases. Specifically, the loop filter unit 120 determines filter characteristics so that the difference in the influence of the filter process becomes large. For example, when d0 > (quantization parameter) × (constant) is satisfied, the loop filter unit 120 sets the weights to 1.4:0.6, and when the above relationship is not satisfied, sets the weights to 1.2:0.8. That is, the loop filter unit 120 compares the difference d0 in pixel values with a threshold based on the quantization parameter, and when the difference d0 in pixel values is larger than the threshold, makes the difference in filter characteristics across the block boundary larger than when the difference d0 in pixel values is smaller than the threshold.

[0323] As another example, for instance, the loop filter unit 120 increases the difference in filter characteristics across the block boundary as the average value b0 of the variances of the pixel values in both blocks sandwiching the block boundary is larger. Specifically, the loop filter unit 120 may determine the filter characteristics so that the difference in the influence of the filter process becomes larger. For example, when b0 > (quantization parameter) × (constant) is satisfied, the loop filter unit 120 sets the weights to 1.4:0.6, and when the above relationship is not satisfied, sets the weights to 1.2:0.8. That is, the loop filter unit 120 compares the variance b0 of the pixel values with a threshold value based on the quantization parameter, and when the variance b0 of the pixel values is larger than the threshold value, may increase the difference in filter characteristics across the block boundary more than when the variance b0 of the pixel values is smaller than the threshold value.

[0324] Note that, among adjacent blocks, which block's weight is to be increased, that is, which block has a larger error, can be specified by the method of Embodiment 2 or 3 described above, or the method of Embodiment 6 to be described later. That is, the loop filter unit 120 determines asymmetric filter characteristics across the block boundary according to a predetermined rule (for example, the method of Embodiment 2, 3, or 6). Next, the loop filter unit 120 changes the determined filter characteristics so that the difference in filter characteristics across the block boundary becomes larger based on the pixel value difference d0. That is, the loop filter unit 120 increases the ratio or difference between the weight of the pixel with a large error and the weight of the pixel with a small error.

[0325] Here, when the pixel value difference d0 is large, there is a possibility that the block boundary coincides with the edge of the object in the image. Therefore, in such a case, by reducing the difference in filter characteristics across the block boundary, it is possible to suppress unnecessary smoothing from being performed.

[0326] Note that, contrary to the above, the loop filter unit 120 may reduce the difference in filter characteristics across the block boundary as the difference d0 in pixel values increases. Specifically, the loop filter unit 120 determines the filter characteristics so that the difference in the influence of the filter process becomes smaller. For example, when d0 > (quantization parameter) × (constant) is satisfied, the loop filter unit 120 sets the weights to 1.2:0.8, and when the above relationship is not satisfied, it sets the weights to 1.4:0.6. Note that when the above relationship is satisfied, the weights may be set to 1:1 (symmetric filter). That is, the loop filter unit 120 compares the pixel value difference d0 with a threshold based on the quantization parameter, and when the pixel value difference d0 is greater than the threshold, it reduces the difference in filter characteristics across the block boundary more than when the pixel value difference d0 is less than the threshold.

[0327] For example, a large difference d0 in pixel values means that the block boundary is likely to be prominent. Therefore, in such a case, by reducing the difference in filter characteristics across the block boundary, it is possible to suppress the weakening of smoothing by the asymmetric filter.

[0328] Note that these two processes may be performed simultaneously. For example, when the pixel value difference d0 is less than the first threshold, the loop filter unit 120 uses the first weight, when the pixel value difference d0 is greater than or equal to the first threshold and less than the second threshold, it uses a second weight with a greater difference than the first weight, and when the pixel value difference d0 is greater than or equal to the second threshold, it may use a third weight with a smaller difference than the second weight.

[0329] Also, the pixel value difference d0 may be the difference in pixel values across the boundary itself, or the average or variance of the pixel value differences. For example, the pixel value difference d0 is obtained by (A×(q1 - p1) - B×(q2 - p2) + C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Also, p1, p2, q1, and q2 are the pixel values of the pixels in the positional relationship shown in FIG. 12 across the block boundary.

[0330] Note that the setting of the difference d0 of the pixel values and the weights may be performed on a pixel-by-pixel basis, on a block boundary unit basis, or on a block group unit basis including a plurality of blocks (for example, on an LCU (Largest Coding Unit) unit basis).

[0331] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Further, a part of the processes described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0332] (Embodiment 5) In this embodiment, the loop filter unit 120 determines filter characteristics according to the intra prediction direction and the block boundary direction.

[0333] FIG. 20 is a flowchart showing an example of deblocking filter processing according to this embodiment. First, the loop filter unit 120 acquires information indicating the angle between the prediction direction of intra prediction and the block boundary as a coding parameter characterizing the error distribution. Based on the angle, the loop filter unit 120 determines filter characteristics that are asymmetric with respect to the block boundary (S151).

[0334] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S152).

[0335] Specifically, the loop filter unit 120 increases the difference in filter characteristics across the block boundary as the angle is closer to vertical, and decreases the difference in filter characteristics across the block boundary as the angle is closer to horizontal. More specifically, when the intra prediction direction is close to perpendicular to the block boundary, the difference in the weights of the filters for the pixels on both sides across the block boundary becomes large, and when the intra prediction direction is close to horizontal with respect to the block boundary, the filter characteristics are determined so that the difference in the weights of the filters for the pixels on both sides across the block boundary becomes small. FIG. 21 is a diagram showing an example of weights with respect to the relationship between the intra prediction direction and the direction of the block boundary.

[0336] Note that, among adjacent blocks, which block's weight is increased, that is, which block has a larger error, can be specified by the method of Embodiment 2 or 3 described above, or the method of Embodiment 6 described later. That is, the loop filter unit 120 determines filter characteristics that are asymmetric across a block boundary according to a predetermined rule (for example, the method of Embodiment 2, 3, or 6). Next, the loop filter unit 120 changes the determined filter characteristics so that the difference in filter characteristics across the block boundary becomes large based on the intra prediction direction and the direction of the block boundary.

[0337] Also, the encoding device 100 and the decoding device 200 specify the intra prediction direction using, for example, the intra prediction mode.

[0338] Note that when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 does not need to consider the direction of the block boundary. For example, when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 may use a predetermined weight or a difference in weights regardless of the direction of the block boundary. Alternatively, when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 may use a symmetric filter.

[0339] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.

[0340] (Embodiment 6) In this embodiment, the loop filter unit 120 determines filter characteristics according to a quantization parameter indicating the width of quantization.

[0341] FIG. 16 is a flowchart showing an example of deblocking filter processing according to the present embodiment. First, the loop filter unit 120 acquires information indicating the quantization parameter used during quantization of the target block as an encoding parameter characterizing the error distribution. The loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary based on the quantization parameter (S161).

[0342] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S162).

[0343] Here, the larger the quantization parameter, the higher the likelihood of a larger error. Therefore, the loop filter unit 120 determines the filter characteristics such that the influence of the filter process becomes greater as the quantization parameter increases.

[0344] FIG. 23 is a diagram showing an example of weights with respect to the quantization parameter. As shown in FIG. 23, the loop filter unit 120 increases the weight for the upper left pixel in the block as the quantization parameter increases. On the other hand, the loop filter unit 120 reduces the increase in the weight for the lower right pixel in the block as the quantization parameter increases. That is, the loop filter unit 120 determines the filter characteristics such that the change in the influence of the filter process accompanying the change in the quantization parameter of the upper left pixel is greater than the change in the influence of the filter process accompanying the change in the quantization parameter of the lower right pixel.

[0345] Here, the upper left pixel in the block is more susceptible to the influence of the quantization parameter than the lower right pixel in the block. Therefore, by performing the above-described process, the error can be appropriately reduced.

[0346] Further, the loop filter unit 120 may determine the weight of each of two blocks sandwiching a boundary based on the quantization parameter of the block, or may calculate the average value of the quantization parameters of the two blocks and determine the weights of the two blocks based on the average value. Alternatively, the loop filter unit 120 may determine the weights of the two blocks based on the quantization parameter of one of the blocks. For example, the loop filter unit 120 uses the above method to determine the weight of one of the blocks based on the quantization parameter of the one block. Next, the loop filter unit 120 determines the weight of the other block according to a predetermined rule based on the determined weight.

[0347] Further, when the quantization parameters of the two blocks are different, or when the difference between the quantization parameters of the two blocks exceeds a threshold, the loop filter unit 120 may use a symmetric filter.

[0348] Also, in FIG. 23, the weight is set using a linear function, but any function other than the linear function or a table may be used. For example, a curve showing the relationship between the quantization parameter and the quantization step (quantization width) may be used.

[0349] Further, when the quantization parameter exceeds a threshold, the loop filter unit 120 may use a symmetric filter instead of an asymmetric filter.

[0350] Also, when the quantization parameter is described with decimal point precision, the loop filter unit 120 may perform operations such as rounding, ceiling, or truncation on the quantization parameter and use the quantized parameter after the operation for the above processing. Alternatively, the loop filter unit 120 may perform the above processing considering up to the decimal point unit.

[0351] As described above, in Embodiments 2 to 6, a plurality of methods for determining an error have been individually described, but two or more of these methods may be combined. In this case, the loop filter unit 120 may perform weighting on two or more combined elements.

[0352] Hereinafter, a modification will be described.

[0353] Examples other than the encoding parameters described above may be used. For example, the encoding parameters may be the type of orthogonal transformation (Wavelet, DFT, redundant transformation, etc.), the block size (the width and height of the block), the direction of the motion vector, the length of the motion vector, or the number of reference pictures used for inter prediction, information indicating the characteristics of the reference filter. Further, they may be used in combination. For example, the loop filter unit 120 may use an asymmetric filter only when the length of the block boundary is 16 pixels or less and the filter target pixel is close to the reference pixel for intra prediction, and use a symmetric filter in other cases. As another example, asymmetric processing may be performed only when a filter of a predetermined type among a plurality of filter candidates is used. For example, an asymmetric filter may be used only when the displacement by the reference filter is calculated by (A×(q1 - p1) - B×(q2 - p2) + C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Further, p1, p2, q1, and q2 are the pixel values of the pixels having the positional relationship shown in FIG. 12 across the block boundary.

[0354] Further, the loop filter unit 120 may perform the above processing on one of the luminance signal and the color difference signal, or may perform the above processing on both. Further, the loop filter unit 120 may perform common processing on the luminance signal and the color difference signal, or may perform different processing. For example, the loop filter unit 120 may use different weights for the luminance signal and the color difference signal, or may determine the weights according to different rules.

[0355] Further, various parameters used in the above processing may be determined in the encoding device 100, or may be preset fixed values.

[0356] Also, whether to perform the above processing, not to perform it, or the content of the above processing may be switched in predetermined units. The predetermined unit is, for example, a slice unit, a tile unit, a wavefront division unit, or a CTU unit. Further, the content of the above processing is which of the plurality of methods shown above to use, or a parameter indicating a weight or the like, or a parameter for determining these.

[0357] Also, the loop filter unit 120 may limit the area where the above processing is performed to the boundary of a CTU, the boundary of a slice, or the boundary of a tile.

[0358] Also, the number of taps of the filter may be different between the symmetric filter and the asymmetric filter.

[0359] Also, the loop filter unit 120 may change whether to perform the above processing or the content of the above processing according to the frame type (I frame, P frame, B frame).

[0360] Also, the loop filter unit 120 may determine whether to perform the above processing or the content of the above processing according to whether a specific processing in the previous stage or the subsequent stage has been performed.

[0361] Also, the loop filter unit 120 may perform different processes according to the type of prediction mode used for a block, or may perform the above process only for a block using a specific prediction mode. For example, the loop filter unit 120 may perform different processes for a block using intra prediction, a block using inter prediction, and a merged block.

[0362] Further, the encoding device 100 may encode filter information, which is a parameter indicating whether to perform the above processing or the content of the above processing. That is, the encoding device 100 may generate an encoded bitstream including the filter information. This filter information may include information indicating whether to perform the above processing on the luminance signal, information indicating whether to perform the above processing on the chrominance signal, or information indicating whether to perform different processing for each prediction mode.

[0363] Also, the decoding device 200 may perform the above processing based on the filter information included in the encoded bitstream. For example, the decoding device 200 may determine whether to perform the above processing or the content of the above processing based on the filter information.

[0364] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processing described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.

[0365] (Embodiment 7) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, etc. Also, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (a dedicated circuit).

[0366] Also, the processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing or distributed processing may be performed.

[0367] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible, and they are also included within the scope of the aspects of the present disclosure.

[0368] Furthermore, here, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, and an image encoding / decoding device having both. Other configurations in the system can be appropriately changed as the case may be.

[0369] [Usage Example] FIG. 24 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, ex110, which are fixed radio stations, are installed in each cell.

[0370] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above elements. The devices may be directly or indirectly connected to each other via a telephone network or short-range wireless etc. without going through the base stations ex106 to ex110, which are fixed radio stations. Also, the streaming server ex103 is connected to devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101 etc. Further, the streaming server ex103 is connected to terminals etc. within a hotspot in an airplane ex117 via a satellite ex116.

[0371] Note that, instead of the base stations ex106 to ex110, a wireless access point, a hot spot, or the like may be used. Further, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0372] The camera ex113 is a device capable of taking still images and videos such as a digital camera. Further, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handyphone System) or the like that generally supports the mobile communication system standards called 2G, 3G, 3.9G, 4G, and in the future, 5G.

[0373] The home appliance ex118 is a device included in a refrigerator or a household fuel cell cogeneration system.

[0374] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 through the base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal in the airplane ex117, etc.) performs the encoding process described in each of the above embodiments on the still image or video content photographed by the user using the terminal, multiplexes the video data obtained by encoding with the audio data obtained by encoding the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present disclosure.

[0375] On the one hand, the streaming server ex103 streams the transmitted content data to the requested client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc., which can decode the encoded data. Each device that receives the distributed data decodes and plays the received data. That is, each device functions as an image decoding device according to an aspect of the present disclosure.

[0376] [Distributed processing] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data in a distributed manner. For example, the streaming server ex103 may be realized by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world. In a CDN, an edge server physically close to the client is dynamically assigned according to the client. Then, by caching and distributing the content to the edge server, the delay can be reduced. Also, when an error occurs or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the distribution, so high-speed and stable distribution can be realized.

[0377] Furthermore, not only the distribution process itself can be decentralized, but the encoding process of the captured data can be performed on each terminal, on the server side, or they can be shared between each other. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image in terms of frames or scenes, or the amount of code is detected. Also, in the second loop, a process is performed to improve the encoding efficiency while maintaining the image quality. For example, if the terminal performs the first encoding process and the server that receives the content performs the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode almost in real time, the already encoded data from the first encoding performed by the terminal can be received and played back by other terminals, enabling more flexible real-time distribution.

[0378] As another example, cameras such as ex113 perform feature extraction from the image, compress the data related to the features as metadata, and transmit it to the server. The server performs compression according to the meaning of the image, such as judging the importance of the object from the features and switching the quantization accuracy. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) can be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed on the server.

[0379] As yet another example, in a stadium, shopping mall, or factory, etc., there may be a case where there are multiple video data captured by multiple terminals with almost the same scene. In this case, using the multiple terminals that performed the shooting, and other terminals and servers that did not perform shooting if necessary, encoding processes are respectively assigned and decentralized processing is performed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture. This can reduce the delay and achieve more real-time performance.

[0380] In addition, since the plurality of video data is in substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can be referenced to each other. Alternatively, the server may receive the encoded data from each terminal, change the reference relationship among the plurality of data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.

[0381] In addition, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based method, or convert H.264 to H.265.

[0382] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the entity performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.

[0383] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and use different scenes captured by terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal are integrated based on the relative positional relationship between the terminals obtained separately, or the regions where the feature points included in the videos match.

[0384] The server may not only encode two-dimensional moving images, but also automatically encode still images based on scene analysis of the moving images or at a time specified by the user, and transmit them to the receiving terminal. Further, when the server can acquire the relative positional relationship between the shooting terminals, it can generate the three-dimensional shape of the scene based on not only two-dimensional moving images but also videos shot from different angles of the same scene. Note that the server may separately encode three-dimensional data generated by a point cloud or the like, or select or reconstruct the video to be transmitted to the receiving terminal from the videos shot by a plurality of terminals based on the result of recognizing or tracking a person or an object using the three-dimensional data.

[0385] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or can also enjoy the content obtained by cutting out the video from an arbitrary viewpoint from the three-dimensional data reconstructed using a plurality of images or videos. Further, similar to the video, sound is also collected from a plurality of different angles, and the server may multiplex and transmit the sound from a specific angle or space together with the video according to the video.

[0386] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become widespread. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by Multi-View Coding (MVC) or the like, or may encode them as separate streams without referring to each other. At the time of decoding the separate streams, it is preferable to synchronize and reproduce them so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0387] In the case of an AR image, the server superimposes virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting virtual object information, the decoding device transmits the movement of the user's viewpoint to the server, and the server creates superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encodes the superimposed data, and distributes it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like and encode it in a state where the portion is transparent. Alternatively, the server may set the RGB value of a predetermined value as a background like a chroma key and generate data with the portion other than the object being the background color.

[0388] Similarly, the decoding process of the distributed data may be performed on each client terminal, on the server side, or shared between them. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request on another terminal, perform the decoding process, and transmit the decoded signal to the device having a display. By dispersing the process and selecting appropriate content regardless of the performance of the communicable terminal itself, it is possible to reproduce data with good image quality. Also, as another example, while receiving large-size image data on a TV or the like, a part of the area such as a tile in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to confirm the area to be in charge of or the area to be confirmed in more detail at hand.

[0389] In the future, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used regardless of indoors or outdoors, by using a distribution system standard such as MPEG-DASH, it is expected to seamlessly receive content while switching appropriate data for the ongoing communication. As a result, the user can freely select not only their own terminal but also a decoding device or display device such as a display installed indoors or outdoors and switch in real time. Also, based on their own location information, etc., decoding can be performed while switching the terminal to be decoded and the terminal to be displayed. This makes it possible to move while displaying map information on a part of the wall surface or ground of the adjacent building where a displayable device is embedded during movement to the destination. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or being copied to an edge server in a content delivery service, it is also possible to switch the bitrate of the received data.

[0390] [Scalable Encoding] Regarding content switching, it will be described using a scalable stream that is compression-encoded by applying the moving image encoding method shown in each of the above embodiments, as shown in FIG. 25. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of a temporally / spatially scalable stream realized by encoding in layers as shown in the figure, a configuration for switching content may be used. That is, by determining up to which layer to decode according to internal factors such as the performance on the decoding side and external factors such as the state of the communication bandwidth, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on a smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.

[0391] Furthermore, as described above, pictures are encoded for each layer, and in addition to the configuration that realizes scalability where an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of the image, etc., and the decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an enlargement of the resolution. The meta information includes information for specifying linear or non-linear filter coefficients used for super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used for super-resolution processing, etc.

[0392] Alternatively, the picture may be divided into tiles, etc. according to the meaning of objects in the image, etc., and the decoding side may decode only a part of the region by selecting the tiles to be decoded. Also, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 26, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0393] Also, the meta information may be stored in units composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining it with the information in picture units, can specify the picture in which the object exists and the position of the object in the picture.

[0394] [Optimization of Web Page] FIG. 27 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 28 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 27 and 28, a web page may include a plurality of link images that are links to image content, and the appearance thereof may vary depending on the device for viewing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) displays a still image or an I picture that each content has as a link image, displays a video like a gif animation with a plurality of still images or I pictures, etc., or receives only the base layer and decodes and displays the video.

[0395] When a link image is selected by the user, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Also, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay between the decoding time and the display time of the leading picture (the delay from the start of content decoding to the start of display) by decoding and displaying only the forward-reference pictures (I pictures, P pictures, B pictures with only forward reference). Further, the display device may boldly ignore the reference relationship of the pictures, roughly decode all B pictures and P pictures with forward reference, and perform normal decoding as the received pictures increase over time.

[0396] [Autonomous driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.

[0397] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal move, the receiving terminal can achieve seamless reception and decoding by transmitting the position information of the receiving terminal at the time of a reception request, while switching between the base stations ex106 to ex110. Further, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication band.

[0398] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information transmitted by the user in real time.

[0399] [Delivery of Personal Content] Further, in the content supply system ex100, not only high-quality and long-duration content by video delivery providers but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. Also, it is considered that such personal content will increase in the future. In order to make personal content into better content, the server may perform an encoding process after performing an editing process. This can be realized, for example, with the following configuration.

[0400] During shooting in real time or accumulating and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection on the original picture or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes unimportant scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone for editing. The server encodes the edited data based on the editing results. Also, it is known that if the shooting time is too long, the viewing rate will decrease. The server may automatically clip not only unimportant scenes but also scenes with little movement within a specific time range according to the shooting time so that the content is within that time range, based on the image processing results. Or, the server may generate and encode a digest based on the result of semantic analysis of the scene.

[0401] Note that in personal content, there may be cases where it contains elements that would directly violate copyright, moral rights of the author, or portrait rights as it is, and there may be inconveniences for individuals such as the sharing scope exceeding the intended scope. Therefore, for example, the server may deliberately change the image to be out of focus for the faces of people in the peripheral part of the screen or inside the house and then encode it. Also, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, the user designates a person or background area that the user wants to process the image from the perspective of copyright, etc., and the server can perform processing such as replacing the designated area with another video or blurring the focus. For a person, the video of the face part can be replaced while tracking the person in the moving image.

[0402] In addition, since viewing personal content with a small data volume requires strong real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and performs decoding and playback. During this time, the decoding device receives the enhancement layer and, when playback is looped or played back two or more times, such as when the enhancement layer is also included, high-quality video may be played back. For a stream encoded in a scalable manner like this, the video is rough when not selected or at the beginning of viewing, but it is possible to provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played back for the first time and a second stream encoded with reference to the first video are configured as one stream.

[0403] [Other Usage Examples] In addition, these encoding or decoding processes are generally processed by the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processing may be performed using that software. Furthermore, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.

[0404] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the content encoding method or has the ability to execute a specific service. If the terminal does not support the content encoding method or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays back the content.

[0405] Moreover, not only the content supply system ex100 via the Internet ex101, but also at least either the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can be incorporated into a digital broadcast system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by loading it on a broadcast radio wave using a satellite or the like, there is a difference in that it is suitable for multicast as opposed to a configuration in which unicast of the content supply system ex100 is easy, but the same application is possible with respect to the encoding process and the decoding process.

[0406] [Hardware Configuration] FIG. 29 is a diagram showing a smartphone ex115. FIG. 30 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by the camera unit ex465 and video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with a SIM ex468 for identifying a user and authenticating access to various data including a network. Note that an external memory may be used instead of the memory unit ex467.

[0407] In addition, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a bus ex470.

[0408] When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack to activate the smartphone ex115 in an operable state.

[0409] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after performing digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified, frequency conversion processing and analog-to-digital conversion processing are performed, spectral inverse spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog voice signal by the voice signal processing unit ex454, it is output from the voice output unit ex457. In the data communication mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body unit, and the same transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while the camera unit ex465 is imaging video or still images, etc., and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.

[0410] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and a video or a still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming is widespread, there may be a situation where it is not socially appropriate to play audio depending on the user's situation. Therefore, as an initial value, it is desirable to have a configuration that plays only video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.

[0411] Also, although the smartphone ex115 has been described as an example here, as a terminal, in addition to a transmission / reception type terminal having both an encoder and a decoder, there are three possible implementation forms: a transmission terminal having only an encoder, and a reception terminal having only a decoder. Furthermore, in the digital broadcast system, although it has been described as receiving or transmitting multiplexed data in which audio data and the like are multiplexed into video data, in the multiplexed data, character data related to the video or the like may be multiplexed in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0412] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, the terminal often has a GPU. Therefore, a configuration may be adopted in which a wide area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so as to be commonly used. As a result, the encoding time can be shortened, real-time performance can be ensured, and low latency can be realized. In particular, it is efficient to perform processes such as motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures using the GPU instead of the CPU.

[0413] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processes described in the flowchart of this aspect, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.

Industrial Applicability

[0414] The present disclosure can be applied to an encoding device, a decoding device, an encoding method, and a decoding method.

Explanation of Signs

[0415] 100 Encoding device 102 Division unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 200 Decoding device 202 Entropy Decoding Unit

Claims

1. A processing circuit; a memory coupled to the processing circuit; The processing circuitry uses the memory to: performing a luminance filtering process on a first boundary between the first luminance block and the second luminance block; performing a chrominance filtering process on a second boundary between the first chrominance block and the second chrominance block; In the luminance filtering process, modifying values ​​of the plurality of luminance samples using a clipping process such that a change amount of values ​​of the plurality of luminance samples included in the first luminance block and the second luminance block aligned in a direction perpendicular to the first boundary does not exceed a first threshold value; The first thresholds are determined based on the sizes of the first luminance block and the second luminance block and a quantization parameter, and are set symmetrically or asymmetrically with respect to the first boundary; In the color difference filtering process, changing values ​​of the plurality of chrominance samples using a clipping process such that amounts of change in values ​​of the plurality of chrominance samples included in the first chrominance block and the second chrominance block, which are aligned in a direction perpendicular to the second boundary, do not exceed second thresholds; Encoding device.

2. A processing circuit; a memory coupled to the processing circuit; The processing circuitry uses the memory to: performing a luminance filtering process on a first boundary between the first luminance block and the second luminance block; performing a chrominance filtering process on a second boundary between the first chrominance block and the second chrominance block; In the luminance filtering process, modifying values ​​of the plurality of luminance samples using a clipping process such that a change amount of values ​​of the plurality of luminance samples included in the first luminance block and the second luminance block aligned in a direction perpendicular to the first boundary does not exceed a first threshold value; The first thresholds are determined based on the sizes of the first luminance block and the second luminance block and a quantization parameter, and are set symmetrically or asymmetrically with respect to the first boundary; In the color difference filtering process, changing values ​​of the plurality of chrominance samples using a clipping process such that amounts of change in values ​​of the plurality of chrominance samples included in the first chrominance block and the second chrominance block, which are aligned in a direction perpendicular to the second boundary, do not exceed second thresholds; Decryption device.

Citation Information

Patent Citations

  • Decoding method and decoder

    JP2014042326A

  • Method for Filter Control and a Filtering Control Device

    US20130294525A1

  • Encoding / decoding apparatus and method using flexible deblocking filtering

    US20140133564A1

  • Video image encoding device and video image decoding device

    WO2012035746A1

  • Deblocking filtering control

    WO2012118421A1