Symbolization method and decoding method
By adjusting pixel values and using asymmetric filter characteristics based on quantization parameters and block sizes, the method improves error suppression at block boundaries in video coding, addressing limitations in existing technologies like HEVC.
Patent Information
- Application Number
- JP2024190270
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-12-27
- Filing Date
- 2024-10-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2037-12-21
AI Technical Summary
Existing video coding technologies, such as HEVC, require further improvements in encoding and decoding methods to enhance the suppression of errors at block boundaries in video images.
The proposed method involves adjusting pixel values in adjacent blocks to ensure change amounts fall within specific clip widths, using asymmetric filter characteristics based on quantization parameters and block sizes to perform deblocking filter processing, thereby reducing errors at boundaries.
This approach effectively suppresses errors at block boundaries by tailoring filter coefficients to pixel error amplitudes and block sizes, enhancing image quality and reducing artifacts.
Smart Images

Figure 0007710587000006 
Figure 0007710587000007 
Figure 0007710587000008
Abstract
Description
Technical Field
[0001] The present disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method.
Background Art
[0002] A video coding standard called HEVC (High-Efficiency Video Coding) has been standardized by JCT-VC (Joint Collaborative Team on Video Coding).
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In such encoding and decoding technologies, further improvements are required.
[0005] Therefore, an object of the present disclosure is to provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvements.
Means for Solving the Problems
[0006] An encoding method according to one aspect of the present disclosure changes values of a plurality of pixels in the first block and values of a plurality of pixels in the second block adjacent to the first block so that a change amount of each value is within a respective clip width, and encodes a third block included in a second picture different from the first picture with reference to a first picture including the first block and the second block after the filtering. The plurality of pixels in the first block and the plurality of pixels in the second block are arranged along a line intersecting the boundary. The plurality of pixels in the first block includes a first pixel at a first position. The plurality of pixels in the second block includes a second pixel at a second position corresponding to the first position across the boundary. The clip width is selected based on a quantization parameter. The clip width includes a first clip width corresponding to the first pixel and a second clip width corresponding to the second pixel. The first clip width and the second clip width are different. The first block and the second block are adjacent to each other on the left and right, and the boundary is a vertical boundary.
[0007] A decoding method according to an aspect of the present disclosure changes values of a plurality of pixels in a first block and values of a plurality of pixels in a second block adjacent to the first block so that a change amount of each value is within a respective clip width, in order to filter a boundary between the first block and the second block. The method decodes a third encoded block included in a second picture different from the first picture, with reference to a first picture including the first block and the second block after the filtering. The plurality of pixels in the first block and the plurality of pixels in the second block are arranged along a line intersecting the boundary. The plurality of pixels in the first block include a first pixel at a first position, and the plurality of pixels in the second block include a second pixel at a second position corresponding to the first position across the boundary. The clip width is selected based on a quantization parameter, and the clip width includes a first clip width corresponding to the first pixel and a second clip width corresponding to the second pixel. The first clip width and the second clip width are different. The first block and the second block are adjacent to each other on the left and right, and the boundary is a vertical boundary.
[0008] Note that these general or specific aspects may be implemented in a system, method, integrated circuit, computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, method, integrated circuit, computer program, and recording medium.
Advantages of the Invention
[0009] The present disclosure can provide an encoding device, a decoding device, an encoding method, or a decoding method that can achieve further improvements.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Embodiments for Carrying Out the Invention
[0011] An encoding device according to one aspect of the present disclosure includes a processing circuit and a memory. The processing circuit uses the memory to convert each block composed of a plurality of pixels into a block composed of a plurality of conversion coefficients using a base for each block composed of a plurality of pixels, and for each block composed of the plurality of conversion coefficients, at least an inverse transform is performed on the block to reconstruct a block composed of a plurality of pixels. Based on the combination of bases used for the conversion of two adjacent reconstructed blocks, filter characteristics for the boundary of the two blocks are determined, and deblocking filter processing having the determined filter characteristics is performed.
[0012] Depending on the combination of bases used for the transformation of two adjacent blocks, the error distribution near the boundary between the two blocks is different. For example, due to the transformation of two blocks, a large error may occur near the boundary of one of the two blocks, while a small error may occur near the boundary of the other block. Note that this error is the difference in pixel values between the original image or input image and the reconstructed image. In such a case, if deblocking filter processing with filter characteristics symmetric with respect to the boundary is performed, the pixel values with small errors may be greatly affected by the pixel values with large errors. That is, there is a possibility that the error cannot be sufficiently suppressed. Therefore, in the encoding device according to one aspect of the present disclosure, based on the combination of bases used for the transformation of each of two adjacent blocks that have been reconstructed, the filter characteristics with respect to the boundary between the two blocks are determined. Thereby, for example, filter characteristics asymmetric with respect to the boundary can be determined. As a result, even when there is a difference in error near the boundary between the two blocks as described above, by performing deblocking filter processing having asymmetric filter characteristics, the possibility of suppressing the error can be increased.
[0013] Further, in determining the filter characteristics, the processing circuit may determine, as the filter characteristics, a smaller filter coefficient for a pixel located at a position where the amplitude of the base used for the transformation of the block is larger.
[0014] For example, the higher the amplitude of the base, the higher the possibility that the pixel value of the pixel has a large error. In the encoding device according to one aspect of the present disclosure, a smaller filter coefficient is determined for a pixel having the pixel value with the large error. Therefore, by the deblocking filter processing having such a filter coefficient, the influence of the pixel value with the large error on the pixel value with the small error can be further suppressed. That is, the possibility of suppressing the error can be further increased.
[0015] Further, the amplitude of the base may be the amplitude of the zero-order base.
[0016] The lower the order of the basis, the greater the influence on the error. Therefore, for pixels located at positions where the amplitude of the 0th-order basis is large, smaller filter coefficients are determined for those pixels, thereby further increasing the possibility of suppressing the error.
[0017] Also, the two blocks consist of a first block and a second block located on the right or below the first block. In determining the filter characteristics, the processing circuit uses the basis used for the transformation of the first block as the first basis and the basis used for the transformation of the second block as the second basis. When this is the case, the first filter coefficient for pixels near the boundary within the first block and the second filter coefficient for pixels near the boundary within the second block may be determined as the filter characteristics based on the first basis and the second basis, respectively. For example, in determining the filter characteristics, when the first basis and the second basis are DST (Discrete Sine Transforms)-VII, the processing circuit may determine the second filter coefficient, which is larger than the first filter coefficient, as the filter characteristics.
[0018] When the first basis and the second basis are DST-VII, there is a high possibility that the error is large near the boundary of the first block and small near the boundary of the second block. Therefore, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and by performing deblocking filter processing with these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be increased.
[0019] Also, in determining the filter characteristics, when the first basis and the second basis are DCT (Discrete Cosine Transforms)-II, the processing circuit may determine the second filter coefficient equal to the first filter coefficient as the filter characteristics.
[0020] When the first basis and the second basis are DCT-II, the error is likely to be equal near the boundary of the first block and near the boundary of the second block. Therefore, in such a case, a second filter coefficient equal to the first filter coefficient is determined, and by performing deblocking filter processing having these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be increased.
[0021] Further, in the determination of the filter characteristics, when the first basis and the second basis are DST (Discrete Sine Transforms)-VII and the size of the second block is smaller than the size of the first block, the processing circuit determines the second filter coefficient larger than the first filter coefficient as the filter characteristics, and the slope of the filter coefficient between the first filter coefficient and the second filter coefficient may be gentler than when the sizes of the first block and the second block are equal.
[0022] When the first basis and the second basis are DST-VII and the size of the second block is smaller than the size of the first block, the error is large near the boundary of the first block, and the error is likely to be at a medium level near the boundary of the second block. That is, the error distribution near the boundary between the first block and the second block is likely to have a gentle gradient.
[0023] In the encoding device according to one aspect of the present disclosure, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and deblocking filter processing having these first and second filter coefficients is performed. Here, the slope of the filter coefficient between the determined first filter coefficient and the second filter coefficient is gentler than when the sizes of the first block and the second block are equal. Therefore, even if the error distribution near the boundary between the first block and the second block has a gentle gradient, the possibility of appropriately suppressing the error near the boundary can be increased.
[0024] Further, in determining the filter characteristics, the processing circuit further determines, based on the combination of bases of the first block and the second block, a first threshold value for the first block and a second threshold value for the second block as the filter characteristics. In the deblocking filter process, an operation is performed on the pixel value of the target pixel using the first filter coefficient and the second filter coefficient to obtain the pixel value of the target pixel after the operation. It is determined whether the amount of change from the pixel value of the target pixel before the operation to the pixel value after the operation is greater than the threshold value of the block to which the target pixel belongs among the first threshold value and the second threshold value. When the amount of change is greater than the threshold value, the pixel value of the target pixel after the operation may be clipped to the sum or difference between the pixel value of the target pixel before the operation and the threshold value.
[0025]
[0026] Also, an encoding device according to an aspect of the present disclosure includes a processing circuit and a memory. The processing circuit uses the memory to determine filter characteristics for a boundary between a first block and a second block adjacent to the first block based on block sizes of the first block and the second block, and performs deblocking filter processing having the determined filter characteristics. For example, in determining the filter characteristics, the processing circuit sets, as the filter characteristics, a first filter coefficient for a pixel near the boundary in the first block and a second filter coefficient for a pixel near the boundary in the second block. When the size of the second block is smaller than the size of the first block, the processing circuit may determine, as the filter characteristics, the second filter coefficient that is larger than the first filter coefficient.
[0027] Thereby, according to the difference in block sizes, for example, it is possible to determine filter characteristics that are asymmetric with respect to the boundary. As a result, even when there is a difference in error near the boundary between two blocks as described above, the possibility of suppressing the error can be increased by performing deblocking filter processing having asymmetric filter characteristics.
[0028] A decoding device according to an aspect of the present disclosure includes a processing circuit and a memory. The processing circuit uses the memory to reconstruct a block composed of a plurality of pixels by performing at least an inverse transform on each block composed of a plurality of transform coefficients obtained by a transform using a basis, and determines filter characteristics for a boundary between the two blocks based on a combination of bases used for the transforms of the two adjacent reconstructed blocks, and performs deblocking filter processing having the determined filter characteristics.
[0029] Depending on the combination of bases used for the transformation of two adjacent blocks, the error distribution near the boundary of the two blocks is different. For example, due to the transformation of two blocks, a large error may occur near the boundary of one of the two blocks, and a small error may occur near the boundary of the other block. Note that this error is the difference in pixel values between the original image or input image and the reconstructed image. In such a case, if deblocking filter processing with filter characteristics symmetric with respect to the boundary is performed, the pixel values with small errors may be greatly affected by the pixel values with large errors. That is, there is a possibility that the error cannot be sufficiently suppressed. Therefore, in the decoding apparatus according to one aspect of the present disclosure, based on the combination of bases used for the transformation of each of the two adjacent blocks that have been reconstructed, the filter characteristics with respect to the boundary of the two blocks are determined. Thereby, for example, filter characteristics asymmetric with respect to the boundary can be determined. As a result, even when there is a difference in error near the boundary of the two blocks as described above, by performing deblocking filter processing having asymmetric filter characteristics, the possibility of suppressing the error can be increased.
[0030] Further, in determining the filter characteristics, the processing circuit may determine, as the filter characteristics, a smaller filter coefficient for a pixel located at a position where the amplitude of the base used for the transformation of the block is larger.
[0031] For example, the higher the amplitude of the base, the higher the possibility that the pixel value of the pixel has a large error. In the decoding apparatus according to one aspect of the present disclosure, a smaller filter coefficient is determined for a pixel having the pixel value with the large error. Therefore, by the deblocking filter processing having such a filter coefficient, the influence of the pixel value with the large error on the pixel value with the small error can be further suppressed. That is, the possibility of suppressing the error can be further increased.
[0032] Further, the amplitude of the base may be the amplitude of the zero-order base.
[0033] The lower the order of the basis, the greater the influence on the error. Therefore, for pixels located at positions where the amplitude of the 0th-order basis is large, by determining smaller filter coefficients for those pixels, the possibility of suppressing the error can be further enhanced.
[0034] Also, the two blocks consist of a first block and a second block located on the right side or the lower side of the first block. In determining the filter characteristics, the processing circuit, when the basis used for the transformation of the first block is the first basis and the basis used for the transformation of the second block is the second basis, may determine, as the filter characteristics, a first filter coefficient for pixels near the boundary within the first block and a second filter coefficient for pixels near the boundary within the second block, based on the first basis and the second basis. For example, in determining the filter characteristics, when the first basis and the second basis are DST (Discrete Sine Transforms)-VII, the processing circuit may determine the second filter coefficient, which is larger than the first filter coefficient, as the filter characteristics.
[0035] When the first basis and the second basis are DST-VII, there is a high possibility that the error is large near the boundary of the first block and small near the boundary of the second block. Therefore, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and by performing deblocking filter processing with these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be enhanced.
[0036] Also, in determining the filter characteristics, when the first basis and the second basis are DCT (Discrete Cosine Transforms)-II, the processing circuit may determine the second filter coefficient equal to the first filter coefficient as the filter characteristics.
[0037] When the first basis and the second basis are DCT-II, the errors are likely to be equal near the boundary of the first block and near the boundary of the second block. Therefore, in such a case, a second filter coefficient equal to the first filter coefficient is determined, and by performing deblocking filter processing having these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be increased.
[0038] Further, in determining the filter characteristics, the processing circuit determines, as the filter characteristics, a second filter coefficient larger than the first filter coefficient when the first basis and the second basis are DST (Discrete Sine Transforms)-VII and the size of the second block is smaller than the size of the first block, and the slope of the filter coefficient between the first filter coefficient and the second filter coefficient may be gentler than when the sizes of the first block and the second block are equal.
[0039] When the first basis and the second basis are DST-VII and the size of the second block is smaller than the size of the first block, the error is large near the boundary of the first block, and the error is likely to be at a medium level near the boundary of the second block. That is, the error distribution near the boundary between the first block and the second block is likely to have a gentle gradient.
[0040] In the decoding apparatus according to one aspect of the present disclosure, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and deblocking filter processing having these first and second filter coefficients is performed. Here, the slope of the filter coefficient between the determined first filter coefficient and the second filter coefficient is gentler than when the sizes of the first block and the second block are equal. Therefore, even if the error distribution near the boundary between the first block and the second block has a gentle gradient, the possibility of appropriately suppressing the error near the boundary can be increased.
[0041] Further, in determining the filter characteristics, the processing circuit further determines, based on a combination of bases of the first block and the second block, a first threshold value for the first block and a second threshold value for the second block as the filter characteristics. In the deblocking filter process, an operation using the first filter coefficient and the second filter coefficient is performed on the pixel value of a target pixel to obtain the pixel value of the target pixel after the operation. It is determined whether a change amount from the pixel value of the target pixel before the operation to the pixel value after the operation is greater than the threshold value of the block to which the target pixel belongs among the first threshold value and the second threshold value. When the change amount is greater than the threshold value, the pixel value of the target pixel after the operation may be clipped to the sum or difference between the pixel value of the target pixel before the operation and the threshold value.
[0042] Thereby, when the change amount of the pixel value of the target pixel after the operation is greater than the threshold value, the pixel value after the operation is clipped to the sum or difference between the pixel value before the operation and the threshold value, so that it is possible to suppress a large change in the pixel value to be processed by the deblocking filter process. Also, the first threshold value for the first block and the second threshold value for the second block are determined based on a combination of bases of the first block and the second block. Therefore, for each of the first block and the second block, a large threshold value can be determined for a pixel located at a position where the amplitude of the base is large, that is, a pixel with a large error, and a small threshold value can be determined for a pixel located at a position where the amplitude of the base is small, that is, a pixel with a small error. As a result, it is possible to permit a large change in the pixel value of a pixel with a large error and prohibit a large change in the pixel value of a pixel with a small error by the deblocking filter process. Therefore, the possibility of appropriately suppressing the error near the boundary between the first block and the second block can be further enhanced.
[0043] Also, a decoding device according to one aspect of the present disclosure includes a processing circuit and a memory. The processing circuit uses the memory to determine filter characteristics for a boundary between a first block and a second block adjacent to the first block based on block sizes of the first block and the second block, and performs deblocking filter processing having the determined filter characteristics. For example, in determining the filter characteristics, the processing circuit sets, as the filter characteristics, a first filter coefficient for a pixel near the boundary in the first block and a second filter coefficient for a pixel near the boundary in the second block. When the size of the second block is smaller than the size of the first block, the processing circuit may determine, as the filter characteristics, the second filter coefficient that is larger than the first filter coefficient.
[0044] Accordingly, filter characteristics that are asymmetric with respect to the boundary, for example, can be determined according to the difference in block sizes. As a result, even when there is a difference in error near the boundary between two blocks as described above, the possibility of suppressing the error can be increased by performing deblocking filter processing having asymmetric filter characteristics.
[0045] These general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0046] Hereinafter, embodiments will be specifically described with reference to the drawings.
[0047] Note that all the embodiments described below show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.
[0048] (Embodiment 1) First, as an example of an encoding device and a decoding device to which the processing and / or configuration described in each aspect of the present disclosure to be described later can be applied, an overview of Embodiment 1 will be described. However, Embodiment 1 is merely an example of an encoding device and a decoding device to which the processing and / or configuration described in each aspect of the present disclosure can be applied, and the processing and / or configuration described in each aspect of the present disclosure can also be implemented in an encoding device and a decoding device different from Embodiment 1.
[0049] When applying the processing and / or configuration described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.
[0050] (1) For the encoding device or the decoding device of Embodiment 1, among the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. (2) For the encoding device or the decoding device of Embodiment 1, after performing any changes such as addition, replacement, deletion, etc. of the functions or processes implemented for some of the plurality of components constituting the encoding device or the decoding device, the component corresponding to the component described in each aspect of the present disclosure is replaced with the component described in each aspect of the present disclosure. For the method implemented by the encoding device or decoding device of Embodiment 1, after adding processing and / or making any changes such as replacement or deletion to some of the multiple processes included in the method, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure. (4) Implement a combination of some of the components constituting the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components implementing a part of the processes implemented by the components described in each aspect of the present disclosure. (5) Implement a combination of components having a part of the functions provided by some of the components constituting the encoding device or decoding device of Embodiment 1 or components implementing a part of the processes implemented by some of the components constituting the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components implementing a part of the processes implemented by the components described in each aspect of the present disclosure. (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure among the multiple processes included in the method. (7) Implement a combination of some of the multiple processes included in the method implemented by the encoding device or decoding device of Embodiment 1 with the processes described in each aspect of the present disclosure.
[0051] Note that the implementation manners of the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Also, the processes and / or configurations described in different aspects may be implemented in combination.
[0052] [Overview of the Encoding Device] First, the overview of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of an encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.
[0053] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0054] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0055] Hereinafter, each component included in the encoding device 100 will be described.
[0056] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These blocks of fixed size are sometimes called Coding Tree Units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These blocks of variable size are sometimes called Coding Units (CUs), Prediction Units (PUs), or Transformation Units (TUs). Note that in this embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be the processing units of CUs, PUs, and TUs.
[0057] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, solid lines represent block boundaries by quadtree block splitting, and dashed lines represent block boundaries by binary tree block splitting.
[0058] Here, block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).
[0059] The upper-left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper-left 64x64 block is split into two 16x64 blocks 11, 12 and a 32x64 block 13.
[0060] The upper-right 64x64 block is horizontally split into two rectangular 64x32 blocks 14, 15 (binary tree block splitting).
[0061] The 64x64 block in the lower left is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the 64x64 block in the lower left is divided into 16 16x32 blocks, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.
[0062] The 64x64 block 23 in the lower right is not divided.
[0063] As described above, in FIG. 2, the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.
[0064] Note that in FIG. 2, one block was divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary-tree block division). Division including such ternary-tree block division is sometimes called MBT (multi type tree) division.
[0065] [Subtraction unit] The subtraction unit 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in units of the blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.
[0066] The original signal is the input signal of the encoding device 100 and is a signal representing the image of each picture constituting a moving image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0067] [Transformation unit] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0068] Note that the transformation unit 106 may adaptively select a transformation type from a plurality of transformation types and transform the prediction error into transformation coefficients using a transformation basis function corresponding to the selected transformation type. Such a transformation may be called an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).
[0069] The plurality of transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the transformation basis functions corresponding to each transformation type. In FIG. 3, N indicates the number of input pixels. The selection of the transformation type from these plurality of transformation types may depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.
[0070] Information indicating whether to apply such an EMT or AMT (for example, called an AMT flag) and information indicating the selected transformation type are signaled at the CU level. Note that the signaling of these information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0071] Further, the conversion unit 106 may re-convert the conversion coefficient (conversion result). Such re-conversion may be referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (e.g., 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0072] Here, separable conversion is a method in which conversion is performed multiple times by separating for each direction by the number of dimensions of the input, and non-separable conversion is a method in which when the input is multi-dimensional, two or more dimensions are regarded as one dimension and conversion is performed collectively.
[0073] For example, as an example of non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.
[0074] Similarly, after regarding a 4×4 input block as an array having 16 elements, a method in which Givens rotation is performed multiple times on the array (Hypercube Givens Transform) is also an example of non-separable conversion.
[0075] [Quantization Unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order, and quantizes the scanned conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0076] The predetermined order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).
[0077] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.
[0078] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable-length encoding on the quantization coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.
[0079] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantization coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized conversion coefficients of the current block to the inverse conversion unit 114.
[0080] [Inverse Conversion Unit] The inverse transform unit 114 restores the prediction error by inversely transforming the transform coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients, thereby restoring the prediction error of the current block. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0081] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes quantization error.
[0082] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 114 and the prediction sample that is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.
[0083] [Block memory] The block memory 118 is a storage unit for storing blocks within the coded target picture (hereinafter referred to as the current picture), which are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.
[0084] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0085] In ALF, a least-squares error filter for removing coding distortion is applied. For example, for each 2x2 sub-block within a current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.
[0086] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (e.g., C = 5D + A) is calculated using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (e.g., 15 or 25 classes).
[0087] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantizing the addition result.
[0088] Based on the results of such classification, a filter for the sub-block is determined from among a plurality of filters.
[0089] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. The information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0090] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance difference, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or the CU level. Note that the signaling of the information indicating the on / off of ALF does not have to be limited to the picture level or the CU level, and it may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0091] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets does not have to be limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0092] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.
[0093] [Intra prediction unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to the blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.
[0094] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes includes one or more non-directional prediction modes and a plurality of directional prediction modes.
[0095] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).
[0096] The plurality of directional prediction modes includes, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.
[0097] In addition, in the intra prediction of the chrominance blocks, a luminance block may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0098] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction. Such intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called a PDPC flag) is signaled, for example, at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0099] [Inter prediction unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0100] The motion information used for motion compensation is signaled. A motion vector predictor may be used for the signaling of the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.
[0101] In addition to the motion information of the current block obtained by motion search, the motion information of adjacent blocks may also be used to generate an inter prediction signal. Specifically, an inter prediction signal may be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).
[0102] In such an OBMC mode, information indicating the size of sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be at other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).
[0103] The OBMC mode will be described in more detail. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of prediction image correction processing by OBMC processing.
[0104] First, a prediction image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.
[0105] Next, the motion vector (MV_L) of the encoded left adjacent block is applied to the block to be coded to obtain a prediction image (Pred_L), and the first correction of the prediction image is performed by weighting and superimposing the prediction image and Pred_L.
[0106] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_U), and the predicted image after the first correction and Pred_U are weighted and superimposed to perform the second correction of the predicted image, which is used as the final predicted image.
[0107] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to configure to perform more corrections than two stages using the right adjacent block or the lower adjacent block.
[0108] Note that the area for superposition may be only a partial area near the block boundary, rather than the pixel area of the entire block.
[0109] Here, the prediction image correction process from a single reference picture has been described. However, the same applies to the case of correcting the prediction image from a plurality of reference pictures. After obtaining the prediction images corrected from each reference picture, the obtained prediction images are further superimposed to obtain the final prediction image.
[0110] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.
[0111] As a method for determining whether to apply OBMC processing, for example, there is a method using an obmc_flag which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as the obmc_flag and encoding is performed by applying OBMC processing. If it does not belong to a region with complex motion, the value 0 is set as the obmc_flag and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, the obmc_flag described in the stream is decoded, and decoding is performed by switching whether to apply OBMC processing according to the value.
[0112] Note that the motion information may be derived on the decoder side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also for example, the motion information may be derived by performing motion search on the decoder side. In this case, the motion search is performed without using the pixel values of the current block.
[0113] Here, a mode in which motion search is performed on the decoder side will be described. This mode in which motion search is performed on the decoder side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.
[0114] An example of the FRUC process is shown in FIG. 5D. First, by referring to the motion vectors of the encoded blocks spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.
[0115] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also for example, pattern matching may be performed in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, whereby a motion vector for the current block may be derived. That is, search may be performed in the same manner for the region around the best candidate MV, and if there is an MV for which the evaluation value is a good value, the best candidate MV may be updated to the MV, and that may be used as the final MV of the current block. Note that a configuration in which the said process is not performed is also possible.
[0116] When performing processing in sub-block units, the same processing may be performed.
[0117] The evaluation value is calculated by obtaining the difference value of the reconstructed image through pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that in addition to the difference value, other information may be used to calculate the evaluation value.
[0118] As the pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.
[0119] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the above-described candidate, the region in another reference picture along the motion trajectory of the current block is used.
[0120] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is advisable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.
[0121] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.
[0122] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidates.
[0123] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of a current block is derived by searching, in a reference picture (Ref0), for a block that most closely matches a block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, a difference is derived between a reconstructed image of an encoded region that is adjacent to the current block on both the left and above or either one of them, and a reconstructed image at an equivalent position in the encoded reference picture (Ref0) specified by a candidate MV, and an evaluation value is calculated using the obtained difference value. It is preferable to select, as the best candidate MV, a candidate MV that has the best evaluation value among a plurality of candidate MVs.
[0124] Information indicating whether or not to apply such an FRUC mode (for example, called an FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (for example, called an FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level, and it may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0125] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode may be called the BIO (bi-directional optical flow) mode.
[0126] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v y( ) indicates the velocity vector, and τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.
[0127] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.
[0128] [Equation]
[0129] Here, I (k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the merge list or the like is corrected in pixel units.
[0130] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0131] Here, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.
[0132] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of the adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (v x , v y ) of each sub-block within the current block is derived by the following formula (2).
[0133]
Equation
[0134] Here, x and y respectively indicate the horizontal position and vertical position of the sub-block, and w indicates a predetermined weight coefficient.
[0135] Such an affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right control points are different. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode does not have to be limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0136] [Prediction Control Unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.
[0137] Here, an example of deriving the motion vector of the picture to be encoded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.
[0138] First, a prediction MV list in which candidates for the prediction MV are registered is generated. Examples of candidates for the prediction MV include a spatial adjacent prediction MV which is an MV of a plurality of encoded blocks located spatially adjacent to the block to be encoded, a temporal adjacent prediction MV which is an MV of a nearby block obtained by projecting the position of the block to be encoded in the encoded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV which is an MV with a value of zero.
[0139] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be encoded.
[0140] Furthermore, in the variable length encoding unit, a merge_idx which is a signal indicating which prediction MV has been selected is described in the stream and encoded.
[0141] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are only examples, and the number may be different from that in the figure, the configuration may not include some of the types of prediction MVs in the figure, or the configuration may include prediction MVs other than the types of prediction MVs in the figure.
[0142] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be encoded derived in the merge mode.
[0143] Here, an example of determining the MV using the DMVR process will be described.
[0144] FIG. 9C is a conceptual diagram for explaining the outline of the DMVR process.
[0145] First, using the optimal MVP set for the processing target block as a candidate MV, reference pixels are respectively obtained from a first reference picture that is a processed picture in the L0 direction and a second reference picture that is a processed picture in the L1 direction according to the candidate MV, and a template is generated by taking the average of each reference pixel.
[0146] Next, using the template, the peripheral areas of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, the MV value, etc.
[0147] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.
[0148] Note that even if it is not the processing itself described here, other processing may be used as long as it is a process capable of searching the periphery of the candidate MV to derive the final MV.
[0149] Here, a mode of generating a predicted image using the LIC process will be described.
[0150] FIG. 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction process by the LIC process.
[0151] First, an MV for obtaining a reference image corresponding to the encoding target block is derived from a reference picture that is an encoded picture.
[0152] Next, for the block to be encoded, information indicating how the luminance values change between the reference picture and the picture to be encoded is extracted using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, and the luminance correction parameter is calculated.
[0153] A predicted image for the block to be encoded is generated by performing luminance correction processing on the reference image in the reference picture specified by the MV using the luminance correction parameter.
[0154] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.
[0155] Also, although the process of generating a predicted image from a single reference picture has been described here, the same applies when generating a predicted image from a plurality of reference pictures. Luminance correction processing is performed on the reference images obtained from each reference picture in the same manner, and then the predicted image is generated.
[0156] As a method for determining whether to apply the LIC process, for example, there is a method using a lic_flag which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply the LIC process according to the value.
[0157] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the selected peripheral encoded block in the derivation of the MV in the merge mode process is encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. In the case of this example, the process in decoding is exactly the same.
[0158] [Overview of Decoder] Next, an overview of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing a functional configuration of a decoder 200 according to Embodiment 1. The decoder 200 is a moving image / image decoder that decodes moving images / images in units of blocks.
[0159] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0160] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0161] Each component included in the decoder 200 will be described below.
[0162] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bit stream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients in block units to the inverse quantization unit 204.
[0163] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0164] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.
[0165] For example, when the information decoded from the encoded bit stream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.
[0166] Also, for example, when the information decoded from the encoded bit stream indicates that NSST is to be applied, the inverse transform unit 206 applies an inverse inverse-transform to the transform coefficients.
[0167] [Addition Unit] The adder 208 reconstructs the current block by adding the prediction error which is the input from the inverse transform unit 206 and the prediction sample which is the input from the predictive control unit 220. Then, the adder 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0168] [Block Memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture) which are blocks referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the adder 208.
[0169] [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder 208, and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.
[0170] When the information indicating the on / off of the ALF read from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.
[0171] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.
[0172] [Intra Prediction Unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bitstream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0173] Note that when an intra prediction mode that refers to a luminance block is selected for intra prediction of a chrominance difference block, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0174] Also, when the information decoded from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0175] [Inter prediction unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bitstream, and outputs the inter prediction signal to the prediction control unit 220.
[0176] Note that when the information decoded from the encoded bitstream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.
[0177] Also, when the information decoded from the encoded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.
[0178] Also, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Further, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0179] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as the prediction signal to the addition unit 208.
[0180] [Deblocking filter processing] Next, the deblocking filter processing performed in the encoding device 100 and the decoding device 200 configured as described above will be specifically described with reference to the drawings. Hereinafter, the operation of the loop filter unit 120 included in the encoding device 100 will be mainly described, but the operation of the loop filter unit 212 included in the decoding device 200 is the same.
[0181] As described above, when encoding an image, the encoding apparatus 100 calculates a prediction error by subtracting a prediction signal generated by the intra prediction unit 124 or the inter prediction unit 126 from the original signal. The encoding apparatus 100 generates quantization coefficients by performing orthogonal transformation processing, quantization processing, and the like on the prediction error. Further, the encoding apparatus 100 restores the prediction error by inverse quantization and inverse orthogonal transformation of the obtained quantization coefficients. Here, since the quantization process is an irreversible process, the restored prediction error has an error (quantization error) with respect to the prediction error before the transformation.
[0182] The deblocking filter process performed by the loop filter unit 120 is a type of filter process performed for purposes such as reducing this quantization error. The deblocking filter process is applied to block boundaries to remove block noise. Hereinafter, this deblocking filter process will also be simply referred to as a filter process.
[0183] FIG. 11 is a flowchart showing an example of the deblocking filter process performed by the loop filter unit 120. For example, the process shown in FIG. 11 is performed for each block boundary.
[0184] First, the loop filter unit 120 calculates a block boundary strength (Bs) in order to determine the behavior of the deblocking filter process (S101). Specifically, the loop filter unit 120 determines Bs using the prediction mode of the block to be filtered, the nature of the motion vector, or the like. For example, if at least one of the blocks sandwiching the boundary is an intra prediction block, Bs = 2 is set. Also, if at least one of the following conditions (1) to (3) is satisfied: (1) at least one of the blocks sandwiching the block boundary includes dominant orthogonal transformation coefficients, (2) the difference between the motion vectors of both blocks sandwiching the block boundary is greater than or equal to a threshold value, and (3) the number of motion vectors or the reference images of both blocks sandwiching the block boundary are different, then Bs = 1 is set. If none of the conditions (1) to (3) are met, Bs = 0 is set.
[0185] Next, the loop filter unit 120 determines whether the set Bs is greater than the first threshold (S102). If Bs is less than or equal to the first threshold (No in S102), the loop filter unit 120 does not perform filter processing (S107).
[0186] On the other hand, if the set Bs is greater than the first threshold (Yes in S102), the loop filter unit 120 calculates the pixel variation d of the boundary region using the pixel values in the blocks on both sides of the block boundary (S103). This process will be described with reference to FIG. 12. If the pixel values at the block boundary are defined as shown in FIG. 12, the loop filter unit 120 calculates, for example, d = |p30 - 2×p20 + p10| + |p33 - 2×p23 + p13| + |q30 - 2×q20 + q10| + |q33 - 2×q23 + q13|.
[0187] Next, the loop filter unit 120 determines whether the calculated d is greater than the second threshold (S104). If d is less than or equal to the second threshold (No in S104), the loop filter unit 120 does not perform filter processing (S107). Note that the first threshold and the second threshold are different.
[0188] If the calculated d is greater than the second threshold (Yes in S104), the loop filter unit 120 determines the filter characteristics (S105) and performs filter processing with the determined filter characteristics (S106). For example, a 5-tap filter such as (1, 2, 2, 2, 1) / 8 is used. That is, for p10 shown in FIG. 12, an operation of (1×p30 + 2×p20 + 2×p10 + 2×q10 + 1×q20) / 8 is performed. Here, during the filter processing, a clip process is performed so that the displacement is within a certain range so as not to cause excessive smoothing. The clip process mentioned here is, for example, a threshold process in which when the threshold of the clip process is tc and the pixel value before filtering is q, the pixel value after filtering can only take values in the range of q ± tc.
[0189] Next, an example of applying an asymmetric filter across a block boundary in the deblocking filter process by the loop filter unit 120 according to the present embodiment will be described.
[0190] FIG. 13 is a flowchart showing an example of the deblocking filter process according to the present embodiment. Note that the process shown in FIG. 13 may be performed for each block boundary or for each unit pixel including one or more pixels.
[0191] First, the loop filter unit 120 acquires encoding parameters, and determines asymmetric filter characteristics across a block boundary using the acquired encoding parameters (S111). In the present disclosure, it is assumed that the acquired encoding parameters characterize, for example, the error distribution.
[0192] Here, the filter characteristics are filter coefficients and parameters used for controlling the filter process, etc. Also, the encoding parameters may be any parameters that can be used to determine the filter characteristics. The encoding parameters may be information indicating the error itself, or information or parameters related to the error (for example, affecting the magnitude relationship of the error).
[0193] Also, hereinafter, a pixel determined to have a large or small error based on the encoding parameters, that is, a pixel highly likely to have a large or small error, will simply be referred to as a pixel with a large or small error.
[0194] Here, it is not necessary to perform the determination process each time, and the process may be performed according to a rule that associates the encoding parameters and the filter characteristics determined in advance.
[0195] Note that even a pixel that is statistically likely to have a small error may have a larger error than the error of a pixel that is likely to have a large error when viewed pixel by pixel.
[0196] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S112).
[0197] Here, the filter characteristics determined in step S111 do not necessarily have to be asymmetric, and symmetric designs are also possible. In the following, a filter having asymmetric filter characteristics across a block boundary is also referred to as an asymmetric filter, and a filter having symmetric filter characteristics across a block boundary is also referred to as a symmetric filter.
[0198] Specifically, filter characteristics are determined in consideration of two points: pixels determined to have a small error are less likely to be affected by pixels with a large error in the surroundings, and pixels determined to have a large error are more likely to be affected by pixels with a small error in the surroundings. That is, the filter characteristics are determined such that the greater the error of a pixel, the greater the influence of the filter process. For example, the filter characteristics are determined such that the greater the error of a pixel, the greater the amount of change in the pixel value before and after the filter process. This can prevent values from deviating from the true value by fluctuating greatly for pixels that are likely to have a small error. Conversely, for pixels that are likely to have a large error, the error can be reduced by strongly being affected by pixels with a small error and varying the value.
[0199] In the following, an element that changes the displacement by the filter is defined as the weight of the filter. In other words, the weight indicates the degree of influence of the filter process on the target pixel. Increasing the weight means that the influence of the filter process on the pixel increases. In other words, it means that the pixel value after the filter process is more likely to be affected by other pixels. Specifically, increasing the weight means determining the filter characteristics such that the amount of change in the pixel value before and after the filter process increases, or such that the filter process is more likely to be performed.
[0200] That is, the loop filter unit 120 increases the weight for pixels with larger errors. Note that increasing the weight for pixels with larger errors does not necessarily mean continuously changing the weight based on the error, but also includes cases where the weight is changed stepwise. That is, the weight of the first pixel only needs to be smaller than the weight of the second pixel with a larger error than the first pixel. The same expression will be used below.
[0201] Note that in the finally determined filter characteristics, it is not necessary that pixels with larger errors have larger weights. That is, the loop filter unit 120 may, for example, modify the reference filter characteristics determined by a conventional method so that pixels with larger errors tend to have larger weights.
[0202] Hereinafter, a plurality of specific methods for asymmetrically changing the weight will be described. Note that any one of the methods shown below may be used, or a method combining a plurality of methods may be used.
[0203] As a first method, the loop filter unit 120 decreases the filter coefficient for pixels with larger errors. For example, the loop filter unit 120 decreases the filter coefficient for pixels with larger errors and increases the filter coefficient for pixels with smaller errors.
[0204] An example of deblocking filter processing performed on pixel p1 shown in FIG. 12 will be described. Hereinafter, without applying this method, for example, a filter determined by a conventional method is called a reference filter. The reference filter is a 5-tap filter perpendicular to the block boundary, and is assumed to be a filter extending over (p3, p2, p1, q1, q2). Also, assume that the filter coefficients are (1, 2, 2, 2, 1) / 8. Also, assume that the error of block P is likely to be large and the error of block Q is likely to be small. In this case, the filter coefficients are set so that block P with a large error is likely to be affected by block Q with a small error. Specifically, the filter coefficients used for pixels with small errors are set large, and the filter coefficients used for pixels with large errors are set small. For example, (0.5, 1.0, 1.0, 2.0, 1.5) / 6 is used as the filter coefficients.
[0205] As another example, 0 may be used for the filter coefficients of pixels with small errors. For example, (0, 0, 1, 2, 2) / 5 may be used as the filter coefficients. That is, the filter taps may be changed. Conversely, the current filter coefficient of 0 may be changed to a value other than 0. For example, (1, 2, 2, 2, 1, 1) / 9 etc. may be used as the filter coefficients. That is, the loop filter unit 120 may extend the filter taps to the side with small errors.
[0206] Note that the reference filter does not have to be a filter symmetric about the target pixel as in (1, 2, 2, 2, 1) / 8 above. In such a case, the loop filter unit 120 further adjusts the filter. For example, the filter coefficients of the reference filter used for the leftmost pixel of block Q are (1, 2, 3, 4, 5) / 15, and the filter coefficients of the reference filter used for the rightmost pixel of block P are (5, 4, 3, 2, 1) / 15. That is, in this case, filter coefficients that are reversed left and right are used between pixels sandwiching the block boundary. Such a filter characteristic that is symmetric with respect to inversion across the block boundary can also be called "a filter characteristic that is symmetric across the block boundary". That is, the filter characteristic that is asymmetric across the block boundary is a filter characteristic that is not symmetric with respect to inversion across the block boundary.
[0207] Also, in the same manner as above, when the error of block P is large and the error of block Q is small, the loop filter unit 120 changes, for example, the filter coefficient (5, 4, 3, 2, 1) / 15, which is the filter coefficient of the reference filter used for the pixel at the right end of block P, to (2.5, 2.0, 1.5, 2.0, 1.0) / 9.
[0208] In this way, in the deblocking filter process, a filter in which the filter coefficients change asymmetrically across the block boundary is used. For example, the loop filter unit 120 determines a reference filter having symmetric filter characteristics across the block boundary according to a predetermined reference. The loop filter unit 120 changes the reference filter so as to have asymmetric filter characteristics across the block boundary. Specifically, the loop filter unit 120 increases the filter coefficient of at least one pixel with a small error and / or decreases the filter coefficient of at least one pixel with a large error among the filter coefficients of the reference filter.
[0209] Next, a second method of changing the weights asymmetrically will be described. First, the loop filter unit 120 performs a filter operation using the reference filter. Next, the loop filter unit 120 performs asymmetric weighting across the block boundary on the reference change amount Δ0, which is the amount of change in the pixel values before and after the filter operation using the reference filter. Hereinafter, for the sake of distinction, the process using the reference filter is called a filter operation, and a series of processes including the filter operation and the subsequent correction process (for example, asymmetric weighting) is called a filter process (deblocking filter process).
[0210] For example, for pixels with small errors, the loop filter unit 120 calculates the corrected change amount Δ1 by multiplying the reference change amount Δ0 by a coefficient smaller than 1. Also, for pixels with large errors, the loop filter unit 120 calculates the corrected change amount Δ1 by multiplying the reference change amount Δ0 by a coefficient larger than 1. Next, the loop filter unit 120 generates the pixel value after the filter process by adding the corrected change amount Δ1 to the pixel value before the filter operation. Note that the loop filter unit 120 may perform only one of the process for pixels with small errors and the process for pixels with large errors.
[0211] For example, similar to the above, assume that the error of block P is large and the error of block Q is small. In this case, for the pixels included in block Q with small errors, the loop filter unit 120 calculates the corrected change amount Δ1 by, for example, multiplying the reference change amount Δ0 by 0.8. Also, for the pixels included in block P with large errors, the loop filter unit 120 calculates the corrected change amount Δ1 by, for example, multiplying the reference change amount Δ0 by 1.2. By doing so, the variation in the values of pixels with small errors can be reduced. Also, the variation in the values of pixels with large errors can be increased.
[0212] Note that as the ratio between the coefficient multiplied by the reference change amount Δ0 of pixels with small errors and the coefficient multiplied by the reference change amount Δ0 of pixels with large errors, a ratio of 1:1 may be selected. In this case, the filter characteristics are symmetric across the block boundary.
[0213] Further, the loop filter unit 120 may calculate a coefficient to be multiplied by the reference change amount Δ0 by multiplying a reference coefficient by a constant. In this case, the loop filter unit 120 uses a larger constant for pixels with a larger error than for pixels with a smaller error. As a result, the amount of change in the pixel value for a pixel with a larger error increases, and the amount of change in the pixel value for a pixel with a high possibility of having a smaller error decreases. For example, the loop filter unit 120 uses 1.2 or 0.8 as a constant for pixels adjacent to the block boundary, and uses 1.1 or 0.9 as a constant for pixels one pixel away from the pixels adjacent to the block boundary. Further, the reference coefficient is obtained, for example, by (A×(q1 - p1)-B×(q2 - p2)+C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Further, p1, p2, q1, and q2 are pixel values of pixels having the positional relationship shown in FIG. 12 with the block boundary in between.
[0214] Next, a third method of changing weights asymmetrically will be described. Similar to the second method, the loop filter unit 120 performs a filter operation using the filter coefficients of the reference filter. Next, the loop filter unit 120 adds an asymmetric offset value across the block boundary to the pixel value after the filter operation. Specifically, the loop filter unit 120 adds a positive offset value to the pixel value of a pixel with a larger error so that the value of the pixel with a larger error approaches the value of a pixel with a high possibility of having a smaller error and the displacement of the pixel with a larger error increases. Also, the loop filter unit 120 adds a negative offset to the pixel value of a pixel with a smaller error so that the value of the pixel with a smaller error does not approach the value of a pixel with a larger error and the displacement of the pixel with a smaller error decreases. As a result, the amount of change in the pixel value for a pixel with a larger error increases, and the amount of change in the pixel value for a pixel with a smaller error decreases. Note that the loop filter unit 120 may perform only one of the processing for pixels with a smaller error and the processing for pixels with a larger error.
[0215] For example, for pixels included in a block with a large error, the loop filter unit 120 calculates a corrected change amount Δ1 by adding a positive offset value (for example, 1) to the absolute value of the reference change amount Δ0. Also, for pixels included in a block with a small error, the loop filter unit 120 calculates a corrected change amount Δ1 by adding a negative offset value (for example, -1) to the absolute value of the reference change amount Δ0. Next, the loop filter unit 120 generates a pixel value after filter processing by adding the corrected change amount Δ1 to the pixel value before filter operation. Note that the loop filter unit 120 may add the offset value not to the change amount but to the pixel value after filter operation. Also, the offset value does not have to be symmetric across the block boundary.
[0216] Also, when the filter taps extend over a plurality of pixels from the block boundary, the loop filter unit 120 may change only the weight for a certain specific pixel or may change the weights for all pixels. Also, the loop filter unit 120 may change the weights according to the distance from the block boundary to the target pixel. For example, the loop filter unit 120 may make the filter coefficients for up to two pixels from the block boundary asymmetric and make the filter coefficients for subsequent pixels symmetric. Also, the weights of the filter may be common to a plurality of pixels or may be set for each pixel.
[0217] Next, a fourth method of changing the weights asymmetrically will be described. The loop filter unit 120 performs a filter operation using the filter coefficients of the reference filter. Next, when the change amount Δ between the pixel values before and after the filter operation exceeds a clip width that is a reference value, the loop filter unit 120 clips the change amount Δ to the clip width. The loop filter unit 120 sets the clip width asymmetrically across the block boundary.
[0218] Specifically, the loop filter unit 120 makes the clip width for pixels with a large error larger than the clip width for pixels with a small error. For example, the loop filter unit 120 makes the clip width for pixels with a large error a constant multiple of the clip width for pixels with a small error. As a result of changing the clip width, the values of pixels with a small error cannot change significantly. Also, the values of pixels with a large error can change significantly.
[0219] Note that the loop filter unit 120 may adjust the absolute value of the clip width instead of specifying the ratio of the clip widths. For example, the loop filter unit 120 fixes the clip width for pixels with a large error to be a multiple of a predetermined reference clip width. The loop filter unit 120 sets the ratio of the clip width for pixels with a large error to the clip width for pixels with a small error to 1.2:0.8. Specifically, for example, assume that the reference clip width is 10 and the change amount Δ before and after the filter operation is 12. In this case, when the reference clip width is used as it is, the change amount Δ is corrected to 10 by threshold processing. On the other hand, when the target pixel is a pixel with a large error, for example, the reference clip width is multiplied by 1.5. As a result, the clip width becomes 15, so threshold processing is not performed and the change amount Δ becomes 12.
[0220] Next, a fifth method of changing weights asymmetrically will be described. The loop filter unit 120 sets the condition for determining whether to perform filter processing asymmetrically across the block boundary. Here, the condition for determining whether to perform filter processing is, for example, the first threshold value or the second threshold value shown in FIG. 11.
[0221] Specifically, the loop filter unit 120 sets the condition so that filter processing is likely to be performed for pixels with a large error, and sets the condition so that filter processing is unlikely to be performed for pixels with a small error. For example, the loop filter unit 120 sets the threshold value for pixels with a small error to be higher than the threshold value for pixels with a large error. For example, the loop filter unit 120 makes the threshold value for pixels with a small error a constant multiple of the threshold value for pixels with a large error.
[0222] Further, the loop filter unit 120 may not only specify the ratio of the threshold values but also adjust the absolute value of the threshold values. For example, the loop filter unit 120 may fix the threshold value for pixels with small errors to a multiple of a predetermined reference threshold value, and set the ratio between the threshold value for pixels with small errors and the threshold value for pixels with large errors to 1.2:0.8.
[0223] Specifically, assume that the reference threshold value of the second threshold value in step S104 is 10 and d calculated from the pixel values in the block is 12. When the reference threshold value is used as the second threshold value as it is, it is determined that filter processing is to be performed. On the other hand, when the target pixel is a pixel with small error, for example, a value obtained by multiplying the reference threshold value by 1.5 is used as the second threshold value. In this case, the second threshold value becomes 15, which is larger than d. As a result, it is determined that filter processing is not to be performed.
[0224] Also, constants or the like indicating weights based on the errors used in the above first to fifth methods may be predetermined values in the encoding device 100 and the decoding device 200, or may be variable. Specifically, this constant is a coefficient multiplied by the filter coefficient in the first method or the filter coefficient of the reference filter, a coefficient multiplied by the reference change amount Δ0 in the second method or a constant multiplied by the reference coefficient, an offset value in the third method, a constant multiplied by the clip width or the reference clip width in the fourth method, and a constant multiplied by the threshold value or the reference threshold value in the fifth method, and the like.
[0225] When the constant is variable, information indicating the constant may be included in the bitstream as a parameter in, for example, sequence or slice units and transmitted from the encoding device 100 to the decoding device 200. Note that the information indicating the constant may be information indicating the constant itself, or may be information indicating the ratio or difference from the reference value.
[0226] In addition, as a method of changing a coefficient or a constant according to an error, for example, there are a method of changing linearly, a method of changing in a quadratic function manner, a method of changing in an exponential function manner, or a method of using a lookup table showing the relationship between the error and the constant.
[0227] Also, when the error is above a reference or when the error is below a reference, a fixed value may be used as a constant. For example, when the error of the loop filter unit 120 is below a predetermined range, the variable is set to a first value, and when the error is above the predetermined range, the variable is set to a second value. When the error is within the predetermined range, the variable may be continuously changed from the first value to the second value according to the error.
[0228] Also, when the error of the loop filter unit 120 exceeds a predetermined reference, a symmetric filter (reference filter) may be used instead of using an asymmetric filter.
[0229] Also, when using a lookup table or the like, the loop filter unit 120 may hold both tables for large and small errors, or may hold only one table and calculate the other constant according to a predetermined rule from the content of the table.
[0230] As described above, the encoding device 100 and the decoding device 200 according to the present embodiment can reduce the error of the reconstructed image by using an asymmetric filter, so that the encoding efficiency can be improved.
[0231] (Embodiment 2) In Embodiments 2 to 6, specific examples of the encoding parameters characterizing the error distribution described above will be described. In the present embodiment, the loop filter unit 120 determines filter characteristics according to the position within the block of the target pixel.
[0232] FIG. 14 is a flowchart showing an example of deblocking filter processing according to the present embodiment. First, the loop filter unit 120 acquires information indicating the position within the block of the target pixel as an encoding parameter characterizing the error distribution. Based on the position, the loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary (S121).
[0233] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S122).
[0234] Here, pixels farther from the intra-prediction reference pixels are more likely to have a larger error than pixels closer to the intra-prediction reference pixels. Therefore, the loop filter unit 120 determines the filter characteristics such that the amount of change in the pixel values before and after the filter process increases as the pixel is farther from the intra-prediction reference pixels.
[0235] For example, in the case of H.265 / HEVC or JEM, as shown in FIG. 15, the pixel close to the reference pixel is the pixel existing in the upper left within the block, and the pixel far from the reference pixel is the pixel existing in the lower right within the block. Therefore, the loop filter unit 120 determines the filter characteristics such that the weight of the pixel in the lower right within the block is larger than the weight of the pixel in the upper left.
[0236] Specifically, for pixels far from the reference pixels of intra prediction, the loop filter unit 120 determines filter characteristics so that the influence of the filtering process becomes large as described in Embodiment 1. That is, the loop filter unit 120 increases the weight of pixels far from the reference pixels of intra prediction. Here, increasing the weight means, as described above, (1) reducing the filter coefficient, (2) increasing the filter coefficient of pixels sandwiching a boundary (i.e., pixels close to the reference pixels of intra prediction), (3) increasing the coefficient multiplied by the change amount, (4) increasing the offset value of the change amount, (5) increasing the clip width, and (6) modifying the threshold value so that the filtering process is likely to be executed, and at least one of these is implemented. On the other hand, for pixels close to the reference pixels of intra prediction, the loop filter unit 120 determines filter characteristics so that the influence of the filtering process becomes small. That is, the loop filter unit 120 reduces the weight of pixels close to the reference pixels of intra prediction. Here, reducing the weight means, as described above, (1) increasing the filter coefficient, (2) reducing the filter coefficient of pixels sandwiching a boundary (i.e., pixels close to the reference pixels of intra prediction), (3) reducing the coefficient multiplied by the change amount, (4) reducing the offset value of the change amount, (5) reducing the clip width, and (6) modifying the threshold value so that the filtering process is unlikely to be executed, and at least one of these is implemented.
[0237] Note that when intra prediction is used, the above processing is performed, and it is not necessary to perform the above processing on blocks using inter prediction. However, since the nature of the intra prediction block may also be affected by inter prediction, the above processing may also be performed on the inter prediction block.
[0238] Also, the loop filter unit 120 may change the weights by arbitrarily specifying positions within a specific block. For example, as described above, the loop filter unit 120 may increase the weight of the pixel in the lower right of the block and decrease the weight of the pixel in the upper left of the block. Note that the loop filter unit 120 may change the weights by arbitrarily specifying positions within the block, not limited to the upper left and lower right.
[0239] Also, as shown in FIG. 15, at the horizontal adjacent block boundary, the error of the left block increases and the error of the right block increases. Therefore, the loop filter unit 120 may increase the weight of the left block and decrease the weight of the right block with respect to the horizontal adjacent block boundary.
[0240] Similarly, at the vertical adjacent block boundary, the error of the upper block increases and the error of the lower block decreases. Therefore, the loop filter unit 120 may increase the weight of the upper block and decrease the weight of the lower block with respect to the vertical adjacent block boundary.
[0241] Also, the loop filter unit 120 may change the weights according to the distance from the reference pixels of the intra prediction. Also, the loop filter unit 120 may determine the weights in units of block boundaries or in units of pixels. The farther away from the reference pixels, the more likely the error is to increase. Therefore, the loop filter unit 120 determines the filter characteristics so that the weight gradient becomes steeper as the distance from the reference pixels increases. Also, the loop filter unit 120 determines the filter characteristics so that the weight gradient on the upper side of the right side of the block is gentler than the weight gradient on the lower side.
[0242] (Embodiment 3) In this embodiment, the loop filter unit 120 determines the filter characteristics according to the orthogonal transform basis.
[0243] FIG. 16 is a flowchart showing an example of deblocking filter processing according to the present embodiment. First, the loop filter unit 120 acquires information indicating the orthogonal transformation basis used for the target block as an encoding parameter characterizing the error distribution. The loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary based on the orthogonal transformation basis (S131).
[0244] Next, the loop filter unit 120 performs filter processing having the determined filter characteristics (S132).
[0245] The encoding device 100 selects one orthogonal transformation basis, which is the transformation basis when performing orthogonal transformation, from among a plurality of candidates. The plurality of candidates includes, for example, a zero-order transformation basis such as DCT-II with a flat basis and a zero-order transformation basis such as DST-VII with a non-flat basis. FIG. 17 is a diagram showing the transformation basis of DCT-II. FIG. 18 is a diagram showing the transformation basis of DCT-VII.
[0246] The zero-order basis of DCT-II is constant regardless of the position within the block. That is, when DCT-II is used, the error within the block is constant. Therefore, when both blocks sandwiching the block boundary are transformed by DCT-II, the loop filter unit 120 performs filter processing using a symmetric filter without using an asymmetric filter.
[0247] On the other hand, the zero-order basis of DST-VII increases in value as the distance from the left or upper block boundary increases. That is, the error is likely to increase as the distance from the left or upper block boundary increases. Therefore, when at least one of the two blocks sandwiching the block boundary is transformed by DST-VII, the loop filter unit 120 uses an asymmetric filter. Specifically, the loop filter unit 120 determines the filter characteristics so that the influence of the filter processing is smaller for pixels with smaller values within the block of the low-order (e.g., zero-order) basis.
[0248] Specifically, when both blocks sandwiching a block boundary are transformed by DST-VII, the loop filter unit 120 determines filter characteristics for the lower-right pixel in the block such that the influence of the filtering process becomes greater by the method described above. Also, the loop filter unit 120 determines filter characteristics for the upper-left pixel in the block such that the influence of the filtering process becomes greater.
[0249] Also, even when DST-VII and DCT-II are vertically adjacent, the loop filter unit 120 determines filter characteristics such that the filter weight for the pixels in the lower part of the upper block where DST-VII is used, which are adjacent to the block boundary, becomes greater than the filter weight for the pixels in the upper part of the lower block where DCT-II is used. However, the difference in the amplitude of the lower-order basis in this case is smaller than the difference in the amplitude of the lower-order basis when DST-VIIs are adjacent. Therefore, the loop filter unit 120 sets the filter characteristics such that the slope of the weight in this case is smaller than the slope of the weight when DST-VIIs are adjacent. The loop filter unit 120 sets, for example, the weight when DCT-II and DCT-II are adjacent to 1:1 (symmetric filter), the weight when DST-VII and DST-VII are adjacent to 1.3:0.7, and the weight when DST-VII and DCT-II are adjacent to 1.2:0.8.
[0250] (Embodiment 4) In this embodiment, the loop filter unit 120 determines filter characteristics according to the pixel values sandwiching the block boundary.
[0251] FIG. 19 is a flowchart showing an example of the deblocking filter process according to this embodiment. First, the loop filter unit 120 acquires information indicating the pixel values in the block sandwiching the block boundary as an encoding parameter characterizing the error distribution. The loop filter unit 120 determines asymmetric filter characteristics across the block boundary based on the pixel values (S141).
[0252] Next, the loop filter unit 120 executes filter processing having the determined filter characteristics (S142).
[0253] For example, the loop filter unit 120 increases the difference in filter characteristics across the block boundary as the difference d0 in pixel values increases. Specifically, the loop filter unit 120 determines the filter characteristics so that the difference in the influence of the filter processing increases. For example, the loop filter unit 120 sets the weights to 1.4:0.6 when d0 > (quantization parameter) × (constant) is satisfied, and sets the weights to 1.2:0.8 when the above relationship is not satisfied. That is, the loop filter unit 120 compares the difference d0 in pixel values with a threshold value based on the quantization parameter, and when the difference d0 in pixel values is greater than the threshold value, it increases the difference in filter characteristics across the block boundary more than when the difference d0 in pixel values is less than the threshold value.
[0254] As another example, for example, the loop filter unit 120 increases the difference in filter characteristics across the block boundary as the average value b0 of the variances of the pixel values in both blocks across the block boundary increases. Specifically, the loop filter unit 120 may determine the filter characteristics so that the difference in the influence of the filter processing increases. For example, the loop filter unit 120 sets the weights to 1.4:0.6 when b0 > (quantization parameter) × (constant) is satisfied, and sets the weights to 1.2:0.8 when the above relationship is not satisfied. That is, the loop filter unit 120 compares the variance b0 of the pixel values with a threshold value based on the quantization parameter, and when the variance b0 of the pixel values is greater than the threshold value, it may increase the difference in filter characteristics across the block boundary more than when the variance b0 of the pixel values is less than the threshold value.
[0255] Note that, among adjacent blocks, which block's weight should be increased, that is, which block has a larger error, can be specified by the method of Embodiment 2 or 3 described above, or the method of Embodiment 6 to be described later. That is, the loop filter unit 120 determines filter characteristics that are asymmetric across a block boundary according to a predetermined rule (for example, the method of Embodiment 2, 3, or 6). Next, the loop filter unit 120 changes the determined filter characteristics based on the pixel value difference d0 so that the difference in filter characteristics across the block boundary becomes larger. That is, the loop filter unit 120 increases the ratio or difference between the weight of pixels with a large error and the weight of pixels with a small error.
[0256] Here, when the pixel value difference d0 is large, it may be the case where the block boundary coincides with the edge of an object in the image. Therefore, in such a case, by reducing the difference in filter characteristics across the block boundary, it is possible to suppress unnecessary smoothing from being performed.
[0257] Note that, conversely, the loop filter unit 120 may make the difference in filter characteristics across the block boundary smaller as the pixel value difference d0 becomes larger. Specifically, the loop filter unit 120 determines the filter characteristics so that the difference in the influence of the filter process becomes smaller. For example, when d0 > (quantization parameter) × (constant) is satisfied, the loop filter unit 120 sets the weight to 1.2:0.8, and when the above relationship is not satisfied, sets the weight to 1.4:0.6. Note that when the above relationship is satisfied, the weight may be set to 1:1 (symmetric filter). That is, the loop filter unit 120 compares the pixel value difference d0 with a threshold value based on the quantization parameter, and when the pixel value difference d0 is larger than the threshold value, makes the difference in filter characteristics across the block boundary smaller than when the pixel value difference d0 is smaller than the threshold value.
[0258] For example, if the difference d0 in pixel values is large, it means that the block boundary is likely to be prominent. Therefore, in such a case, by reducing the difference in filter characteristics across the block boundary, it is possible to suppress the weakening of smoothing by the asymmetric filter.
[0259] Note that these two processes may be performed simultaneously. For example, when the difference d0 in pixel values is less than the first threshold, the loop filter unit 120 uses the first weight; when the difference d0 in pixel values is greater than or equal to the first threshold and less than the second threshold, the loop filter unit 120 uses a second weight with a greater difference than the first weight; and when the difference d0 in pixel values is greater than or equal to the second threshold, the loop filter unit 120 may use a third weight with a smaller difference than the second weight.
[0260] Also, the difference d0 in pixel values may be the difference in pixel values across the boundary itself, or the average or variance of the differences in pixel values. For example, the difference d0 in pixel values is obtained by (A×(q1 - p1) - B×(q2 - p2) + C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Also, p1, p2, q1, and q2 are the pixel values of the pixels in the positional relationship shown in FIG. 12 across the block boundary.
[0261] Note that the setting of the difference d0 in pixel values and the weights may be performed on a pixel-by-pixel basis, on a block boundary-by-block boundary basis, or on a block group basis including a plurality of blocks (for example, on an LCU (Largest Coding Unit) basis).
[0262] (Embodiment 5) In this embodiment, the loop filter unit 120 determines filter characteristics according to the intra prediction direction and the block boundary direction.
[0263] FIG. 20 is a flowchart showing an example of deblocking filter processing according to the present embodiment. First, the loop filter unit 120 acquires information indicating an angle between a prediction direction of intra prediction and a block boundary as encoding parameters characterizing an error distribution. Based on the angle, the loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary (S151).
[0264] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S152).
[0265] Specifically, the loop filter unit 120 increases a difference in filter characteristics across the block boundary as the angle is closer to vertical, and decreases the difference in filter characteristics across the block boundary as the angle is closer to horizontal. More specifically, when the intra prediction direction is closer to perpendicular to the block boundary, the difference in filter weights for pixels on both sides across the block boundary becomes larger, and when the intra prediction direction is closer to horizontal with respect to the block boundary, the filter characteristics are determined so that the difference in filter weights for pixels on both sides across the block boundary becomes smaller. FIG. 21 is a diagram showing an example of weights with respect to the relationship between the intra prediction direction and the direction of the block boundary.
[0266] Note that which block among adjacent blocks has larger weights, that is, which block has a larger error, can be specified by the method of Embodiment 2 or 3 described above, or the method of Embodiment 6 described later. That is, the loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary according to a predetermined rule (for example, the method of Embodiment 2, 3, or 6). Next, the loop filter unit 120 changes the determined filter characteristics so that the difference in filter characteristics across the block boundary becomes larger based on the intra prediction direction and the direction of the block boundary.
[0267] Also, the encoding device 100 and the decoding device 200 specify the intra prediction direction using, for example, an intra prediction mode.
[0268] In addition, when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 may not consider the direction of the block boundary. For example, when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 may use a predetermined weight or a difference in weights regardless of the direction of the block boundary. Alternatively, when the intra prediction mode is the Planar mode or the DC mode, the loop filter unit 120 may use a symmetric filter.
[0269] (Embodiment 6) In this embodiment, the loop filter unit 120 determines filter characteristics according to a quantization parameter indicating the width of quantization.
[0270] FIG. 22 is a flowchart showing an example of the deblocking filter process according to this embodiment. First, the loop filter unit 120 acquires information indicating the quantization parameter used during quantization of the target block as an encoding parameter characterizing the error distribution. Based on the quantization parameter, the loop filter unit 120 determines filter characteristics that are asymmetric across the block boundary (S161).
[0271] Next, the loop filter unit 120 executes a filter process having the determined filter characteristics (S162).
[0272] Here, the larger the quantization parameter, the higher the possibility that the error becomes larger. Therefore, the loop filter unit 120 determines the filter characteristics so that the influence of the filter process becomes larger as the quantization parameter becomes larger.
[0273] FIG. 23 is a diagram showing an example of weights for quantization parameters. As shown in FIG. 23, the loop filter unit 120 increases the weight for the upper left pixel in the block as the quantization parameter increases. On the other hand, the loop filter unit 120 reduces the increase in weight for the lower right pixel in the block as the quantization parameter increases. That is, the loop filter unit 120 determines the filter characteristics such that the change in the influence of the filter process associated with the change in the quantization parameter of the upper left pixel is greater than the change in the influence of the filter process associated with the change in the quantization parameter of the lower right pixel.
[0274] Here, the upper left pixel in the block is more susceptible to the influence of the quantization parameter than the lower right pixel in the block. Therefore, by performing the above-described processing, the error can be appropriately reduced.
[0275] Further, the loop filter unit 120 may determine the weight of each of the two blocks sandwiching the boundary based on the quantization parameter of the block, or may calculate the average value of the quantization parameters of the two blocks and determine the weights of the two blocks based on the average value. Alternatively, the loop filter unit 120 may determine the weights of the two blocks based on the quantization parameter of one of the blocks. For example, the loop filter unit 120 uses the above method to determine the weight of one of the blocks based on the quantization parameter of the one block. Next, the loop filter unit 120 determines the weight of the other block according to a predetermined rule based on the determined weight.
[0276] Further, when the quantization parameters of the two blocks are different, or when the difference between the quantization parameters of the two blocks exceeds a threshold value, the loop filter unit 120 may use a symmetric filter.
[0277] Also, in FIG. 23, the weights are set using a linear function, but any function other than the linear function or a table may be used. For example, a curve showing the relationship between the quantization parameter and the quantization step (quantization width) may be used.
[0278] Also, when the quantization parameter exceeds the threshold value, the loop filter unit 120 may use a symmetric filter instead of an asymmetric filter.
[0279] Also, when the quantization parameter is described with decimal precision, the loop filter unit 120 may perform operations such as rounding, ceiling, or truncation on the quantization parameter and use the quantized parameter after the operation for the above processing. Alternatively, the loop filter unit 120 may perform the above processing considering up to the decimal unit.
[0280] As described above, in Embodiments 2 to 6, a plurality of methods for determining an error have been individually described, but two or more of these methods may be combined. In this case, the loop filter unit 120 may perform weighting on two or more combined elements.
[0281] Hereinafter, a modification example will be described.
[0282] Examples other than those of the encoding parameters described above may be used. For example, the encoding parameters may be the type of orthogonal transformation (such as Wavelet, DFT, or redundant transformation), the block size (width and height of the block), the direction of the motion vector, the length of the motion vector, or the number of reference pictures used for inter prediction, or information indicating the characteristics of the reference filter. Also, they may be used in combination. For example, the loop filter unit 120 may use an asymmetric filter only when the length of the block boundary is 16 pixels or less and the filter target pixel is close to the reference pixel for intra prediction, and use a symmetric filter in other cases. As another example, asymmetric processing may be performed only when a filter of a predetermined type among a plurality of filter candidates is used. For example, an asymmetric filter may be used only when the displacement by the reference filter is calculated by (A×(q1 - p1) - B×(q2 - p2) + C) / D. Here, A, B, C, and D are constants. For example, A = 9, B = 3, C = 8, and D = 16. Also, p1, p2, q1, and q2 are the pixel values of the pixels in the positional relationship shown in FIG. 12 across the block boundary.
[0283] Also, the loop filter unit 120 may perform the above processing for either the luminance signal or the color difference signal, or for both. Also, the loop filter unit 120 may perform common processing for the luminance signal and the color difference signal, or different processing. For example, the loop filter unit 120 may use different weights for the luminance signal and the color difference signal, or determine the weights according to different rules.
[0284] Also, the various parameters used in the above processing may be determined in the encoding device 100, or may be preset fixed values.
[0285] Furthermore, whether to perform the above processing, or not to perform it, or the content of the above processing may be switched in predetermined units. The predetermined unit is, for example, a slice unit, a tile unit, a wavefront division unit, or a CTU unit. Also, the content of the above processing refers to which of the plurality of methods shown above to use, or a parameter indicating a weight or the like, or a parameter for determining these.
[0286] In addition, the loop filter unit 120 may limit the area where the above processing is performed to the boundary of a CTU, the boundary of a slice, or the boundary of a tile.
[0287] Also, the number of taps of the filter may be different between the symmetric filter and the asymmetric filter.
[0288] In addition, the loop filter unit 120 may change whether to perform the above processing or the content of the above processing according to the frame type (I-frame, P-frame, B-frame).
[0289] In addition, the loop filter unit 120 may determine whether to perform the above processing or the content of the above processing according to whether a specific processing in the previous stage or the subsequent stage has been performed.
[0290] In addition, the loop filter unit 120 may perform different processes according to the type of prediction mode used for the block, or may perform the above processing only for the block using a specific prediction mode. For example, the loop filter unit 120 may perform different processes for a block using intra prediction, a block using inter prediction, and a merged block.
[0291] In addition, the encoding device 100 may encode filter information, which is a parameter indicating whether to perform the above processing or the content of the above processing. That is, the encoding device 100 may generate an encoded bitstream including the filter information. This filter information may include information indicating whether to perform the above processing on the luminance signal, information indicating whether to perform the above processing on the chrominance signal, or information indicating whether to perform different processes for each prediction mode.
[0292] Further, the decoding device 200 may perform the above processing based on the filter information included in the encoded bit stream. For example, the decoding device 200 may determine whether to perform the above processing or the content of the above processing based on the filter information.
[0293] (Embodiment 7) In the present embodiment, similar to the above-described Embodiment 3, the loop filter unit 120 determines filter characteristics according to the orthogonal transform basis. In the present embodiment, the configuration and processing in the above-described Embodiment 3 are shown more specifically, and in particular, the configuration and processing for determining filter characteristics according to the combination of the orthogonal transform bases of adjacent blocks will be described. Further, the loop filter unit 212 in the decoding device 200 has the same configuration as the loop filter unit 120 of the encoding device 100 and performs the same processing operation as the loop filter unit 120. Therefore, in the present embodiment, the configuration and processing operation of the loop filter unit 120 of the encoding device 100 will be described, and the detailed description of the configuration and processing operation of the loop filter unit 212 of the decoding device 200 will be omitted.
[0294] Various orthogonal transform bases are used for the orthogonal transform used in image encoding. Therefore, there are cases where the error distribution is not spatially uniform. The orthogonal transform basis is also referred to as a transform basis or simply a basis.
[0295] Specifically, in image encoding, the residual between the prediction signal generated by inter prediction or intra prediction and the original signal is orthogonally transformed and quantized. Thereby, the data amount is reduced. Since quantization is an irreversible process, an error, that is, a deviation from the image before encoding, occurs in the encoded image.
[0296] However, the error distribution generated during encoding does not necessarily become spatially uniform even if the quantization parameter is constant. This error distribution is considered to depend on the basis of the orthogonal transform.
[0297] That is, the conversion unit 106 selects a conversion basis for performing orthogonal conversion from among a plurality of candidates. At this time, when the 0th-order conversion basis is a flat basis, for example, DCT-II may be selected, or when the 0th-order conversion basis is not a flat basis, for example, DST-VII may be selected.
[0298] FIG. 24 is a diagram showing DCT-II, which is an example of a basis. Note that the horizontal axis of the graph in FIG. 24 indicates the position in the one-dimensional space, and the vertical axis indicates the value (i.e., amplitude) of the basis. Here, k indicates the order of the basis, n indicates the position in the one-dimensional space, and N indicates the number of pixels to be orthogonally converted. Note that the position n in the one-dimensional space is a horizontal position or a vertical position, and indicates a larger value as it goes from left to right in the horizontal direction or from top to bottom in the vertical direction. Further, x n represents the pixel value (specifically, the residual) of the pixel at position n, and Xk represents the frequency conversion result at the kth order, i.e., the conversion coefficient.
[0299] In DCT-II, when k = 0, the conversion coefficient X0 is represented by the following equation (3).
[0300]
Equation
[0301] Also, in DCT-II, when 1 ≦ k ≦ N - 1, the conversion coefficient Xk is represented by the following equation (4).
[0302]
Equation
[0303] FIG. 25 is a diagram showing DST-VII, which is an example of a basis. Note that the horizontal axis in FIG. 25 indicates the position in the one-dimensional space, and the vertical axis indicates the value (i.e., amplitude) of the basis.
[0304] In DST-VII, when 0 ≦ k ≦ N - 1, the conversion coefficient Xk is represented by the following formula (5).
[0305] [Number]
[0306] Thus, the conversion coefficient is basically determined by Σ(pixel value × conversion basis). Also, the conversion coefficients of lower-order bases tend to be larger than those of higher-order bases. Therefore, if DST-VII with a non-flat 0th-order basis is used as the conversion basis for a block, even if the same quantization error is added to the conversion coefficients, a bias will occur in the error distribution according to the values (i.e., amplitudes) of the lower-order bases. That is, within the block, the error tends to be small in the upper or left region where the value of the lower-order basis is small, and conversely, the error tends to be large in the lower or right region where the value of the lower-order basis is large.
[0307] FIG. 26 is a diagram showing the error distributions of four adjacent blocks and the error distributions after deblocking filter processing for them.
[0308] As shown on the left side of FIGS. 26(a) and (b), when DST-VII is used for the orthogonal transformation of each of the four blocks, the error is small in the upper or left region within these blocks, and conversely, the error is large in the lower or right region within the block.
[0309] When deblocking filter processing with filter characteristics symmetric with respect to the block boundary is performed in such a case where the error distribution is not uniform, there is a problem that a region where the error becomes larger occurs, as shown on the right side of FIG. 26(a). That is, when a region with a large error and a region with a small error are adjacent, an extra error is added to the pixels that originally had a small error.
[0310] Therefore, in the present embodiment, based on the basis used for the orthogonal transformation of the block, the error distribution is estimated, and the deblocking filter process is performed based on the result. As a result, as shown on the left side of (b) in FIG. 26, it is possible to suppress the error in the pixels originally having a large error without adding the error to the pixels originally having a small error.
[0311] FIG. 27 is a block diagram showing the main configuration of the loop filter unit 120 according to the present embodiment.
[0312] The loop filter unit 120 includes an error distribution estimation unit 1201, a filter characteristic determination unit 1202, and a filter processing unit 1203.
[0313] The error distribution estimation unit 1201 estimates the error distribution based on the error-related parameter. The error-related parameter is a parameter that affects the magnitude relationship of the error, and for example, indicates the type of basis applied to the orthogonal transformation of each of the two blocks sandwiching the block boundary to be subjected to the deblocking filter process.
[0314] The filter characteristic determination unit 1202 determines the filter characteristics based on the error distribution estimated by the error distribution estimation unit 1201.
[0315] The filter processing unit 1203 performs a deblocking filter process having the filter characteristics determined by the filter characteristic determination unit 1202 on the vicinity of the block boundary.
[0316] FIG. 28 is a flowchart showing the schematic processing operation of the loop filter unit 120 according to the present embodiment.
[0317] First, the error distribution estimation unit 1201 acquires error-related parameters. These error-related parameters are parameters that affect the magnitude relationship of errors. In other words, they are information that characterizes the error distribution in the region targeted for deblocking filter processing. Specifically, the error-related parameters indicate the types of bases applied to the orthogonal transforms of two blocks sandwiching the block boundary to be processed by the deblocking filter, that is, the combination of bases of the two blocks. Then, based on the error-related parameters, the error distribution estimation unit 1201 estimates the error distribution in the region targeted for deblocking filter processing (step S1201). Specifically, the error distribution estimation unit 1201 selects the i-th (1 ≤ i ≤ N) error distribution corresponding to the error-related parameters from the N-classified error distributions. Thereby, the error distribution is estimated.
[0318] Next, the filter characteristic determination unit 1202 determines filter characteristics according to the estimated error distribution (step S1202). That is, the filter characteristic determination unit 1202 refers to a table in which filter characteristics are associated with each of the N error distributions. Then, the filter characteristic determination unit 1202 finds the filter characteristics associated with the error distribution estimated in step S1201 from the table. Thereby, the filter characteristics are determined.
[0319] Finally, the filter processing unit 1203 performs deblocking filter processing in which the filter characteristics determined in step S1202 are reflected on the image indicated by the input signal (step S1203). Note that the image indicated by the input signal is, for example, a reconstructed image.
[0320] In the present embodiment, the error-related parameter indicates the type of basis used for the transformation of the block. Therefore, in the present embodiment, deblocking filter processing is performed based on the basis. For example, the loop filter unit 120 determines at least one of the filter coefficients and the threshold value of the deblocking filter processing as filter characteristics according to the combination of the orthogonal transformation bases used for two adjacent blocks. That is, the filter characteristics are designed based on the magnitude relationship of the errors. Then, the loop filter unit 120 performs deblocking filter processing having the determined filter characteristics on the target pixel.
[0321] That is, the encoding device 100 in the present embodiment includes, for example, a processing circuit and a memory. Then, the processing circuit performs the following processing using the memory. That is, the processing circuit transforms each block composed of a plurality of pixels into a block composed of a plurality of transform coefficients using a basis. Next, the processing circuit reconstructs a block composed of a plurality of pixels by performing at least an inverse transform on each block composed of the plurality of transform coefficients. Next, the processing circuit determines the filter characteristics for the boundary between the two blocks based on the combination of the bases used for the respective transforms of the two adjacent reconstructed blocks. Then, the processing circuit performs deblocking filter processing having the determined filter characteristics.
[0322] Note that the processing circuit is composed of, for example, a CPU (Central Processing Unit) or a processor, and functions as the loop filter unit 120 shown in FIG. 1. The memory may be a block memory 118 or a frame memory 122, or other memories.
[0323] Similarly, the decoding device 200 in the present embodiment includes, for example, a processing circuit and a memory. Then, the processing circuit performs the following processing using the memory. That is, the processing circuit reconstructs a block composed of a plurality of pixels by performing at least an inverse transform on each block composed of a plurality of transform coefficients obtained by transform using a basis. Next, the processing circuit determines the filter characteristics for the boundary between the two reconstructed adjacent blocks based on the combination of the bases used for the respective transforms of the two blocks. Then, the processing circuit performs a deblocking filter process having the determined filter characteristics.
[0324] Note that the processing circuit is composed of, for example, a CPU (Central Processing Unit) or a processor, and functions as the loop filter unit 212 shown in FIG. 10. The memory may be a block memory 210 or a frame memory 214, or other memories.
[0325] Thereby, since the filter characteristics for the boundary between the two reconstructed adjacent blocks are determined based on the combination of the bases used for the respective transforms of the two blocks, for example, it is possible to determine asymmetric filter characteristics for the boundary. As a result, even when there is a difference in the error of each pixel value of the pixels sandwiching the boundary between the two blocks, the possibility of suppressing the error can be increased by performing a deblocking filter process having asymmetric filter characteristics.
[0326] FIG. 29 is a block diagram showing a detailed configuration of the loop filter unit 120 according to the present embodiment.
[0327] The loop filter unit 120 includes an error distribution estimation unit 1201, a filter characteristic determination unit 1202, and a filter processing unit 1203, similar to the configuration shown in FIG. 27. Further, the loop filter unit 120 includes switches 1205, 1207, and 1209, a boundary determination unit 1204, a filter determination unit 1206, and a processing determination unit 1208.
[0328] The boundary determination unit 1204 determines whether the pixel to be deblocking-filtered (i.e., the target pixel) exists near the block boundary. Then, the boundary determination unit 1204 outputs the determination result to the switch 1205 and the process determination unit 1208.
[0329] When it is determined by the boundary determination unit 1204 that the target pixel exists near the block boundary, the switch 1205 outputs the image before the filter process to the switch 1207. Conversely, when it is determined by the boundary determination unit 1204 that the target pixel does not exist near the block boundary, the switch 1205 outputs the image before the filter process to the switch 1209.
[0330] The filter determination unit 1206 determines whether to perform deblocking filter processing on the target pixel based on the pixel values of at least one neighboring pixel around the target pixel and the error distribution estimated by the error distribution estimation unit 1201. Then, the filter determination unit 1206 outputs the determination result to the switch 1207 and the process determination unit 1208.
[0331] When it is determined by the filter determination unit 1206 that deblocking filter processing is to be performed on the target pixel, the switch 1207 outputs the image before the filter process obtained via the switch 1205 to the filter processing unit 1203. Conversely, when it is determined by the filter determination unit 1206 that deblocking filter processing is not to be performed on the target pixel, the switch 1207 outputs the image before the filter process obtained via the switch 1205 to the switch 1209.
[0332] When the filter processing unit 1203 obtains the image before the filter process via the switches 1205 and 1207, it performs deblocking filter processing having the filter characteristics determined by the filter characteristic determination unit 1202 on the target pixel. Then, the filter processing unit 1203 outputs the pixel after the filter process to the switch 1209.
[0333] According to the control by the processing determination unit 1208, the switch 1209 selectively outputs pixels that have not been subjected to deblocking filter processing and pixels that have been subjected to deblocking filter processing by the filter processing unit 1203.
[0334] The processing determination unit 1208 controls the switch 1209 based on the respective determination results of the boundary determination unit 1204 and the filter determination unit 1206. That is, when the processing determination unit 1208 determines by the boundary determination unit 1204 that the target pixel exists near the block boundary, and determines by the filter determination unit 1206 that deblocking filter processing is to be performed on the target pixel, the switch 1209 is caused to output the pixel that has been subjected to deblocking filter processing. Also, in cases other than the above, the processing determination unit 1208 causes the switch 1209 to output the pixel that has not been subjected to deblocking filter processing. By repeatedly performing such pixel output, the image after the filter processing is output from the switch 1209.
[0335] FIG. 30 is a diagram showing an example of a deblocking filter having filter characteristics symmetric with respect to a block boundary.
[0336] In the deblocking filter processing of HEVC, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as shown in FIG. 30, when there are pixels p0 - p2 and pixels q0 - q2 sandwiching the block boundary, the respective pixel values of the pixels q0 - q2 are changed to pixel values q'0 - q'2 by performing the operations shown in the following equations.
[0337] q’0=(p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8 q’1=(p0 + q0 + q1 + q2 + 2) / 4 q’2=(p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8
[0338] In the above equations, p0 - p2 and q0 - q2 are the respective pixel values of pixels p0 - p2 and pixels q0 - q2. Also, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the side opposite to the block boundary. Further, in the right - hand side of each of the above equations, the coefficient multiplied by the pixel value of each pixel used in the de - blocking filter process is the filter coefficient.
[0339] Furthermore, in the de - blocking filter process of HEVC, clip processing is performed so that the pixel value after the operation does not change beyond the threshold value. In this clip processing, the pixel value after the operation by the above equation is clipped to "the pixel value before the operation ± 2×threshold value" using the threshold value determined from the quantization parameter. Thereby, excessive smoothing can be prevented.
[0340] However, in such a de - blocking filter process, the change in the pixel value is determined from the surrounding pixel values and the quantization parameter, and its filter characteristics are designed without reflecting the non - uniformity of the error distribution within the block. Therefore, problems as shown in Fig. 26(a) may occur.
[0341] Therefore, the loop filter unit 120 in the present embodiment determines the filter characteristics of the de - blocking filter process while reflecting the non - uniformity of the error distribution within the block. Specifically, the error distribution estimation unit 1201 estimates that the error in the pixel value of a pixel located at a position with a larger base amplitude is larger. Next, the filter characteristic determination unit 1202 determines the filter characteristics based on the error distribution estimated by the error distribution estimation unit 1201. Here, when the filter characteristic determination unit 1202 determines the filter coefficient among the filter characteristics, for a pixel with a large error, a small filter coefficient is determined so that a pixel with a small error is less likely to be affected by pixels with large surrounding errors. Also, for a pixel with a small error, a large filter coefficient is determined so that a pixel with a large error is more likely to be affected by pixels with small surrounding errors.
[0342] That is, the filter coefficients or thresholds determined by the loop filter unit 120 do not necessarily have to be symmetric with respect to the block boundary. When performing deblocking filter processing on a target pixel with a large error, the loop filter unit 120 determines a large filter coefficient for a pixel with a small error. Also, when performing deblocking filter processing on a target pixel with a small error among two pixels sandwiching a block boundary, the loop filter unit 120 determines a small filter coefficient for a pixel with a large error.
[0343] That is, in the present embodiment, in determining the filter characteristics, the processing circuit determines, as the filter characteristics, a smaller filter coefficient for a pixel located at a position where the amplitude of the basis used for the block transformation is larger, with respect to that pixel. Also, as described above, the transformation coefficients of the lower-order bases tend to be larger than those of the higher-order bases. Therefore, the amplitude of the basis described above is, for example, the amplitude of the lower-order basis, and in particular, the amplitude of the 0th-order basis.
[0344] For example, the greater the amplitude of the basis at the position of a pixel, the higher the likelihood that the pixel value of that pixel has a large error. In the encoding device 100 according to the present embodiment, a small filter coefficient is determined for a pixel having a pixel value with such a large error. Therefore, by the deblocking filter processing having such a filter coefficient, the influence of the pixel value with a large error on the pixel value with a small error can be further suppressed. That is, the possibility of suppressing the error can be further increased. Also, the lower-order bases have a greater influence on the error. Therefore, by determining a smaller filter coefficient for a pixel located at a position where the amplitude of the 0th-order basis is large, with respect to that pixel, the possibility of suppressing the error can be further increased.
[0345] Note that the error distribution estimation unit 1201 may use at least one of the orthogonal transformation basis, the block size, the presence or absence of a previous-stage filter, the intra prediction direction, the number of reference frames for inter prediction, and the quantization parameter, as an error-related parameter.
[0346] Figures 31 to 35 are diagrams showing examples of bases for orthogonal transformation for each block size. That is, these diagrams show examples of bases for orthogonal transformation for block sizes N = 32, 16, 8, and 4, respectively. Specifically, Figure 31 shows bases of the 0th to 5th orders among the bases of DCT-II, Figure 32 shows bases of the 0th to 5th orders among the bases of DCT-V, and Figure 33 shows bases of the 0th to 5th orders among the bases of DCT-VIII. Figure 34 shows bases of the 0th to 5th orders among the bases of DST-I, and Figure 35 shows bases of the 0th to 5th orders among the bases of DST-VII. Note that the horizontal axis of each graph in Figures 31 to 35 indicates the position in one-dimensional space, and the vertical axis indicates the value (i.e., amplitude) of the basis.
[0347] The error distribution estimation unit 1201 estimates the error distribution based on bases such as DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII shown in Figures 31 to 35. At this time, the error distribution estimation unit 1201 may further estimate the error distribution based on the size of the block to which the orthogonal transformation of these bases is applied, that is, the number of pixels N.
[0348] <Specific Example of DST-VII / DST-VII> Figure 36 is a diagram showing an example of the determined filter coefficients.
[0349] For example, as shown in Figure 36, the loop filter unit 120 performs deblocking filter processing on the target pixel p0. Note that the block P and the block Q are adjacent to each other, for example, in the horizontal direction, and the target pixel p0 exists at a position close to the boundary (i.e., the block boundary) with the block Q in the block P.
[0350] Here, for example, each of the block P and the block Q is a block orthogonally transformed by DST-VII. In such a case, as shown in Figure 35, in the upper and left regions within the block, the amplitude of the lower-order (specifically, the 0th order) basis is small, and in the lower and right regions within the block, the amplitude of the lower-order basis is large.
[0351] Therefore, when block P and block Q are adjacent horizontally, near the block boundary of the left block P, that is, on the right side within the block P, the amplitude of the lower-order (specifically, the 0th-order) basis is large. Also, near the block boundary of the right block Q, that is, on the left side within the block Q, the amplitude of the lower-order basis is small.
[0352] As a result, the error distribution estimation unit 1201 estimates a large error near the block boundary within block P and a small error near the block boundary within block Q. Thereby, the error distribution near the block boundary is estimated.
[0353] Based on the estimated error distribution, the filter characteristic determination unit 1202 determines, for example, the filter coefficients of a 5-tap deblocking filter as filter characteristics.
[0354] Note that the 5-tap deblocking filter is a deblocking filter that uses five horizontally arranged pixels. When pixel p0 is the target pixel, the five pixels are pixels p2, p1, p0, q0, and q1. Also, the filter coefficients of the reference 5-tap deblocking filter are, for example, (1, 2, 2, 2, 1) / 8. By performing an operation using this reference filter coefficient, that is, p’0 = (1×p2 + 2×p1 + 2×p0 + 2×q0 + 1×q1) / 8, the pixel value p’0 of the target pixel p0 after the operation is calculated.
[0355] When the error distribution as described above is estimated, the filter characteristic determination unit 1202 in the present embodiment determines filter coefficients different from the above criteria as filter characteristics. Specifically, for pixels at positions where the error in block P is estimated to be large, small filter coefficients are determined, and for pixels at positions where the error in block Q is estimated to be small, large filter coefficients are determined. More specifically, as shown in FIG. 36, the filter characteristic determination unit 1202 determines the filter coefficients of pixels p2, p1, p0, q0, q1 for performing deblocking filter processing on the target pixel p0 of block P as (0.5, 1.0, 1.0, 2.0, 1.5) / 6. In this case, the filter coefficient determined for the pixel p0 at the position where the error in block P is estimated to be large is "1.0 / 6", and the filter coefficient determined for the pixel q0 at the position where the error in block Q is estimated to be small is "2.0 / 6". That is, the filter coefficient of pixel p0 is smaller than the filter coefficient of pixel q0, and the filter coefficient of pixel q0 is larger than the filter coefficient of pixel p0.
[0356] As a result, the filter processing unit 1203 calculates the pixel value p'0 after the operation of the target pixel p0 by performing an operation using the filter coefficients determined in this way, that is, p'0 = (0.5 × p2 + 1.0 × p1 + 1.0 × p0 + 2.0 × q0 + 1.5 × q1) / 6. The pixel value p'0 after this operation is the pixel value after the deblocking filter processing of the target pixel p0.
[0357] Here, the filter characteristic determination unit 1202 in the present embodiment may further determine the threshold value of the clip process as the filter characteristic based on the estimated error distribution. Note that the threshold value is the above-described reference value or clip width. For example, the filter characteristic determination unit 1202 determines a large threshold value for the pixels at the position where the amplitude of the base is large, that is, the position where the error is estimated to be large. Conversely, the filter characteristic determination unit 1202 determines a small threshold value for the pixels at the position where the amplitude of the base is small, that is, the position where the error is estimated to be small. Note that the amplitude of this base is, for example, the amplitude of a lower-order base, specifically, the amplitude of the 0th-order base. For example, when the reference threshold value is 10, the filter characteristic determination unit 1202 determines 12 as the threshold value for the right-side pixels near the block boundary of block P, and determines 8 as the threshold value for the left-side pixels near the block boundary of block Q.
[0358] When the threshold value is determined for the target pixel p0, the filter processing unit 1203 performs clip processing on the pixel value p'0 after the operation of the target pixel p0. The threshold value determined for the target pixel p0 is, for example, 12. Therefore, when the amount of change from the pixel value p0 before the operation to the pixel value p'0 after the operation is greater than the threshold value 12, the filter processing unit 1203 clips the pixel value p'0 after the operation to the pixel value (p0 + 12) or the pixel value (p0 - 12). More specifically, when (p0 - p'0)>12, the filter processing unit 1203 clips the pixel value p'0 after the operation to (p0 - 12), and when (p0 - p'0)<-12, the filter processing unit 1203 clips the pixel value p'0 after the operation to (p0 + 12). As a result, the pixel value (p0 + 12) or the pixel value (p0 - 12) becomes the pixel value p'0 after the deblocking filter processing of the target pixel p0. On the other hand, when the amount of change is 12 or less, the pixel value p'0 after the operation becomes the pixel value after the deblocking filter processing of the target pixel p0.
[0359] As described above, in this embodiment, the two blocks are composed of a first block and a second block located on the right or below the first block. When determining the filter characteristics, the processing circuit determines, as the filter characteristics, a first filter coefficient for pixels near the boundary in the first block and a second filter coefficient for pixels near the boundary in the second block, based on the first basis used for the transformation of the first block being the first basis and the basis used for the transformation of the second block being the second basis. Specifically, when determining the filter characteristics, the processing circuit determines, as the filter characteristics, a second filter coefficient larger than the first filter coefficient when the first basis and the second basis are DST (Discrete Sine Transforms)-VII.
[0360] When the first basis and the second basis are DST-VII, there is a high possibility that the error is large near the boundary of the first block and small near the boundary of the second block. Therefore, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and by performing deblocking filter processing with these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be increased.
[0361] Further, when determining the filter characteristics, the processing circuit also determines, as the filter characteristics, a first threshold value for the first block and a second threshold value for the second block, based on the combination of the bases of the first block and the second block. Then, in the deblocking filter processing, the processing circuit obtains the pixel value of the target pixel after the operation by performing an operation on the pixel value of the target pixel using the first filter coefficient and the second filter coefficient. Next, the processing circuit determines whether the amount of change from the pixel value of the target pixel before the operation to the pixel value after the operation is greater than the threshold value of the block to which the target pixel belongs among the first threshold value and the second threshold value. And when the amount of change is greater than the threshold value, the processing circuit clips the pixel value of the target pixel after the operation to the sum or difference between the pixel value of the target pixel before the operation and the threshold value.
[0362] Thus, when the change amount of the pixel value after the operation of the target pixel is larger than the threshold value, the pixel value after the operation is clipped to the sum or difference between the pixel value before the operation and the threshold value, so that it is possible to suppress a large change in the pixel value to be processed by the deblocking filter process. Also, the first threshold value for the first block and the second threshold value for the second block are determined based on the combination of the bases of the first block and the second block. Therefore, for each of the first block and the second block, a large threshold value can be determined for a pixel located at a position where the amplitude of the base is large, that is, a pixel with a large error, and a small threshold value can be determined for a pixel located at a position where the amplitude of the base is small, that is, a pixel with a small error. As a result, it is possible to permit a large change in the pixel value of a pixel with a large error and prohibit a large change in the pixel value of a pixel with a small error by the deblocking filter process. Therefore, the possibility of appropriately suppressing the error near the boundary between the first block and the second block can be further enhanced.
[0363] <Specific Example of DST-VII / DCT-II> FIG. 37 is a diagram showing another example of the filter coefficients to be determined.
[0364] For example, similar to the example shown in FIG. 36, the loop filter unit 120 performs deblocking filter processing on the target pixel p0 as shown in FIG. 37.
[0365] Here, for example, block P is a block orthogonally transformed by DST-VII, and block Q is a block orthogonally transformed by DCT-II. In such a case, as shown in FIG. 35, in the upper and left regions within block P, the amplitude of the low-order (specifically, the 0th order) basis is small, and in the lower and right regions within block P, the amplitude of the low-order basis is large. On the other hand, as shown in FIG. 31, within block Q, the amplitude of the low-order basis is constant, but it is larger than the amplitude in the upper and left regions of block P and smaller than the amplitude in the lower and right regions of block P. That is, within block Q, the amplitude of the low-order basis is constant at a medium level.
[0366] Therefore, when block P and block Q are horizontally adjacent, near the block boundary of the left block P, that is, on the right side within block P, the amplitude of the low-order (specifically, the 0th order) basis is large. Also, near the block boundary of the right block Q, that is, on the left side within block Q, the amplitude of the low-order basis is at a medium level.
[0367] As a result, the error distribution estimation unit 1201 estimates a large error near the block boundary within block P and estimates a medium-level error near the block boundary within block Q. Thereby, the error distribution near the block boundary is estimated. That is, an error distribution is estimated that shows a gentle error change in the direction perpendicular to the block boundary, compared to the example shown in FIG. 36.
[0368] Based on the estimated error distribution, the filter characteristic determination unit 1202 determines, as filter characteristics, the filter coefficients of, for example, a 5-tap deblocking filter. That is, the filter characteristic determination unit 1202 determines the filter coefficients such that the target pixel p0 located at a position with a large error in block P is more likely to be affected by the pixels having medium-level errors in block Q. Also, since the error distribution near the block boundary shows a gentle error change in the direction perpendicular to the block boundary, the filter characteristic determination unit 1202 determines five filter coefficients with smaller differences from each other than in the example shown in FIG. 36. Specifically, the filter characteristic determination unit 1202 determines the filter coefficients of pixels p2, p1, p0, q0, q1 for performing deblocking filter processing on the target pixel p0 of block P as (0.5, 1.0, 1.5, 1.75, 1.25) / 6. In this case, the filter coefficient determined for the pixel p0 at the position where the error in block P is estimated to be large is "1.5 / 6", and the filter coefficient determined for the pixel q0 at the position where the error in block Q is estimated to be at a medium level is "1.75 / 6". That is, the filter coefficient of pixel p0 is smaller than the filter coefficient of pixel q0, and the filter coefficient of pixel q0 is larger than the filter coefficient of pixel p0. Also, the difference between the filter coefficient of pixel p0 and the filter coefficient of pixel q0 is smaller than in the example shown in FIG. 36.
[0369] Also, even when block P and block Q are adjacent in the vertical direction, the error distribution estimation unit 1201 estimates the error distribution in the same manner as described above, and the filter characteristic determination unit 1202 determines the filter coefficients based on the error distribution. Further, the filter characteristic determination unit 1202 may determine the threshold for the clip process in the same manner as the example shown in FIG. 36, and the filter processing unit 1203 may perform the clip process using the threshold.
[0370] <Specific Example of DST-I / DST-I> FIG. 38 is a diagram showing still another example of the determined filter coefficients.
[0371] For example, similar to the examples shown in FIGS. 36 and 37, the loop filter unit 120 performs deblocking filter processing on the target pixel p0 as shown in FIG. 38.
[0372] Here, for example, each of block P and block Q is a block orthogonally transformed by DST-I. In such a case, as shown in FIG. 34, the amplitude of the low-order (specifically, 0th-order) basis in the upper and left regions within the block is equal to the amplitude of the low-order basis in the lower and right regions within the block.
[0373] Therefore, when block P and block Q are adjacent in the horizontal direction, the amplitude of the low-order basis near the block boundary of block P is equal to the amplitude of the low-order basis near the block boundary of block Q.
[0374] As a result, the error distribution estimation unit 1201 estimates a symmetric error distribution with respect to the block boundary as the error distribution near the block boundaries of block P and block Q.
[0375] In such a case, the filter characteristic determination unit 1202 determines the filter characteristics based on the error distribution as the above-described reference filter characteristics, that is, filter characteristics symmetric with respect to the block boundary.
[0376] Even when each of block P and block Q is a block orthogonally transformed by DCT-II, the amplitudes of the lower-order bases near the block boundary of block P and the amplitudes of the lower-order bases near the block boundary of block Q are equal, similar to the above. Therefore, even in such a case, the error distribution estimation unit 1201 estimates a symmetric error distribution with respect to the block boundary as the error distribution near the block boundaries of block P and block Q. Then, the filter characteristic determination unit 1202 determines, as the filter characteristics based on the error distribution, the above-described reference filter characteristics, that is, filter characteristics symmetric with respect to the block boundary. The filter coefficients of the reference 5-tap deblocking filter are, for example, (1, 2, 2, 2, 1) / 8. In this case, the filter coefficient determined for the pixel p0 in block P is "2 / 8", and the filter coefficient determined for the pixel q0 in block Q is "2 / 8". That is, these filter coefficients are symmetric with respect to the block boundary.
[0377] As described above, in this embodiment, when the first basis and the second basis are DCT (Discrete Cosine Transforms)-II, the second filter coefficient equal to the first filter coefficient is determined as the filter characteristics.
[0378] When the first basis and the second basis are DCT-II, there is a high possibility that the errors are equal near the boundary of the first block and near the boundary of the second block. Therefore, in such a case, the second filter coefficient equal to the first filter coefficient is determined, and by performing deblocking filter processing having these first and second filter coefficients, the possibility of appropriately suppressing the error near the boundary can be increased.
[0379] <Specific example of block size> FIG. 39 is a diagram for explaining the relationship between the block size and the error. Note that the block size is the block width or the number of pixels in the block. Specifically, FIG. 39(a) shows the 0th to 5th order bases of DST-VII at a block size N = 32, and FIG. 39(b) shows the 0th to 5th order bases of DST-VII at a block size N = 4. In FIG. 39, the horizontal axis of each graph indicates the position in the one-dimensional space, and the vertical axis indicates the value (i.e., amplitude) of the base.
[0380] The amplitude of the base at the block boundary varies depending on the block size, which is the number of pixels in the block to be orthogonally transformed. Therefore, the error distribution estimation unit 1201 in the present embodiment estimates the error distribution based also on the block size. Thereby, the accuracy of the determined filter coefficients can be improved.
[0381] For example, as shown in FIG. 39, even for blocks orthogonally transformed by DST-VII, the amplitude of the base at the block boundary differs between the case of a block size N = 4 and the case of a block size N = 32. Specifically, in the case of a block size N = 32, as shown in FIG. 39(a), the amplitudes of the 0th to 5th order bases at position n = 32 of DST-VII are all 1. Note that the position n = 32 is the right or lower end within the block of block size N = 32. On the other hand, in the case of a block size N = 4, as shown in FIG. 39(b), there are amplitudes of the 0th to 3rd order bases at position n = 4 of DST-VII that are smaller than 1. Note that the position n = 4 is the right or lower end within the block of block size N = 4.
[0382] Therefore, the error distribution estimation unit 1201 in the present embodiment estimates a smaller error as the error on the right and lower sides of the block orthogonally transformed by DST-VII in the case of a block size N = 4, and estimates a larger error than that in the case of a block size N = 32.
[0383] FIG. 40 is a diagram showing still another example of the determined filter coefficients.
[0384] For example, similar to the examples shown in FIGS. 36 to 38, the loop filter unit 120 performs filter processing on the target pixel p0 as shown in FIG. 40.
[0385] Here, for example, block P is a block of block size N = 32 orthogonally transformed by DST-VII, and block Q is a block of block size N = 4 orthogonally transformed by DST-VII. In such a case, as shown in FIG. 39(a), in the upper and left regions within block P, the amplitude of the low-order (specifically, 0th-order) basis is small, and in the lower and right regions within block P, the amplitude of the low-order basis is large. On the other hand, as shown in FIG. 39(b), in the upper and left regions within block Q, the amplitude of the low-order basis is at a medium level.
[0386] As a result, the error distribution estimation unit 1201 estimates a large error near the block boundary within block P and estimates a medium-level error near the block boundary within block Q. Thereby, the error distribution near the block boundary is estimated. That is, an error distribution showing a gentle error change in the direction perpendicular to the block boundary is estimated as compared with the example shown in FIG. 36.
[0387] Based on the estimated error distribution, the filter characteristic determination unit 1202 determines, for example, the filter coefficients of a 5-tap filter as filter characteristics. That is, the filter characteristic determination unit 1202 determines the filter coefficients such that the target pixel p0 with a large error in block P is more likely to be affected by the pixels with medium-level errors in block Q. Also, since the error distribution near the block boundary shows a gentle error change in the direction perpendicular to the block boundary, the filter characteristic determination unit 1202 determines five filter coefficients with smaller differences from each other than in the example shown in FIG. 36. Specifically, the filter characteristic determination unit 1202 determines the filter coefficients of pixels p2, p1, p0, q0, q1 for performing deblocking filter processing on the target pixel p0 of block P as (0.5, 1.0, 1.5, 1.75, 1.25) / 6. In this case, the filter coefficient determined for the pixel p0 at the position where the error in block P is estimated to be large is "1.5 / 6", and the filter coefficient determined for the pixel q0 at the position where the error in block Q is estimated to be at a medium level is "1.75 / 6". That is, the filter coefficient of pixel p0 is smaller than the filter coefficient of pixel q0, and the filter coefficient of pixel q0 is larger than the filter coefficient of pixel p0. Also, the difference between the filter coefficient of pixel p0 and the filter coefficient of pixel q0 is smaller than in the example shown in FIG. 36.
[0388] As described above, in this embodiment, in determining the filter characteristics, when the first basis and the second basis are DST (Discrete Sine Transforms)-VII and the size of the second block is smaller than the size of the first block, the processing circuit determines, as filter characteristics, a second filter coefficient larger than the first filter coefficient. The slope of the filter coefficients between the determined first filter coefficient and the second filter coefficient is gentler than when the sizes of the first block and the second block are equal.
[0389] When the first basis and the second basis are DST-VII and the size of the second block is smaller than the size of the first block, the error is large near the boundary of the first block, and the error is likely to be at a medium level near the boundary of the second block. That is, the error distribution near the boundary between the first block and the second block is likely to have a gentle gradient.
[0390] In the encoding device 100 according to the present embodiment, in such a case, a second filter coefficient larger than the first filter coefficient is determined, and deblocking filter processing having these first and second filter coefficients is performed. Here, the gradient of the filter coefficients between the determined first filter coefficient and the second filter coefficient is gentler than when the sizes of the first block and the second block are equal. Therefore, even if the error distribution near the boundary between the first block and the second block has a gentle gradient, the possibility of appropriately suppressing the error near the boundary can be increased.
[0391] <Modification 1> In the above-described Embodiment 7, the filter processing unit 1203 performs deblocking filter processing on pixels with small errors, but the deblocking filter processing on pixels with small errors may be turned off. Note that turning off the deblocking filter processing is equivalent to setting the filter coefficient for the target pixel to 1 and setting the filter coefficients for pixels other than the target pixel to zero.
[0392] Also, in the above-described Embodiment 7, the filter determination unit 1206 and the filter characteristic determination unit 1202 perform processing based on the error distribution estimated by the error distribution estimation unit 1201. However, the filter determination unit 1206 may determine whether to perform filter processing using only the quantization parameter, and the filter characteristic determination unit 1202 may determine the filter characteristics based on the quantization parameter and the basis of the orthogonal transformation.
[0393] Also, in this modification example, deblocking filter processing may be performed on each of the luminance signal and the color difference signals. In this case, the loop filter unit 120 may be designed to independently design the deblocking filter for the luminance signal and the deblocking filter for the color difference signal, or may be designed to be dependent on each other. For example, the loop filter unit 120 may perform the deblocking filter processing of the seventh embodiment only on one of the luminance signal and the color difference signal, and may perform other deblocking filter processing on the other signal.
[0394] Also, in this modification example, for example, the deblocking filter processing of the seventh embodiment may be performed only on the intra prediction block. Alternatively, the deblocking filter processing of the seventh embodiment may be performed on both the intra prediction block and the inter prediction block.
[0395] Also, the loop filter unit 120 may switch on and off the deblocking filter processing of the seventh embodiment in units of slices, tiles, wavefront segmentation units, or CTUs.
[0396] Also, there is a technique of increasing the bias of coefficients in the frequency domain and improving the compression efficiency by further performing a transformation after the orthogonal transformation (for example, Non-Separable Secondary Transform in JVET). Also at this time, the loop filter unit 120 may determine the filter characteristics based on the transformation basis of the orthogonal transformation.
[0397] Also, for example, for the first block, orthogonal transformation by DST-VII is performed in each of the first and second times, and for the second block, orthogonal transformation by DST-VII may be performed the first time and orthogonal transformation by DCT-II may be performed the second time. In such a case, the loop filter unit 120 estimates a steeper error distribution for the first block than for the second block. That is, the slopes of the error distributions in the horizontal and vertical directions of the first block are steeper than the slopes of the error distribution of the second block. Then, the loop filter unit 120 determines filter characteristics based on the steep error distribution.
[0398] Also, the block sizes may be different between the first transformation and the second transformation. In such a case, the loop filter unit 120 performs deblocking filter processing at at least one boundary of blocks having two different block sizes.
[0399] Also, when performing the deblocking filter processing of the seventh embodiment on an inter prediction block, it is expected that the distribution or absolute value of coefficients for each frequency after orthogonal transformation will differ depending on the prediction method. Examples of this prediction method include Uni-pred (single-reference prediction) and Bi-pred (two-reference prediction).
[0400] Therefore, the loop filter unit 120 may determine filter coefficients according to the prediction method. For example, the loop filter unit 120 makes the filter weight for a Uni-pred block having a tendency for the absolute value of the coefficient to be large smaller than the filter weight for a Bi-pred block.
[0401] Also, the loop filter unit 120 may independently determine the filter weight for a block to which the merge mode is applied. For example, the loop filter unit 120 makes the filter weight for a block to which the merge mode is applied larger or smaller than the filter weight for a block to which a prediction other than the merge mode is applied.
[0402] Also, the deblocking filter process in the above-described Embodiment 7 not only makes the image to be processed closer to the original image, but also makes the block boundary less prominent, similar to the conventional deblocking filter process. Therefore, if importance is placed not only on objective evaluation but also on subjective evaluation, it is effective to change the characteristics of the filter for each block size.
[0403] Specifically, since block noise is prominent in blocks with large errors, the loop filter unit 120 may set the filter coefficients, threshold values, or the number of taps of the filter on both sides sandwiching the block boundary to be larger than those for blocks with small errors.
[0404] Basically, since the correlation of pixel values is higher between pixels with closer distances, it is considered that the objective evaluation deteriorates as more pixels farther from the target pixel are used for the deblocking filter process of the target pixel. However, subjectively, the block noise becomes less prominent as more pixels are used for the deblocking filter process. Therefore, based on the trade-off between objective evaluation and subjective evaluation, the loop filter unit 120 may use such pixels farther from the target pixel for the deblocking filter process.
[0405] Also, in the above-described Embodiment 7, the error distribution is estimated according to bases such as DCT and DST, and the filter characteristics are determined based on the estimated error distribution. However, instead of those bases, the error distribution may be estimated according to other conversion methods such as KLT (Karhunen - Loeve Transform), DFT (Discrete Fourier transform), Wavelet transform, and redundant transform, and the filter characteristics may be determined based on the estimated error distribution.
[0406] Further, the encoding device 100 in the above-described Embodiment 7 includes an error distribution estimation unit 1201 that estimates an error distribution, but it is not necessary to include this error distribution estimation unit 1201. That is, the encoding device 100 may directly determine filter characteristics from the basis used for block transformation without estimating the error distribution.
[0407] <Modification Example 2> In the above-described Embodiment 7 or Modification Example 1 thereof, deblocking filter processing having filter characteristics determined based on a combination of bases is performed, but other loop filter processing other than deblocking filter processing may be performed.
[0408] For example, the loop filter unit 120 may determine filter coefficients of SAO (Sample Adaptive Offset) using the basis used for transformation and the position within the block of the target pixel. Alternatively, the loop filter unit 120 may perform trilateral filter processing based on three parameters. These three parameters include, for example, the difference in pixel values, the distance between pixels, and the error distribution estimated from the basis of orthogonal transformation. Or, the loop filter unit 120 may perform filter processing that applies the deblocking filter processing of the above-described Embodiment 7 or Modification Example 1 thereof within the intra processing loop. Also, the loop filter unit 120 does not necessarily need to change the pixel value of the target pixel based on information of surrounding pixels, and may give an offset corresponding to the error distribution for each target pixel.
[0409] Further, if the loop filter unit 120 can predict the error distribution, it is not necessary to acquire error-related parameters for each filter process. For example, in an intra prediction block by the software of JEM (Joint Exploration Model) 4.0, due to the above-described EMT design, errors are less likely to occur on the upper side and left side of the block, and errors are more likely to occur on the lower side and right side of the block. Therefore, the loop filter unit 120 may perform deblocking filter processing with weak filter strength on the upper side and left side of the JEM4.0 intra prediction block in consideration of this in advance.
[0410] In Embodiment 7 or Modification Example 1 thereof described above, the filter determination unit 1206 determines whether or not to perform deblocking filter processing on a target pixel. However, without performing such determination, deblocking filter processing may be performed on all block boundaries.
[0411] Also, in Embodiment 7 or Modification Example 1 thereof described above, the loop filter unit 120 performs deblocking filter processing having filter characteristics determined based on a combination of bases on block boundaries of the reconstructed block output from the addition unit 116. Here, for example, the encoding device 100 may include a filter unit different from the loop filter unit 120. That is, the filter unit performs filter processing on the reconstructed block, and the loop filter unit 120 performs deblocking filter processing on block boundaries of the reconstructed block that has been filter processed by the filter unit. In such a case, the loop filter unit 120 may determine the filter characteristics of the deblocking filter processing based on the combination of bases and also on the filter characteristics of the filter unit. Further, the filter unit may perform deblocking filter processing having filter characteristics symmetric with respect to the block boundary. In this case, the filter coefficients used in the filter unit are set to be smaller than the filter coefficients when the loop filter unit 120 in Embodiment 7 or Modification Example 1 thereof is not provided in the encoding device 100. Also, the filter unit may perform bilateral filter processing based on two parameters as deblocking filter processing having filter characteristics symmetric with respect to the block boundary. The two parameters are composed of, for example, the difference in pixel values and the distance between pixels. In this case, the loop filter unit 120 may determine, for example, filter coefficients that are overall smaller than the filter coefficients shown in FIGS. 36 to 38 and FIG. 40.
[0412] In addition, in the seventh embodiment or the first modification thereof described above, the loop filter unit 120 performs deblocking filter processing with respect to the block boundary, but may perform filter processing on a region that is not the block boundary within the block. For example, the loop filter unit 120 may change the pixel value of a pixel with a large error by using the pixel value of a pixel with a small error within one block. Hereinafter, such filter processing is referred to as in-block filter processing.
[0413] Specifically, as shown in FIG. 35, within the block orthogonally transformed by DST-VII, the error of the pixel on the upper side or the left side is small, and the error of the pixel on the lower side or the right side is large. Therefore, the loop filter unit 120 changes the pixel value of the pixel on the lower side or the right side by using the pixel value of the pixel on the upper side or the left side by performing in-block filter processing. Thereby, the possibility of reducing the error can be increased.
[0414] FIG. 41 is a diagram showing gradients of bases different according to the block size. In FIG. 41, the horizontal axis of each graph indicates the position in the one-dimensional space, and the vertical axis indicates the value (that is, amplitude) of the base.
[0415] As shown in FIG. 41(a), the amplitude of the 0th base of DST-VII at the block size N = 32 increases approximately 10-fold from the position n = 1 to n = 32 in the one-dimensional space. On the other hand, the amplitude of the 0th base of DST-VII at the block size N = 4 increases approximately 3-fold from the position n = 1 to n = 4 in the one-dimensional space.
[0416] Therefore, even if the bases used for the conversion of each of the two blocks are the same, if the block sizes are different, the gradient of the base within the block is steeper for the small block and gentler for the large block.
[0417] That is, as described above, when the loop filter unit 120 performs in-block filtering, the effect of enhancing the possibility of error reduction can be more effectively exerted on a block with a steep gradient in the basis of the lower order of the orthogonal transform, that is, a small block.
[0418] Also, even when the block size of a block is large, if the distance dependence or direction dependence of the correlation of pixel values within the block is weak, the loop filter unit 120 can increase the possibility of error reduction by performing in-block filtering. For example, the intra prediction direction affects the distance dependence or direction dependence of the correlation of pixel values. The distance dependence is the property that the closer the distance between two pixels, the higher the correlation of their pixel values. The direction dependence is the property that the correlation of the pixel values of these pixels varies according to the direction from one pixel to the other pixel. Therefore, the loop filter unit 120 may change the pixel value of a pixel with a large error by performing in-block filtering according to the intra prediction direction.
[0419] [Implementation Example] FIG. 42 is a block diagram showing an implementation example of the encoding device 100 according to each of the above embodiments. The encoding device 100 includes a processing circuit 160 and a memory 162. For example, a plurality of components of the encoding device 100 shown in FIG. 1 are implemented by the processing circuit 160 and the memory 162 shown in FIG. 42.
[0420] The processing circuit 160 is a circuit that performs information processing and is a circuit that can access the memory 162. For example, the processing circuit 160 is a dedicated or general-purpose electronic circuit that encodes moving images. The processing circuit 160 may be a processor such as a CPU. Also, the processing circuit 160 may be an aggregate of a plurality of electronic circuits. Also, for example, the processing circuit 160 may perform the roles of a plurality of components of the encoding device 100 shown in FIG. 1, excluding the components for storing information.
[0421] Memory 162 is a general-purpose or dedicated memory in which information for the processing circuit 160 to encode a moving image is stored. Memory 162 may be an electronic circuit and may be connected to the processing circuit 160. Also, memory 162 may be included in the processing circuit 160. Further, memory 162 may be an aggregate of a plurality of electronic circuits. Also, memory 162 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or a recording medium, etc. Also, memory 162 may be a non-volatile memory or a volatile memory.
[0422] For example, memory 162 may store the moving image to be encoded, or may store a bit string corresponding to the encoded moving image. Also, a program for the processing circuit 160 to encode a moving image may be stored in memory 162.
[0423] Also, for example, memory 162 may serve as a component for storing information among a plurality of components of the encoding device 100 shown in FIG. 1. Specifically, memory 162 may serve as the block memory 118 and the frame memory 122 shown in FIG. 1. More specifically, memory 162 may store processed sub-blocks, processed blocks, processed pictures, etc.
[0424] Note that in the encoding device 100, not all of the plurality of components shown in FIG. 1 etc. need to be implemented, and not all of the plurality of processes described above need to be performed. A part of the plurality of components shown in FIG. 1 etc. may be included in other devices, and a part of the plurality of processes described above may be executed by other devices. And in the encoding device 100, by implementing a part of the plurality of components shown in FIG. 1 etc. and performing a part of the plurality of processes described above, a moving image can be appropriately processed with a small amount of code.
[0425] FIG. 43 is a block diagram showing an implementation example of the decoding device 200 according to each of the above embodiments. The decoding device 200 includes a processing circuit 260 and a memory 262. For example, a plurality of components of the decoding device 200 shown in FIG. 10 are implemented by the processing circuit 260 and the memory 262 shown in FIG. 43.
[0426] The processing circuit 260 is a circuit that performs information processing and is a circuit that can access the memory 262. For example, the processing circuit 260 is a general-purpose or dedicated electronic circuit that decodes a moving image. The processing circuit 260 may be a processor such as a CPU. Also, the processing circuit 260 may be an aggregate of a plurality of electronic circuits. Further, for example, the processing circuit 260 may play the roles of a plurality of components of the decoding device 200 shown in FIG. 10, excluding the components for storing information.
[0427] The memory 262 is a general-purpose or dedicated memory in which information for the processing circuit 260 to decode a moving image is stored. The memory 262 may be an electronic circuit and may be connected to the processing circuit 260. Also, the memory 262 may be included in the processing circuit 260. Further, the memory 262 may be an aggregate of a plurality of electronic circuits. Also, the memory 262 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or a recording medium, etc. Also, the memory 262 may be a non-volatile memory or a volatile memory.
[0428] For example, the memory 262 may store a bit string corresponding to the encoded moving image, or may store a moving image corresponding to the decoded bit string. Also, the memory 262 may store a program for the processing circuit 260 to decode a moving image.
[0429] Also, for example, the memory 262 may serve as a component for storing information among the plurality of components of the decoding device 200 shown in FIG. 10. Specifically, the memory 262 may serve as the block memory 210 and the frame memory 214 shown in FIG. 10. More specifically, processed sub-blocks, processed blocks, processed pictures, etc. may be stored in the memory 262.
[0430] In the decoding device 200, not all of the plurality of components shown in FIG. 10 etc. need to be implemented, and not all of the plurality of processes described above need to be performed. A part of the plurality of components shown in FIG. 10 etc. may be included in other devices, and a part of the plurality of processes described above may be executed by other devices. And in the decoding device 200, by implementing a part of the plurality of components shown in FIG. 10 etc. and performing a part of the plurality of processes described above, a moving image can be appropriately processed with a small amount of code.
[0431] [Supplementary Explanation] The encoding device 100 and the decoding device 200 in each of the above embodiments may be used as an image encoding device and an image decoding device, respectively, or may be used as a moving image encoding device and a moving image decoding device. Alternatively, the encoding device 100 and the decoding device 200 may each be used as an inter prediction device. That is, the encoding device 100 and the decoding device 200 may each correspond only to the inter prediction unit 126 and the inter prediction unit 218.
[0432] Also, in each of the above embodiments, the prediction block is encoded or decoded as an encoding target block or a decoding target block, but the encoding target block or the decoding target block is not limited to the prediction block, and may be a sub-block or another block.
[0433] In each of the above-described embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0434] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device that is electrically connected to the processing circuit and accessible from the processing circuit.
[0435] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using the storage device. Further, when the processing circuit includes a program execution unit, the storage device stores a software program executed by the program execution unit.
[0436] Here, the software that realizes the encoding device 100 or the decoding device 200 in each of the above-described embodiments is the following program.
[0437] That is, this program causes a computer to execute processing according to the flowchart shown in any of FIGS. 5B, 5D, 11, 13, 14, 16, 19, 20, 22, and 28.
[0438] Also, each component may be a circuit as described above. These circuits may constitute one circuit as a whole, or may be separate circuits. Further, each component may be realized by a general-purpose processor or a dedicated processor.
[0439] Also, another component may execute the processing executed by a specific component. Also, the order in which the processing is executed may be changed, or a plurality of processes may be executed in parallel. Further, the encoding / decoding apparatus may include the encoding apparatus 100 and the decoding apparatus 200.
[0440] The ordinal numbers such as the first and second used in the description may be changed as appropriate. Also, an ordinal number may be newly given to or removed from components and the like.
[0441] As described above, the aspects of the encoding apparatus 100 and the decoding apparatus 200 have been described based on each embodiment, but the aspects of the encoding apparatus 100 and the decoding apparatus 200 are not limited to these embodiments. Without departing from the spirit of the present disclosure, forms in which various modifications conceived by those skilled in the art are applied to the embodiments or forms constructed by combining components in different embodiments may also be included within the scope of the aspects of the encoding apparatus 100 and the decoding apparatus 200.
[0442] (Embodiment 8) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, and the like. Also, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (a dedicated circuit).
[0443] Also, the processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing or distributed processing may be performed.
[0444] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible and are also included within the scope of the aspects of the present disclosure.
[0445] Furthermore, here, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, and an image encoding / decoding device having both. Other configurations in the system can be appropriately changed as the case may be.
[0446] [Usage Example] FIG. 44 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, ex110, which are fixed radio stations, are installed in each cell.
[0447] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above elements. The devices may be directly or indirectly connected to each other via a telephone network or short-range wireless etc. without going through the base stations ex106 to ex110, which are fixed radio stations. Also, the streaming server ex103 is connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101 etc. Further, the streaming server ex103 is connected to terminals etc. within a hotspot in an airplane ex117 via a satellite ex116.
[0448] Note that, instead of the base stations ex106 to ex110, a wireless access point, a hot spot, or the like may be used. Also, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.
[0449] The camera ex113 is a device capable of taking still images and videos such as a digital camera. Also, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handyphone System) or the like that supports the mobile communication system standards generally called 2G, 3G, 3.9G, 4G, and in the future 5G.
[0450] The home appliance ex118 is a device included in a refrigerator or a household fuel cell cogeneration system.
[0451] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 through the base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal in the airplane ex117, etc.) performs the encoding process described in the above embodiments on the still image or video content photographed by the user using the terminal, multiplexes the video data obtained by encoding with the audio data obtained by encoding the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present disclosure.
[0452] On the one hand, the streaming server ex103 streams the transmitted content data to the requested client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc., which is capable of decrypting the encoded data. Each device that receives the distributed data decrypts and plays back the received data. That is, each device functions as an image decrypting device according to an aspect of the present disclosure.
[0453] [Distributed processing] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data in a distributed manner. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world. In a CDN, an edge server physically close to the client is dynamically assigned according to the client. Then, by caching and distributing the content to the edge server, the delay can be reduced. Also, when some error occurs or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the distribution, so high-speed and stable distribution can be realized.
[0454] Furthermore, not limited to the distributed processing of the distribution itself, the encoding process of the captured data may be performed on each terminal, on the server side, or shared between them. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image or the amount of code is detected in units of frames or scenes. In the second loop, a process is performed to improve the encoding efficiency while maintaining the image quality. For example, by having the terminal perform the first encoding process and the server side that receives the content perform the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode almost in real time, the already encoded data from the first encoding performed by the terminal can be received and played back by other terminals, enabling more flexible real-time distribution.
[0455] As another example, cameras such as the ex113 perform feature extraction from the image, compress the data related to the features as metadata, and transmit it to the server. The server performs compression according to the meaning of the image, such as judging the importance of the object from the features and switching the quantization accuracy. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.
[0456] As yet another example, in a stadium, shopping mall, factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using the plurality of terminals that performed the shooting, and other terminals and servers that did not perform the shooting if necessary, encoding processes are respectively assigned and distributed, for example, in units of GOP (Group of Picture), picture, or tiles obtained by dividing the picture. This can reduce the delay and achieve more real-time performance.
[0457] Also, since the multiple video data is of almost the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can refer to each other. Or, the encoded data from each terminal may be received by the server, and the reference relationship may be changed among the multiple data, or the picture itself may be corrected or replaced and re-encoded. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.
[0458] Also, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert the MPEG-based encoding method to the VP-based one, or convert H.264 to H.265.
[0459] In this way, the encoding process can be performed by the terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the subject performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.
[0460] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and use different scenes captured by terminals such as multiple cameras ex113 and / or smartphones ex115 that are almost synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal are integrated based on the relative positional relationship between the terminals obtained separately or the regions where the feature points included in the videos match.
[0461] The server may not only encode two-dimensional moving images, but also automatically encode still images based on scene analysis of the moving images or at the time specified by the user, and transmit them to the receiving terminal. Further, when the server can obtain the relative positional relationship between the shooting terminals, it can generate the three-dimensional shape of the scene based not only on two-dimensional moving images, but also on videos shot from different angles of the same scene. Note that the server may separately encode three-dimensional data generated by point clouds or the like, or select or reconstruct the video to be transmitted to the receiving terminal based on the result of recognizing or tracking a person or an object using the three-dimensional data from the videos shot by a plurality of terminals.
[0462] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or enjoy the content obtained by cutting out the video from an arbitrary viewpoint from the three-dimensional data reconstructed using a plurality of images or videos. Further, similar to the video, sound is also collected from a plurality of different angles, and the server may multiplex and transmit the sound from a specific angle or space together with the video according to the video.
[0463] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become widespread. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by Multi-View Coding (MVC) or the like, or encode them as separate streams without referring to each other. At the time of decoding the separate streams, they may be played back in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.
[0464] In the case of an AR image, the server superimposes virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold the virtual object information and the three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting the virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like, and encode it in a state where the portion is transparent. Alternatively, the server may set the RGB value of a predetermined value like a chroma key as the background, and generate data with the portion other than the object being the background color.
[0465] Similarly, the decoding process of the distributed data may be performed on each client terminal, on the server side, or they may be shared. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request on another terminal, perform the decoding process, and send the decoded signal to the device having a display. By dispersing the process regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with good image quality. Also, as another example, while receiving large-size image data on a TV or the like, a part of the area such as a tile in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to check the area that one is in charge of or the area that one wants to check in more detail at hand.
[0466] In the future, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used regardless of indoors or outdoors, by using a distribution system standard such as MPEG-DASH, appropriate data can be switched for the ongoing communication, and it is expected that content can be received seamlessly. As a result, the user can freely select not only their own terminal but also a decoding device or display device such as a display installed indoors or outdoors and switch in real time. Also, based on their own location information, etc., decoding can be performed while switching the terminal to be decoded and the terminal to be displayed. This makes it possible to move while displaying map information on a part of the wall surface or ground of the adjacent building where a displayable device is embedded during movement to the destination. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or being copied to an edge server in a content delivery service, it is also possible to switch the bitrate of the received data.
[0467] [Scalable Encoding] Regarding content switching, it will be described using a scalable stream compressed and encoded by applying the moving image encoding method shown in each of the above embodiments shown in FIG. 45. The server may have a plurality of streams with the same content but different qualities as individual streams, but as shown in the figure, by encoding in layers, it is possible to utilize the characteristics of the temporal / spatial scalable stream realized, and it may be configured to switch content. That is, by determining up to which layer to decode according to internal factors such as performance on the decoding side and external factors such as the state of the communication bandwidth, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on your smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.
[0468] Furthermore, as described above, pictures are encoded layer by layer. In addition to the configuration that realizes scalability where an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of an image or the like, and the decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta information includes information for specifying linear or non-linear filter coefficients used in the super-resolution process, or information for specifying parameter values in filter processing, machine learning, or least-squares operation used in the super-resolution process.
[0469] Alternatively, a picture may be divided into tiles or the like according to the meaning of an object or the like in the image, and the decoding side may decode only a part of the area by selecting the tile to be decoded. Also, by storing the attributes of the object (such as a person, a car, a ball) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 46, the meta information is stored using a data storage structure different from pixel data such as an SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.
[0470] Also, the meta information may be stored in units composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining it with the information per picture, can specify the picture in which the object exists and the position of the object in the picture.
[0471] [Optimization of Web Page] FIG. 47 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 48 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 47 and 48, a web page may include a plurality of link images that are links to image contents, and the appearance thereof may vary depending on the device being viewed. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) may display a still image or an I picture that each content has as a link image, display a video like a gif animation with a plurality of still images or I pictures, etc., or receive only the base layer and decode and display the video.
[0472] When a link image is selected by the user, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Also, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay between the decoding time and the display time of the leading picture (the delay from the start of content decoding to the start of display) by decoding and displaying only the forward reference pictures (I pictures, P pictures, B pictures with only forward reference). Further, the display device may deliberately ignore the reference relationship of the pictures and roughly decode all B pictures and P pictures as forward references, and perform normal decoding as the received pictures increase over time.
[0473] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.
[0474] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal move, the receiving terminal can achieve seamless reception and decoding while switching between the base stations ex106 to ex110 by transmitting the position information of the receiving terminal at the time of a reception request. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication band.
[0475] As described above, in the content supply system ex100, the client can receive, decode, and reproduce the encoded information transmitted by the user in real time.
[0476] [Delivery of Personal Content] Also, in the content supply system ex100, not only high-quality and long-duration content by video delivery providers but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. Also, it is considered that such personal content will increase in the future. In order to make personal content into better content, the server may perform encoding processing after performing editing processing. This can be realized, for example, with the following configuration.
[0477] During shooting in real time or accumulating and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection on the original image or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone for editing. The server encodes the edited data based on the editing results. Also, it is known that if the shooting time is too long, the viewing rate will decrease. The server may automatically clip scenes with little movement as well as less important scenes as described above so that the content within a specific time range is obtained according to the shooting time, based on the image processing results. Or, the server may generate and encode a digest based on the result of semantic analysis of the scene.
[0478] Note that there are cases where personal content may contain elements that, as they are, would infringe copyright, moral rights of the author, or portrait rights, etc., and there may be inconveniences for individuals, such as the sharing range exceeding the intended range. Therefore, for example, the server may deliberately change the image to be out of focus for the faces of people in the peripheral part of the screen or inside the house and then encode it. Also, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, the user designates a person or background area that the user wants to process the image from the perspective of copyright, etc., and the server can perform processing such as replacing the designated area with another video or blurring the focus. In the case of a person, the video of the face part can be replaced while tracking the person in the moving image.
[0479] In addition, since viewing personal content with a small amount of data requires strong real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and decodes and plays it back. During this time, the decoding device receives the enhancement layer, and when playback is looped or played back two or more times, such as when the enhancement layer is also included, high-quality video may be played back. In the case of a stream encoded in a scalable manner like this, the video is rough when not selected or at the beginning of viewing, but it is possible to provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played back for the first time and a second stream encoded with reference to the first video are configured as one stream.
[0480] [Other usage examples] In addition, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processing may be performed using the software. Further, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.
[0481] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads the codec or application software and then acquires and plays back the content.
[0482] Moreover, not limited to the content supply system ex100 via the Internet ex101, at least either the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcast system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by loading it on broadcast radio waves using a satellite or the like, there is a difference in that it is suitable for multicast as opposed to the configuration of the content supply system ex100 that is easy to perform unicast, but the same application is possible for encoding processing and decoding processing.
[0483] [Hardware Configuration] FIG. 49 is a diagram showing a smartphone ex115. FIG. 50 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by the camera unit ex465 and video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with a SIM ex468 for identifying the user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.
[0484] In addition, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a bus ex470.
[0485] When the power key is turned on by the user's operation, the power supply circuit unit ex461 supplies power to each unit from the battery pack to activate the smartphone ex115 to an operable state.
[0486] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after being subjected to digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, spectrally despread by the modulation / demodulation unit ex452, converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. In the data communication mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body unit, and the same transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while the camera unit ex465 is capturing video or still images, etc., and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.
[0487] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and a video or a still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming is widespread, there may be a situation where it is not socially appropriate to play audio depending on the user's situation. Therefore, as an initial value, it is desirable to have a configuration that plays only the video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.
[0488] Also, although the smartphone ex115 has been described as an example here, as a terminal, in addition to a transceiver type terminal having both an encoder and a decoder, there are three implementation forms: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. Further, in the digital broadcast system, although it has been described as receiving or transmitting multiplexed data in which audio data and the like are multiplexed in video data, the multiplexed data may include character data related to video in addition to audio data, or the video data itself may be received or transmitted instead of the multiplexed data.
[0489] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, the terminal often has a GPU. Therefore, a configuration in which a wide area is processed in a batch by taking advantage of the performance of the GPU may be adopted using a memory shared by the CPU and the GPU or a memory whose addresses are managed so as to be commonly used. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform quantization in units such as pictures using the GPU instead of the CPU.
Industrial Applicability
[0490] The present disclosure can be applied to an encoding device, a decoding device, an encoding method, and a decoding method.
Description of Signs
[0491] 100 Encoding device 102 Splitting unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 160, 260 Processing circuit 162, 262 Memory 200 Decoding device 202 Entropy decoding unit
Claims
1. To filter the boundary between the first block and the second block adjacent to the first block, the values of a plurality of pixels of the first block and the values of a plurality of pixels of the second block are changed such that the amount of change in each value is within each clip width, With reference to a first picture including the first block and the second block after the filtering, a third block included in a second picture different from the first picture is encoded, The plurality of pixels of the first block and the plurality of pixels of the second block are arranged along a line intersecting the boundary, The plurality of pixels of the first block include a first pixel at a first position, and the plurality of pixels of the second block include a second pixel at a second position corresponding to the first position across the boundary, The clip width is selected based on a quantization parameter, The clip width includes a first clip width corresponding to the first pixel and a second clip width corresponding to the second pixel, and the first clip width and the second clip width are different, The first block and the second block are adjacent to each other on the left and right, and the boundary is a vertical boundary. Encoding method.
2. To filter the boundary between the first block and the second block adjacent to the first block, the values of a plurality of pixels of the first block and the values of a plurality of pixels of the second block are changed such that the amount of change in each value is within each clip width, With reference to a first picture including the first block and the second block after the filtering, a third encoded block included in a second picture different from the first picture is decoded, The plurality of pixels of the first block and the plurality of pixels of the second block are arranged along a line intersecting the boundary, The plurality of pixels of the first block include a first pixel at a first position, and the plurality of pixels of the second block include a second pixel at a second position corresponding to the first position across the boundary, The clip width is selected based on a quantization parameter, The clip width includes a first clip width corresponding to the first pixel and a second clip width corresponding to the second pixel, and the first clip width and the second clip width are different, The first block and the second block are adjacent to each other on the left and right, and the boundary is a vertical boundary. Decoding method.
Citation Information
Patent Citations
Information processor
JP2007180767A
Image processor, moving image decoding device, moving image encoding device, image processing method, moving image decoding method, and, moving image encoding method
JP2010081368A
Measurement-based and scalable deblock filtering of image data
JP2010141883A
Image processing device, image processing method, and program
WO2007013437A1
Image processor and image processing method
WO2011145601A1