Symbolizing device and decoding device

The encoding device optimizes prediction processing for moving images by using motion vectors in block and sub-block units, reducing processing volume while maintaining efficiency through selective correction processing, thus addressing the challenge of refined prediction without increased load.

JP7712431B2Active Publication Date: 2025-07-23PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Patent Information

Application Number
JP2024099396
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-02-06
Filing Date
2024-06-20
Publication Date
2025-07-23
Estimated Expiration
2039-01-30

AI Technical Summary

Technical Problem

Existing encoding technologies face challenges in performing refined prediction processing for moving images while keeping the processing amount in check.

Method used

An encoding device that performs prediction processing using motion vectors in block units or sub-block units, with the option to perform correction processing based on spatial gradients, and selectively omits correction processing in certain modes to manage processing volume.

Benefits of technology

The solution enables more refined prediction processing with reduced processing volume, enhancing encoding efficiency by combining motion compensation in block units with motion prediction in sub-block units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712431000003
    Figure 0007712431000003
  • Figure 0007712431000004
    Figure 0007712431000004
  • Figure 0007712431000005
    Figure 0007712431000005
Patent Text Reader

Abstract

To provide an encoding device and a decoding device that can perform more finely divided prediction processing while suppressing an increase in the amount of processing in video encoding and the like.SOLUTION: An encoding device 100 includes a circuit 160 and a memory 162. The circuit 160 determines whether to use one of a plurality of modes including a first mode in which prediction processing is performed on the basis of block-based motion vectors in a video image, and a second mode in which prediction processing is performed on the basis of sub-block-based motion vectors obtained by dividing a block for prediction processing. When prediction processing is performed in the first mode, the circuit 160 determines whether or not to perform correction processing of the predicted image using a spatial gradient of pixel values in the predicted image obtained by the prediction processing, and performs the correction processing when it is determined that the correction processing is to be performed. When prediction processing is performed in the second mode, the circuit 160 does not perform the correction processing, and the second mode is included in the merge mode.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an encoding device that encodes moving images and the like.

Background Art

[0002] Conventionally, as a standard for encoding moving images, there is H.265, also called HEVC (High Efficiency Video Coding) (Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in encoding moving images and the like, it is not easy to perform more refined prediction processing while suppressing an increase in the processing amount.

[0005] Therefore, the present disclosure provides an encoding device and the like that can perform more refined prediction processing while suppressing an increase in the processing amount in encoding moving images and the like.

Means for Solving the Problems

[0006] An encoding device according to one aspect of the present disclosure is an encoding device that performs prediction processing to encode a moving image, and includes a circuit and a memory. The circuit uses the memory to perform the prediction processing based on motion vectors in block units in the moving image in a first mode, and based on motion vectors in sub-block units obtained by dividing the block in a second mode. Among a plurality of modes including the first mode and the second mode, it determines which mode to use for the prediction processing. When performing the prediction processing in the first mode, it determines whether to perform correction processing on the prediction image using the spatial gradient of pixel values in the prediction image obtained by performing the prediction processing. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction processing in the second mode, the correction processing is not performed. The second mode is included in a merge mode that uses the motion vector of an adjacent block adjacent to the block as a motion vector.

[0007] Note that these general or specific aspects may be implemented by a system, apparatus, method, integrated circuit, computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.

Effect of the Invention

[0008] An encoding device and the like according to one aspect of the present disclosure can perform more refined prediction processing while suppressing an increase in the processing amount in encoding a moving image and the like.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

DETAILED DESCRIPTION OF THE INVENTION

[0010] (Knowledge underlying the present disclosure) For example, when an encoding device that encodes a moving image encodes a moving image that performs more refined prediction processing while suppressing an increase in processing volume in the encoding of the moving image or the like, a prediction error is derived by subtracting a predicted image from an image constituting the moving image. Then, the encoding device performs frequency conversion and quantization on the prediction error, and encodes the result as image data. At this time, if motion prediction processing is performed in units of blocks or in units of sub-blocks constituting the blocks for encoding target units such as blocks included in the moving image, and further motion correction processing is performed in minute units, the encoding accuracy is improved.

[0011] However, in the encoding or the like of blocks included in a moving image, if appropriate subdivision prediction processing is not performed, it leads to an increase in processing volume and a decrease in encoding efficiency.

[0012] Therefore, an encoding device according to an aspect of the present disclosure is an encoding device that performs prediction processing to encode a moving image, and includes a circuit and a memory. The circuit uses the memory to determine which mode to use for the prediction processing among a plurality of modes including a first mode in which the prediction processing is performed based on a motion vector in units of blocks in the moving image, and a second mode in which the prediction processing is performed based on a motion vector in units of sub-blocks obtained by dividing the blocks. When performing the prediction processing in the first mode, it is determined whether to perform correction processing on the predicted image using the spatial gradient of pixel values in the predicted image obtained by performing the prediction processing. When it is determined that the correction processing is to be performed, the correction processing is performed. When performing the prediction processing in the second mode, it may be an encoding device that does not perform the correction processing.

[0013] As a result, the encoding device uses motion compensation in tiny units in combination with motion prediction in block units, thereby improving the encoding efficiency. Also, since the motion prediction in sub-block units has a larger processing volume than the motion prediction in block units, when the encoding device performs motion prediction in sub-block units, it does not perform motion compensation in tiny units. Therefore, the encoding device can reduce the processing volume while maintaining the encoding efficiency by executing motion prediction in tiny units only for the motion prediction in block units. Accordingly, the encoding device can perform more refined prediction processing while suppressing an increase in the processing volume.

[0014] For example, the first mode and the second mode may be included in a merge mode that is a mode using a predicted motion vector as a motion vector.

[0015] As a result, the encoding device can speed up the process for deriving a predicted sample set in the merge mode.

[0016] Also, for example, when the circuit performs the prediction process in the first mode, it encodes determination result information indicating a determination result as to whether to perform the prediction process, and when performing the prediction process in the second mode, it may not encode the determination result information.

[0017] As a result, the encoding device can reduce the amount of code.

[0018] Also, for example, the correction process may be a BIO (BI-directional Optical flow) process.

[0019] As a result, the encoding device can correct the predicted image using correction values in tiny units in the predicted image generated by deriving a motion vector in block units.

[0020] Also, for example, the second mode may be an ATMVP (Advanced Temporal Motion Vector Prediction) mode.

[0021] As a result, since the encoding device does not need to perform motion compensation processing in minute units in the ATMVP mode, the processing amount is reduced.

[0022] Also, for example, the second mode may be an STMVP (Spatial-Temporal Motion Vector Prediction) mode.

[0023] As a result, since the encoding device does not need to perform motion compensation processing in minute units in the STMVP mode, the processing amount is reduced.

[0024] Also, for example, the second mode may be an affine motion compensation prediction mode.

[0025] As a result, since the encoding device does not need to perform motion compensation processing in minute units in the affine mode, the processing amount is reduced.

[0026] Also, a decoding device according to an aspect of the present disclosure is a decoding device that performs prediction processing to decode a moving image, and includes a circuit and a memory. The circuit uses the memory to perform the prediction processing based on a motion vector in block units in the moving image, and a first mode for performing the prediction processing based on a motion vector in sub-block units obtained by dividing the block. Among a plurality of modes including a second mode, it is determined which mode to use for the prediction processing. When performing the prediction processing in the first mode, it is determined whether to perform correction processing on the prediction image using a spatial gradient of pixel values in the prediction image obtained by performing the prediction processing. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction processing in the second mode, it may be a decoding device that does not perform the correction processing.

[0027] As a result, the decoding device can improve the coding efficiency by using motion compensation in tiny units in combination with motion prediction in block units. Also, since the motion prediction in sub-block units has a larger processing volume than the motion prediction in block units, when the decoding device performs motion prediction in sub-block units, it does not perform motion compensation in tiny units. Therefore, the decoding device can reduce the processing volume while maintaining the coding efficiency by performing motion prediction in tiny units only for the motion prediction in block units. Thus, the decoding device can perform more refined prediction processing while suppressing an increase in the processing volume.

[0028] For example, the first mode and the second mode may be included in a merge mode which is a mode using a predicted motion vector as a motion vector.

[0029] As a result, the decoding device can speed up the process for deriving a predicted sample set in the merge mode.

[0030] Also, for example, when performing the prediction process in the first mode, decode determination result information indicating a determination result as to whether to perform the correction process, and when performing the prediction process in the second mode, it may not be necessary to decode the determination result information.

[0031] As a result, the decoding device can improve the processing efficiency.

[0032] Also, for example, the correction process may be a BIO process.

[0033] As a result, the decoding device can correct a predicted image using a correction value in tiny units in the predicted image generated by deriving a motion vector in block units.

[0034] Also, for example, the second mode may be an ATMVP mode.

[0035] As a result, since the decoding device does not need to perform motion correction processing in minute units in the ATMVP mode, the processing amount is reduced.

[0036] Also, for example, the second mode may be the STMVP mode.

[0037] As a result, since the decoding device does not need to perform motion correction processing in minute units in the STMVP mode, the processing amount is reduced.

[0038] Also, for example, the second mode may be the affine mode.

[0039] As a result, since the decoding device does not need to perform motion correction in minute units in the affine mode, the processing amount is reduced.

[0040] Further, an encoding method according to an aspect of the present disclosure is an encoding method for encoding a moving image by performing prediction processing, including a first mode of performing the prediction processing based on a motion vector in block units in the moving image, and a second mode of performing the prediction processing based on a motion vector in sub-block units obtained by dividing the block. Among a plurality of modes, it is determined which mode is used to perform the prediction processing. When performing the prediction processing in the first mode, it is determined whether to perform correction processing on the prediction image using the spatial gradient of pixel values in the prediction image obtained by performing the prediction processing. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction processing in the second mode, it may be an encoding method that does not perform the correction processing.

[0041] As a result, by using the motion correction process for minute units in combination with the motion prediction process for block units, the encoding efficiency is improved. Also, since the motion prediction for sub-block units has a larger processing volume than the motion prediction for block units, in the encoding method, when performing the motion prediction process for sub-block units, the motion correction for minute units is not performed. Therefore, according to the encoding method, by executing the motion prediction process for minute units only when performing the motion prediction process for block units, it is possible to reduce the processing volume while maintaining the encoding efficiency. Thus, according to the encoding method, it is possible to perform more refined prediction processing while suppressing an increase in the processing volume.

[0042] Also, a decoding method according to an aspect of the present disclosure is a decoding method for performing prediction processing to decode a moving image, including a first mode of performing the prediction processing based on a motion vector for block units in the moving image, and a second mode of performing the prediction processing based on a motion vector for sub-block units obtained by dividing the block, determining which mode among a plurality of modes to use for performing the prediction processing, when performing the prediction processing in the first mode, determining whether to perform correction processing on the prediction image using the spatial gradient of pixel values in the prediction image obtained by performing the prediction processing, performing the correction processing when it is determined to perform the correction processing, and when performing the prediction processing in the second mode, it may be a decoding method that does not perform the correction processing.

[0043] As a result, by using the motion correction process for minute units in combination with the motion prediction process for block units, the encoding efficiency is improved. Also, since the motion prediction for sub-block units has a larger processing volume than the motion prediction for block units, in the decoding method, when performing the motion prediction for sub-block units, the motion correction for minute units is not performed. Therefore, according to the decoding method, by executing the motion prediction process for minute units only when performing the motion prediction process for block units, it is possible to reduce the processing volume while maintaining the encoding efficiency. Thus, according to the decoding method, it is possible to perform more refined prediction processing while suppressing an increase in the processing volume.

[0044] Further, for example, an encoding device according to an aspect of the present disclosure is an encoding device that encodes a moving image, and may include a splitting unit, an intra prediction unit, an inter prediction unit, a conversion unit, a quantization unit, an entropy encoding unit, and a loop filter unit.

[0045] The splitting unit may split a picture included in the moving image into a plurality of blocks. The intra prediction unit may perform intra prediction on the blocks included in the plurality of blocks. The inter prediction unit may perform inter prediction on the blocks. The conversion unit may convert a prediction error between a prediction image obtained by the intra prediction or the inter prediction and an original image to generate conversion coefficients. The quantization unit may quantize the conversion coefficients to generate quantized coefficients. The entropy encoding unit may encode the quantized coefficients to generate an encoded bit stream. The loop filter unit may apply a filter to a reconstructed image generated using the prediction image.

[0046] Then, for example, the inter prediction unit determines which mode to use for the prediction process among a plurality of modes including a first mode in which the prediction process is performed based on a motion vector in units of blocks in the moving image, and a second mode in which the prediction process is performed based on a motion vector in units of sub-blocks obtained by splitting the blocks. When performing the prediction process in the first mode, it is determined whether to perform correction processing on the prediction image using a spatial gradient of pixel values in the prediction image obtained by performing the prediction process. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction process in the second mode, the correction processing is not performed.

[0047] Further, for example, a decoding device according to an aspect of the present disclosure is a decoding device that decodes a moving image, and may include an entropy decoding unit, an inverse quantization unit, an inverse conversion unit, an intra prediction unit, an inter prediction unit, and a loop filter unit.

[0048] The entropy decoding unit may decode the quantization coefficients of blocks within a picture from the encoded bitstream. The inverse quantization unit may perform inverse quantization on the quantization coefficients to obtain transform coefficients. The inverse transform unit may perform inverse transform on the transform coefficients to obtain prediction errors. The intra prediction unit may perform intra prediction on the block. The inter prediction unit may perform inter prediction on the block. The loop filter unit may apply a filter to a reconstructed image generated using the prediction image obtained by the intra prediction or the inter prediction and the prediction error.

[0049] Then, for example, the inter prediction unit determines which mode to use for the prediction process among a plurality of modes including a first mode of performing the prediction process based on the motion vector in units of blocks in the moving image, and a second mode of performing the prediction process based on the motion vector in units of sub-blocks obtained by dividing the block. When performing the prediction process in the first mode, it is determined whether to perform correction processing on the prediction image using the spatial gradient of the pixel values in the prediction image obtained by performing the prediction process. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction process in the second mode, the correction processing is not performed.

[0050] Furthermore, these general or specific aspects may be implemented by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0051] Hereinafter, embodiments will be specifically described with reference to the drawings.

[0052] Note that all the embodiments described below show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are just examples and are not intended to limit the scope of the claims. Also, among the components in the following embodiments, the components not described in the independent claims indicating the top-level concept are described as optional components.

[0053] (Embodiment 1) First, as an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure to be described later are applicable, an overview of Embodiment 1 will be described. However, Embodiment 1 is only an example of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable, and the processes and / or configurations described in each aspect of the present disclosure can also be implemented in encoding devices and decoding devices different from Embodiment 1.

[0054] When applying the processes and / or configurations described in each aspect of the present disclosure to Embodiment 1, for example, any of the following may be performed.

[0055] (1) For the encoding device or the decoding device of Embodiment 1, replacing the component corresponding to the component described in each aspect of the present disclosure among the plurality of components constituting the encoding device or the decoding device with the component described in each aspect of the present disclosure (2) After making any changes such as addition, replacement, deletion, etc. of the functions or processes implemented for some of the components among the plurality of components constituting the encoding device or the decoding device of Embodiment 1, replacing the component corresponding to the component described in each aspect of the present disclosure with the component described in each aspect of the present disclosure For the method implemented by the encoding device or decoding device of Embodiment 1, after adding processing and / or making any changes such as replacement or deletion to some of the multiple processes included in the method, replace the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure. (4) Combine and implement some of the components among the multiple components constituting the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components implementing a part of the processes implemented by the components described in each aspect of the present disclosure. (5) Combine and implement components having a part of the functions provided by some of the components among the multiple components constituting the encoding device or decoding device of Embodiment 1, or components implementing a part of the processes implemented by some of the components among the multiple components constituting the encoding device or decoding device of Embodiment 1 with the components described in each aspect of the present disclosure, components having a part of the functions provided by the components described in each aspect of the present disclosure, or components implementing a part of the processes implemented by the components described in each aspect of the present disclosure. (6) For the method implemented by the encoding device or decoding device of Embodiment 1, replace, among the multiple processes included in the method, the processes corresponding to the processes described in each aspect of the present disclosure with the processes described in each aspect of the present disclosure. (7) Implement some of the multiple processes included in the method implemented by the encoding device or decoding device of Embodiment 1 in combination with the processes described in each aspect of the present disclosure.

[0056] Note that the implementation manners of the processes and / or configurations described in each aspect of the present disclosure are not limited to the above examples. For example, it may be implemented in a device used for a purpose different from the moving image / image encoding device or moving image / image decoding device disclosed in Embodiment 1, or the processes and / or configurations described in each aspect may be implemented alone. Also, the processes and / or configurations described in different aspects may be implemented in combination.

[0057] [Overview of the Encoding Device] First, the overview of the encoding device according to Embodiment 1 will be described. FIG. 1 is a block diagram showing the functional configuration of an encoding device 100 according to Embodiment 1. The encoding device 100 is a moving image / image encoding device that encodes moving images / images in block units.

[0058] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0059] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0060] Hereinafter, each component included in the encoding device 100 will be described.

[0061] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (for example, 128x128). These blocks of fixed size are sometimes called coding tree units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of variable size (for example, 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These blocks of variable size are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in the present embodiment, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be processing units of CUs, PUs, and TUs.

[0062] FIG. 2 is a diagram showing an example of block splitting in Embodiment 1. In FIG. 2, the solid lines represent block boundaries by quadtree block splitting, and the dashed lines represent block boundaries by binary tree block splitting.

[0063] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first split into four square 64x64 blocks (quadtree block splitting).

[0064] The upper-left 64x64 block is further vertically split into two rectangular 32x64 blocks, and the left 32x64 block is further vertically split into two rectangular 16x64 blocks (binary tree block splitting). As a result, the upper-left 64x64 block is split into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0065] The upper-right 64x64 block is horizontally split into two rectangular 64x32 blocks 14 and 15 (binary tree block splitting).

[0066] The bottom - left 64x64 block is divided into four square 32x32 blocks (quad - tree block division). Among the four 32x32 blocks, the upper - left block and the lower - right block are further divided. The upper - left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary - tree block division). The lower - right 32x32 block is horizontally divided into two 32x16 blocks (binary - tree block division). As a result, the bottom - left 64x64 block is divided into 16 16x32 blocks 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0067] The bottom - right 64x64 block 23 is not divided.

[0068] As described above, in FIG. 2, block 10 is divided into 13 variable - size blocks 11 - 23 based on recursive quad - tree and binary - tree block division. Such a division is sometimes called QTBT (quad - tree plus binary tree) division.

[0069] In FIG. 2, one block is divided into four or two blocks (quad - tree or binary - tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary - tree block division). A division including such a ternary - tree block division is sometimes called MBT (multi - type tree) division.

[0070] [Subtraction unit] The subtraction unit 104 subtracts the predicted signal (predicted sample) from the original signal (original sample) in block units divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error to the conversion unit 106.

[0071] The original signal is the input signal of the encoding device 100, and is a signal representing the images of each picture constituting a moving image (for example, a luma signal and two chroma signals). Hereinafter, the signal representing an image may also be referred to as a sample.

[0072] [Conversion Unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain, and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0073] Note that the conversion unit 106 may adaptively select a conversion type from a plurality of conversion types, and convert the prediction error into conversion coefficients using a transform basis function corresponding to the selected conversion type. Such a conversion may be referred to as an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).

[0074] The plurality of conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the conversion basis functions corresponding to each conversion type. In FIG. 3, N indicates the number of input pixels. The selection of the conversion type from these plurality of conversion types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.

[0075] Information indicating whether to apply such an EMT or AMT (for example, called an AMT flag) and information indicating the selected conversion type are signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0076] Further, the conversion unit 106 may re-convert the conversion coefficient (conversion result). Such re-conversion is sometimes referred to as AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (e.g., 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are signaled at the CU level. Note that the signaling of these pieces of information is not necessarily limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0077] Here, a separable conversion is a method of performing multiple conversions by separating them for each direction by the number of dimensions of the input, and a non-separable conversion is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing the conversion collectively.

[0078] For example, as an example of a non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.

[0079] Similarly, after regarding a 4×4 input block as an array having 16 elements, a method of performing multiple Givens rotations on the array (Hypercube Givens Transform) is also an example of a non-separable conversion.

[0080] [Quantization unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order, and quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0081] The predetermined order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0082] The quantization parameter is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0083] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) by performing variable-length encoding on the quantization coefficients that are the input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantization coefficients and performs arithmetic encoding on the binary signal.

[0084] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantization coefficients that are the input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized conversion coefficients of the current block to the inverse conversion unit 114.

[0085] [Inverse Conversion Unit] The inverse transformation unit 114 restores the prediction error by inversely transforming the transformation coefficients that are the input from the inverse quantization unit 112. Specifically, the inverse transformation unit 114 restores the prediction error of the current block by performing an inverse transformation corresponding to the transformation by the transformation unit 106 on the transformation coefficients. Then, the inverse transformation unit 114 outputs the restored prediction error to the addition unit 116.

[0086] Note that since information is lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error includes a quantization error.

[0087] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction error that is the input from the inverse transformation unit 114 and the prediction sample that is the input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.

[0088] [Block memory] The block memory 118 is a storage unit for storing blocks within the coded target picture (hereinafter referred to as the current picture) that are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.

[0089] [Loop filter unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0090] In ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 sub-block within a current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.

[0091] Specifically, first, sub-blocks (for example, 2x2 sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes (for example, 15 or 25 classes).

[0092] The gradient direction value D is derived, for example, by comparing gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding gradients in a plurality of directions and quantizing the addition result.

[0093] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0094] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is signaled at the picture level. Note that the signaling of the information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (for example, sequence level, slice level, tile level, CTU level, or CU level).

[0095] The on / off of ALF is determined, for example, at the picture level or the CU level. For example, for luminance, it is determined whether to apply ALF at the CU level, and for chrominance, it is determined whether to apply ALF at the picture level. The information indicating the on / off of ALF is signaled at the picture level or the CU level. Note that the signaling of the information indicating the on / off of ALF does not have to be limited to the picture level or the CU level, and it may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0096] The coefficient sets of a plurality of selectable filters (e.g., filters up to 15 or 25) are signaled at the picture level. Note that the signaling of the coefficient sets does not have to be limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0097] [Frame Memory] The frame memory 122 is a storage unit for storing reference pictures used for inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0098] [Intra Prediction Unit] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to the blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to the samples (e.g., luminance values, chrominance values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0099] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0100] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC (High-Efficiency Video Coding) standard (Non-Patent Document 1).

[0101] The plurality of directional prediction modes include, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 5A is a diagram showing 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions.

[0102] In addition, in the intra prediction of the chrominance blocks, the luminance blocks may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0103] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of reference pixels in the horizontal / vertical direction. Intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of PDPC application (for example, called a PDPC flag) is signaled at, for example, the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0104] [Inter prediction unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. Inter prediction is performed in units of the current block or sub-blocks (for example, 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or sub-block. Then, the inter prediction unit 126 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (for example, motion vector) obtained by the motion estimation. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0105] The motion information used for motion compensation is signaled. A motion vector predictor may be used for signaling the motion vector. That is, the difference between the motion vector and the predicted motion vector may be signaled.

[0106] In addition to the motion information of the current block obtained by motion search, the motion information of adjacent blocks may also be used to generate an inter prediction signal. Specifically, an inter prediction signal may be generated for each sub-block in the current block by weighted addition of a prediction signal based on the motion information obtained by motion search and a prediction signal based on the motion information of adjacent blocks. Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0107] In such an OBMC mode, information indicating the size of the sub-blocks for OBMC (for example, called OBMC block size) is signaled at the sequence level. Also, information indicating whether or not to apply the OBMC mode (for example, called OBMC flag) is signaled at the CU level. Note that the signaling levels of these pieces of information do not necessarily have to be limited to the sequence level and the CU level, and may be other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).

[0108] The OBMC mode will be described in more detail. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process by OBMC processing.

[0109] First, a predicted image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded.

[0110] Next, the motion vector (MV_L) of the encoded left adjacent block is applied to the block to be coded to obtain a predicted image (Pred_L), and the first correction of the predicted image is performed by weighting and superimposing the predicted image and Pred_L.

[0111] Similarly, the motion vector (MV_U) of the encoded upper adjacent block is applied to the block to be encoded to obtain a predicted image (Pred_U). The predicted image after the first correction and Pred_U are weighted and superimposed to perform the second correction of the predicted image, which is used as the final predicted image.

[0112] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to configure to perform more corrections than two stages using the right adjacent block or the lower adjacent block.

[0113] Note that the area for superimposing may be only a partial area near the block boundary, rather than the pixel area of the entire block.

[0114] Here, the prediction image correction process from a single reference picture has been described. However, the same applies to the case of correcting the prediction image from multiple reference pictures. After obtaining the prediction images corrected from each reference picture, the obtained prediction images are further superimposed to obtain the final prediction image.

[0115] Note that the block to be processed may be in units of prediction blocks or in units of sub-blocks obtained by further dividing the prediction blocks.

[0116] As a method for determining whether to apply the OBMC process, for example, there is a method using an obmc_flag, which is a signal indicating whether to apply the OBMC process. As a specific example, in the encoding device, it is determined whether the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, the value 1 is set as the obmc_flag and the OBMC process is applied for encoding. If it does not belong to a region with complex motion, the value 0 is set as the obmc_flag and encoding is performed without applying the OBMC process. On the other hand, in the decoding device, the obmc_flag described in the stream is decoded, and decoding is performed by switching whether to apply the OBMC process according to the value.

[0117] Note that the motion information may be derived on the decoder side without being signaled. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also, for example, the motion information may be derived by performing motion search on the decoder side. In this case, the motion search is performed without using the pixel values of the current block.

[0118] Here, the mode in which motion search is performed on the decoder side will be described. The mode in which motion search is performed on the decoder side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.

[0119] An example of FRUC processing is shown in FIG. 5D. First, by referring to the motion vectors of the encoded blocks spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate included in the candidate list is calculated, and one candidate is selected based on the evaluation value.

[0120] Then, based on the motion vector of the selected candidate, the motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, the motion vector for the current block may be derived. That is, search is performed in the same manner for the region around the best candidate MV, and if there is an MV for which the evaluation value is a good value, the best candidate MV may be updated to the MV, and that may be used as the final MV of the current block. Note that a configuration in which the said process is not performed is also possible.

[0121] The same processing may be performed even when processing is performed in sub-block units.

[0122] The evaluation value is calculated by obtaining the difference value of the reconstructed image by pattern matching between the region in the reference picture corresponding to the motion vector and a predetermined region. Note that the evaluation value may be calculated using information other than the difference value.

[0123] As the pattern matching, first pattern matching or second pattern matching is used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.

[0124] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the candidate described above, the region in another reference picture along the motion trajectory of the current block is used.

[0125] FIG. 6 is a diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and an evaluation value is calculated using the obtained difference value. It is preferable to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV.

[0126] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0127] In the second pattern matching, pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as the predetermined region for calculating the evaluation value of the above-described candidate.

[0128] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of a current block is derived by searching, in a reference picture (Ref0), for a block that most closely matches a block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, a difference is derived between a reconstructed image of an encoded region of both or either of the left and upper adjacent blocks and a reconstructed image at an equivalent position in the encoded reference picture (Ref0) specified by a candidate MV, and an evaluation value is calculated using the obtained difference value. It is preferable to select, as the best candidate MV, a candidate MV having the best evaluation value among a plurality of candidate MVs.

[0129] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) is signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating a pattern matching method (first pattern matching or second pattern matching) (for example, called a FRUC mode flag) is signaled at the CU level. Note that the signaling of this information does not necessarily have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0130] Here, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode may be called the BIO (bi - directional optical flow) mode.

[0131] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v yshows the velocity vector, where τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.

[0132] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0133]

Equation

[0134] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from the merge list, etc., are corrected in pixel units.

[0135] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0136] Here, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be referred to as an affine motion compensation prediction mode.

[0137] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, based on the motion vectors of the adjacent blocks, the motion vector v0 of the upper left control point of the current block is derived, and based on the motion vectors of the adjacent sub-blocks, the motion vector v1 of the upper right control point of the current block is derived. Then, using the two motion vectors v0 and v1, the motion vector (v x ,v y ) of each sub-block within the current block is derived by the following equation (2).

[0138] [Equation]

[0139] Here, x and y indicate the horizontal position and vertical position of the sub-block, respectively, and w indicates a predetermined weight coefficient.

[0140] Such an affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right control points are different. Information indicating such an affine motion compensation prediction mode (for example, called an affine flag) is signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode is not necessarily limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0141] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal or an inter prediction signal, and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0142] Here, an example of deriving the motion vector of the picture to be coded in the merge mode will be described. FIG. 9B is a diagram for explaining the outline of the motion vector derivation process in the merge mode.

[0143] First, a prediction MV list in which candidates for the prediction MV are registered is generated. Examples of candidates for the prediction MV include a spatial adjacent prediction MV which is the MV of a plurality of coded blocks located spatially adjacent to the block to be coded, a temporal adjacent prediction MV which is the MV of a nearby block obtained by projecting the position of the block to be coded in the coded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, and a zero prediction MV which is an MV with a value of zero.

[0144] Next, one prediction MV is selected from among the plurality of prediction MVs registered in the prediction MV list, and is determined as the MV of the block to be coded.

[0145] Furthermore, in the variable length coding unit, a merge_idx which is a signal indicating which prediction MV has been selected is described in the stream and coded.

[0146] Note that the prediction MVs registered in the prediction MV list described in FIG. 9B are only examples, and the number may be different from that in the figure, or the configuration may not include some of the types of prediction MVs in the figure, or may include prediction MVs other than the types of prediction MVs in the figure.

[0147] Note that the final MV may be determined by performing the DMVR process described later using the MV of the block to be coded derived in the merge mode.

[0148] Here, an example of determining the MV using the DMVR process will be described.

[0149] FIG. 9C is a conceptual diagram for explaining the outline of the DMVR process.

[0150] First, using the optimal MVP set for the processing target block as a candidate MV, according to the candidate MV, reference pixels are respectively obtained from the first reference picture which is the processed picture in the L0 direction and the second reference picture which is the processed picture in the L1 direction, and a template is generated by taking the average of each reference pixel.

[0151] Next, using the template, the peripheral areas of the candidate MVs of the first reference picture and the second reference picture are respectively searched, and the MV with the minimum cost is determined as the final MV. Note that the cost value is calculated using the difference value between each pixel value of the template and each pixel value of the search area, the MV value, etc.

[0152] Note that in the encoding device and the decoding device, the outline of the processing described here is basically common.

[0153] Note that even if it is not the processing itself described here, other processing may be used as long as it is a processing that can search the periphery of the candidate MV to derive the final MV.

[0154] Here, the mode of generating a predicted image using the LIC process will be described.

[0155] FIG. 9D is a diagram for explaining the outline of a predicted image generation method using the luminance correction process by the LIC process.

[0156] First, an MV for obtaining a reference image corresponding to the encoding target block is derived from the reference picture which is the encoded picture.

[0157] Next, for the block to be encoded, using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, information indicating how the luminance values change between the reference picture and the picture to be encoded is extracted to calculate the luminance correction parameter.

[0158] By performing luminance correction processing on the reference image in the reference picture specified by the MV using the luminance correction parameter, a predicted image for the block to be encoded is generated.

[0159] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0160] Also, although the process of generating a predicted image from a single reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures. Luminance correction processing is performed on the reference images obtained from each reference picture in the same manner and then a predicted image is generated.

[0161] As a method for determining whether to apply the LIC process, for example, there is a method using lic_flag, which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the block to be encoded belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value 0 is set as lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding is performed by switching whether to apply the LIC process according to the value.

[0162] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to the peripheral blocks. As a specific example, when the block to be encoded is in the merge mode, it is determined whether the peripheral encoded blocks selected at the time of deriving the MV in the merge mode process are encoded by applying the LIC process, and encoding is performed by switching whether to apply the LIC process according to the result. In the case of this example, the process in decoding is exactly the same.

[0163] [Overview of the Decoder] Next, an overview of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing the functional configuration of the decoder 200 according to the first embodiment. The decoder 200 is a moving image / image decoder that decodes moving images / images in units of blocks.

[0164] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0165] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0166] Hereinafter, each component included in the decoder 200 will be described.

[0167] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bit stream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. As a result, the entropy decoding unit 202 outputs quantization coefficients in block units to the inverse quantization unit 204.

[0168] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block), which is the input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0169] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.

[0170] For example, when the information decoded from the encoded bit stream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the decoded transform type.

[0171] Also, for example, when the information decoded from the encoded bit stream indicates that NSST is to be applied, the inverse transform unit 206 applies an inverse reverse transform to the transform coefficients.

[0172] [Addition Unit] The adder 208 reconstructs the current block by adding the prediction error, which is the input from the inverse transform unit 206, and the prediction sample, which is the input from the predictive control unit 220. Then, the adder 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0173] [Block Memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture) that are referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the adder 208.

[0174] [Loop Filter Unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.

[0175] When the information indicating the on / off of the ALF read from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0176] [Frame Memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0177] [Intra Prediction Unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0178] Note that when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0179] Also, when the information decoded from the encoded bit stream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0180] [Inter Prediction Unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bit stream, and outputs the inter prediction signal to the prediction control unit 220.

[0181] Note that when the information decoded from the encoded bit stream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.

[0182] Also, when the information decoded from the encoded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation using the derived motion information.

[0183] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0184] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the addition unit 208 as the prediction signal.

[0185] [Details of inter prediction processing] Next, the details of the inter prediction processing will be described.

[0186] For example, in the encoding device 100, the inter prediction unit 126 determines which of a plurality of modes including a first mode that performs prediction processing based on the motion vector in block units in a moving image and a second mode that performs prediction processing based on the motion vector in sub-block units obtained by dividing the block is used for prediction processing. Then, when performing prediction processing in the first mode, the inter prediction unit 126 determines whether to perform correction processing on the prediction image using the spatial gradient of the pixel values in the prediction image obtained by performing the prediction processing, and performs the correction processing when it is determined to perform the correction processing. On the other hand, when performing prediction processing in the second mode, the inter prediction unit 126 does not perform correction processing.

[0187] Then, based on the above prediction process, the inter prediction unit 126 derives a prediction sample set for the CU to be coded. Thereafter, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy coding unit 110, etc. code the CU to be coded using the prediction sample set.

[0188] Also, the inter prediction process in the decoding apparatus 200 is the same as the inter prediction process in the encoding apparatus 100. For example, in the decoding apparatus 200, the inter prediction unit 218 determines which mode among a plurality of modes including a first mode of performing a prediction process based on a motion vector in block units in a moving image and a second mode of performing a prediction process based on a motion vector in sub-block units obtained by dividing a block is used for the prediction process. Then, when performing the prediction process in the first mode, the inter prediction unit 126 determines whether to perform correction processing on the prediction image using the spatial gradient of the pixel values in the prediction image obtained by performing the prediction process, and performs the correction processing when it is determined to perform the correction processing. On the other hand, when performing the prediction process in the second mode, the inter prediction unit 126 does not perform the correction processing. Note that in the decoding apparatus 200, when performing the prediction process in the first mode, the inter prediction unit 218 determines whether to perform the correction processing using the determination result information indicating the determination result of whether to perform the correction processing on the prediction image using the spatial gradient of the pixel values in the prediction image obtained by performing the prediction process in the encoding apparatus 100.

[0189] Then, based on the above prediction process, the inter prediction unit 218 derives a prediction sample set for the CU to be coded. Thereafter, the entropy decoding unit 202, the inverse quantization unit 204, the inverse conversion unit 206, the addition unit 208, etc. decode the CU to be coded using the prediction sample set.

[0190] Hereinafter, the details of the inter prediction process will be described more specifically.

[0191] FIG. 11 is a flowchart showing an example of operations performed by the encoding apparatus 100 and the decoding apparatus 200 in the first aspect.

[0192] The operations performed by the encoding device 100 will be described below. The operations performed by the decoding device 200 are the same as those performed by the encoding device 100. Also, the processing related to inter prediction is mainly performed by the inter prediction unit 126 in the encoding device 100, and mainly performed by the inter prediction unit 218 in the decoding device 200.

[0193] As shown in FIG. 11, the operation of the inter prediction unit 126 of the encoding device in the first aspect is characterized by the operation in the merge mode. In inter prediction, there are a mode (so-called, the first mode) for generating candidate MVs corresponding to the CU to be encoded in units of CUs, and a mode (so-called, the second mode) for generating candidate MVs in units of sub-CUs obtained by dividing the CU into NxN blocks. In the mode (the second mode) for generating candidate MVs in units of sub-CUs, motion prediction is performed in units of sub-CUs, and motion prediction at the pixel level is not performed in the subsequent stage. In other words, in the second mode, the MV is derived in units of sub-blocks, and after generating a predicted image by performing motion compensation processing (MC processing) in units of sub-blocks using the derived MV, the predicted image is not corrected at the pixel level. Therefore, when encoding the CU to be encoded in the second mode, it is not necessary to encode the determination result information (for example, called a flag) regarding whether or not to perform motion correction processing at the pixel level in the predicted image.

[0194] On the other hand, in the mode (the first mode) of generating candidate MVs in CU units, after performing motion prediction in CU units, motion prediction in pixel units is carried out to correct the motion prediction result in CU units in pixel units. In other words, in the first mode, MVs are derived in block units, and after generating a predicted image by performing motion compensation processing (MC processing) in block units using the derived MVs, correction processing of the predicted image is performed using the spatial gradient of pixel values in the predicted image. Here, whether to perform motion prediction in pixel units may be a selectable configuration. Therefore, when encoding the CU to be encoded in the first mode, a flag regarding whether to perform motion correction processing in pixel units in the predicted image may be encoded. Also, the CU may be non-square such as MxN, and the sub-CU may be a unit obtained by dividing the CU into an arbitrary shape. As the motion correction processing in pixel units, for example, a method such as BIO (Bi-directional optical flow) can be used. Note that the motion correction processing in pixel units may be performed for each pixel, or may be corrected in multiple pixel units. For example, the multiple pixel units may be in block units or sub-block units.

[0195] The motion prediction in pixel units has a great effect of improving the encoding efficiency when used in combination with the motion prediction in CU units. The motion prediction in sub-CU units has a larger processing amount than the motion prediction in CU units. According to the inter prediction processing of the encoding apparatus 100 in the first aspect, by making it possible to perform the motion prediction in pixel units only for the motion prediction in CU units, there is a possibility of reducing the processing amount in the merge mode while maintaining the encoding efficiency.

[0196] Hereinafter, with reference to FIG. 11, the operation example of the encoding apparatus 100 will be described more specifically.

[0197] When motion prediction is not performed using the merge mode (No in S100), the encoding apparatus 100 performs motion prediction using a predetermined mode different from the merge mode (S105). The mode different from the merge mode may be, for example, a normal inter mode for deriving the difference between the candidate MV and the MV.

[0198] On the one hand, when motion prediction is performed using the merge mode (Yes in S100), when generating candidate MVs in units of sub-CUs (Yes in S101), the encoding device 100 performs motion prediction in units of sub-CUs based on the candidate MVs for each sub-CU (S102).

[0199] Also, when motion prediction is performed using the merge mode (Yes in S100), when the encoding device 100 does not generate candidate MVs in units of sub-CUs (No in S101), it performs motion prediction in units of CUs based on the candidate MVs for each CU (S103). Next, the encoding device 100 determines whether to perform motion compensation in units of CUs (not shown). When the encoding device 100 determines to perform motion compensation in units of pixels (not shown), it performs motion compensation in units of pixels (S104). On the other hand, when the encoding device 100 determines not to perform motion compensation in units of pixels (not shown), it does not perform motion compensation in units of pixels (not shown). Note that the encoding device 100 may encode determination result information regarding whether motion compensation in units of pixels is performed.

[0200] FIG. 12 is a flowchart showing another example of operations performed by the encoding device 100 and the decoding device 200 in the first aspect. In FIG. 11, a method of switching whether to perform motion prediction in units of pixels in the merge mode has been described, but it is not limited to the merge mode. For example, in an operation flow as shown in FIG. 12, the encoding device 100 may switch whether to execute motion prediction processing in units of pixels. Also, in FRUC, when only performing motion search in units of CUs, it may be possible to perform motion prediction in units of pixels, and when performing motion search up to units of sub-CUs, it may also be possible to perform motion prediction in units of pixels. Further, after the motion prediction processing in units of sub-CUs, motion prediction processing with lower processing than the motion prediction processing in units of pixels performed at the subsequent stage of the motion prediction in units of CUs may be performed to correct the motion in the predicted image.

[0201] As shown in FIG. 12, the encoding device 100 determines whether to generate candidate MVs in units of sub-CUs (S101). When it is determined that the encoding device 100 generates candidate MVs in units of sub-CUs (Yes in S101), motion prediction is performed in units of sub-CUs based on the candidate MVs for each sub-CU (S102). Examples of methods for performing motion prediction in units of sub-CUs include, for example, a method based on the merge mode and a method based on the affine mode. Further, the merge mode includes the ATMVP mode and the STMVP mode. Details of the ATMVP mode and the STMVP mode will be described later.

[0202] On the other hand, when it is determined that the encoding device 100 does not generate candidate MVs in units of sub-CUs (No in S101), motion prediction is performed in units of CUs based on the candidate MVs for each CU (S103). Examples of methods for performing motion prediction in units of CUs include, for example, a method based on the normal inter mode, a method based on the merge mode, and a method based on the affine mode.

[0203] Next, the encoding device 100 determines whether to perform motion compensation in units of CUs (not shown). When it is determined that the encoding device 100 performs motion compensation in units of pixels (not shown), motion compensation in units of pixels is performed (S104). On the other hand, when it is determined that the encoding device 100 does not perform motion compensation in units of pixels (not shown), motion compensation in units of pixels is not performed (not shown). Note that the encoding device 100 may encode determination result information regarding whether motion compensation in units of pixels is performed.

[0204] In the above, an operation example of the encoding method has been shown. During decoding as well, it can operate in the same manner. When performing motion prediction in units of sub-CUs based on candidate MVs for each sub-CU, motion compensation at the pixel level is not performed. In other words, when generating candidate MVs in units of CUs (i.e., in the case of the first mode), motion compensation (so-called correction processing) is permitted, and when generating candidate MVs in units of sub-CUs (i.e., in the case of the second mode), motion compensation (so-called correction processing) is prohibited. The decoding apparatus 200 determines whether to perform motion compensation at the pixel level using the determination result information determined by the encoding apparatus 100 when generating candidate MVs in units of CUs. Note that the decoding apparatus 200 may decode the determination result information regarding whether motion compensation at the pixel level is to be performed.

[0205] Subsequently, as an example of a mode for determining an MV in units of sub-CUs, a method for determining an MV (sub-CU MV: motion vector in units of sub-blocks) in units of sub-CUs in the ATMVP mode and the STMVP mode will be described. As described above, the ATMVP mode and the STMVP mode are included in the merge mode, which is a mode that uses a predicted motion vector as a motion vector. In the merge mode, one candidate MV is selected from among the MV candidate lists generated by referring to processed blocks to determine the MV of the block to be encoded. There are the ATMVP mode and the STMVP mode as the modes for registering in this MV candidate list.

[0206] FIG. 13 is a diagram showing an example of a method for determining a motion vector in units of sub - blocks in the ATMVP mode. First, a temporal MV for the CU to be coded (the target CU in FIG. 13) is selected from the MVs of the CUs adjacent to the CU to be coded. The temporal MV can be selected from the blocks that are merge candidates. For example, the MVs of the blocks of available merge candidates are searched in order from the merge candidate with the smallest index number, and the MV of the block of the available merge candidate is set as the temporal MV of the CU to be coded. Next, the position of the reference CU in the reference picture is determined according to the temporal MV, and the MV in units of sub - CUs in the reference CU is obtained. Then, the MV in units of sub - CUs in the obtained reference CU is used as the MV in units of sub - CUs corresponding to the sub - CU in the target CU (hereinafter also referred to as the MV in units of sub - CUs of the target CU). When the sub - CUs in the reference CU have multiple MVs (L0, L1), if the reference picture can be obtained, the multiple MVs (L0, L1) are used as the MV in units of sub - CUs of the target CU. Here, the processing in encoding has been described, but the same processing is also performed in decoding.

[0207] FIG. 14 is a diagram showing an example of a method for determining a motion vector in units of sub - blocks in the STMVP mode. In the STMVP mode, the motion vector in units of sub - blocks is determined by averaging, or weighted addition, etc., of the MV of the N×N block adjacent spatially and the MV obtained from a reference picture different temporally, in units of sub - CUs. More specifically, in the STMVP mode, first, a temporal MV reference block at the same position as the block to be coded in the encoded reference picture is specified. Next, for each sub - block in the block to be coded, the MV of the block adjacent spatially above, the MV of the block adjacent spatially to the left, and the MV used when the temporal MV reference block was coded are specified. Then, the MV of each sub - block is obtained by calculating the average of the values obtained by scaling these MVs according to the time interval.

[0208] In the example of FIG. 14, the sub-CU of A determines a spatial MV based on the MVs of spatially adjacent blocks (c or d) and spatially left-adjacent blocks (b or a), and determines a temporal MV based on the MV of an N×N block that is co-located with the sub-CU of D in the reference picture. Then, the spatial MV and the temporal MV can be averaged to obtain the MV of the sub-CU of A. Here, for sub-CUs such as B, C, and D, the spatial MV can be determined using the MV of the encoded or decoded sub-CU. For example, for the sub-CU of B, the sub-CU of A can be used as the spatially left-adjacent block. Here, the processing in encoding has been described, but the same processing applies to decoding.

[0209] [Implementation Example] FIG. 15 is a block diagram showing an implementation example of the encoding apparatus 100. The encoding apparatus 100 includes a circuit 160 and a memory 162. For example, a plurality of components of the encoding apparatus 100 shown in FIG. 1 are implemented by the circuit 160 and the memory 162 shown in FIG. 15.

[0210] The circuit 160 is an electronic circuit accessible to the memory 162 and performs information processing. For example, the circuit 160 is a dedicated or general-purpose electronic circuit that encodes moving images using the memory 162. The circuit 160 may be a processor such as a CPU. Also, the circuit 160 may be an aggregate of a plurality of electronic circuits.

[0211] Further, for example, the circuit 160 may perform the roles of a plurality of components of the encoding apparatus 100 shown in FIG. 1, excluding the components for storing information. That is, the circuit 160 may perform the operations described above as the operations of these components.

[0212] The memory 162 is a dedicated or general-purpose memory in which information for the circuit 160 to encode moving images is stored. The memory 162 may be an electronic circuit, may be connected to the circuit 160, or may be included in the circuit 160.

[0213] Further, the memory 162 may be an aggregate of a plurality of electronic circuits or may be composed of a plurality of sub - memories. Also, the memory 162 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or a recording medium, etc. Further, the memory 162 may be a non - volatile memory or a volatile memory.

[0214] For example, the memory 162 may play the role of a component for storing information among the plurality of components of the encoding device 100 shown in FIG. 1. Specifically, the memory 162 may play the roles of the block memory 118 and the frame memory 122 shown in FIG. 1.

[0215] Also, the memory 162 may store the moving image to be encoded or may store the bit string corresponding to the encoded moving image. Also, the memory 162 may store a program for the circuit 160 to encode the moving image.

[0216] Note that in the encoding device 100, not all of the plurality of components shown in FIG. 1 need to be implemented, and not all of the plurality of processes described above need to be performed. A part of the plurality of components shown in FIG. 1 may be included in other devices, and a part of the plurality of processes described above may be executed by other devices. And in the encoding device 100, by implementing a part of the plurality of components shown in FIG. 1 and performing a part of the plurality of processes described above, it is possible to perform more refined prediction processing while suppressing an increase in the processing amount.

[0217] FIG. 16 is a flowchart showing an operation example of the encoding device 100 shown in FIG. 15. For example, when encoding a moving image, the encoding device 100 shown in FIG. 15 performs the operations shown in FIG. 16. Specifically, the circuit 160 performs the following operations using the memory 162.

[0218] First, circuit 160 determines which mode to use for prediction processing among a plurality of modes including a first mode that performs prediction processing based on motion vectors in block units in a moving image and a second mode that performs prediction processing based on motion vectors in sub-block units obtained by dividing a block (S201).

[0219] Next, when performing prediction processing in the first mode, circuit 160 determines whether to perform correction processing on the predicted image using the spatial gradient of pixel values in the predicted image obtained by performing prediction processing, and performs correction processing when it is determined to perform correction processing (S202). Then, when performing prediction processing in the second mode, circuit 160 does not perform correction processing (S203).

[0220] Thereby, since encoding device 100 uses motion compensation in minute units (e.g., pixel units) in combination with motion prediction in block units, the encoding efficiency is improved. Also, since motion prediction in sub-block units has a larger processing amount than motion prediction in block units, the encoding device does not perform motion compensation in minute units when performing motion prediction in sub-block units. Therefore, the encoding device can reduce the processing amount while maintaining the encoding efficiency by executing motion prediction in minute units only for motion prediction in block units. Thus, the encoding device can perform more refined prediction processing while suppressing an increase in the processing amount.

[0221] For example, the first mode and the second mode are included in a merge mode which is a mode that uses a predicted motion vector as a motion vector.

[0222] Thereby, encoding device 100 can speed up the process for deriving a predicted sample set in the merge mode.

[0223] Also, for example, when the circuit 160 performs prediction processing in the first mode, it encodes determination result information indicating a determination result of whether to perform correction processing, and when performing prediction processing in the second mode, it does not encode the determination result information. Thereby, the encoding device 100 can reduce the amount of encoding.

[0224] Also, for example, the correction processing may be BIO (BI - directional Optical flow) processing. Thereby, the encoding device 100 can correct the prediction image using correction values of minute units in the prediction image generated by deriving motion vectors in block units.

[0225] Also, for example, the second mode may be an ATMVP (Advanced Temporal Motion Vector Prediction) mode. Thereby, since the encoding device 100 does not need to perform motion correction processing in minute units in the ATMVP mode, the processing amount is reduced.

[0226] Also, for example, the second mode may be an STMVP (Spatial - Temporal Motion Vector Prediction) mode. Thereby, since the encoding device 100 does not need to perform motion correction processing in minute units in the STMVP mode, the processing amount is reduced.

[0227] Also, for example, the second mode may be an affine (affine compensation prediction) mode. Thereby, since the encoding device 100 does not need to perform motion correction processing in minute units in the affine mode, the processing amount is reduced.

[0228] FIG. 17 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a circuit 260 and a memory 262. For example, a plurality of components of the decoding device 200 shown in FIG. 10 are implemented by the circuit 260 and the memory 262 shown in FIG. 17.

[0229] Circuit 260 is an electronic circuit that can access memory 262 and performs information processing. For example, circuit 260 is a dedicated or general-purpose electronic circuit that decodes moving images using memory 262. Circuit 260 may be a processor such as a CPU. Also, circuit 260 may be an aggregate of multiple electronic circuits.

[0230] Also, for example, circuit 260 may perform the roles of multiple components of the decoding device 200 shown in FIG. 10, excluding the components for storing information. That is, circuit 260 may perform the operations described above as the operations of these components.

[0231] Memory 262 is a dedicated or general-purpose memory in which information for circuit 260 to decode moving images is stored. Memory 262 may be an electronic circuit, may be connected to circuit 260, or may be included in circuit 260.

[0232] Also, memory 262 may be an aggregate of multiple electronic circuits or may be composed of multiple sub-memories. Also, memory 262 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or recording medium, etc. Also, memory 262 may be a non-volatile memory or a volatile memory.

[0233] For example, memory 262 may perform the role of a component for storing information among the multiple components of the decoding device 200 shown in FIG. 10. Specifically, memory 262 may perform the roles of the block memory 210 and the frame memory 214 shown in FIG. 10.

[0234] Also, a bit string corresponding to the encoded moving image may be stored in memory 262, or the decoded moving image may be stored. Also, a program for circuit 260 to decode moving images may be stored in memory 262.

[0235] Note that in the decoding device 200, not all of the plurality of components shown in FIG. 10 need to be implemented, and not all of the plurality of processes described above need to be performed. Some of the plurality of components shown in FIG. 10 may be included in other devices, and some of the plurality of processes described above may be executed by other devices. In the decoding device 200, by implementing some of the plurality of components shown in FIG. 10 and performing some of the plurality of processes described above, it is possible to perform more refined prediction processing while suppressing an increase in the processing amount.

[0236] FIG. 18 is a flowchart showing an operation example of the decoding device 200 shown in FIG. 17. For example, when decoding a moving image, the decoding device 200 shown in FIG. 17 performs the operations shown in FIG. 18. Specifically, the circuit 260 performs the following operations using the memory 262.

[0237] First, the circuit 260 determines which mode to use for the prediction processing among a plurality of modes including a first mode for performing the prediction processing based on the motion vector in block units in the moving image and a second mode for performing the prediction processing based on the motion vector in sub-block units obtained by dividing the block (S301).

[0238] Next, when the circuit 260 performs the prediction processing in the first mode, it determines whether to perform correction processing on the prediction image using the spatial gradient of the pixel values in the prediction image obtained by performing the prediction processing, and performs the correction processing when it is determined to perform the correction processing (S302). When the circuit 260 performs the prediction processing in the second mode, it does not perform the correction processing (S303).

[0239] As a result, the decoding device 200 uses motion compensation in minute units (e.g., pixel units) in combination with motion prediction in block units, thereby improving the encoding efficiency. Also, since the motion prediction in sub-block units has a larger processing volume than the motion prediction in block units, the decoding device does not perform motion compensation in minute units when performing motion prediction in sub-block units. Therefore, the decoding device can reduce the processing volume while maintaining the encoding efficiency by performing motion prediction in minute units only for the motion prediction in block units. Accordingly, the decoding device can perform more refined prediction processing while suppressing an increase in the processing volume.

[0240] For example, the circuit 260 includes the first mode and the second mode in the merge mode that uses the predicted motion vector as the motion vector.

[0241] As a result, the decoding device 200 can speed up the process for deriving the predicted sample set in the merge mode.

[0242] Also, for example, when performing prediction processing in the first mode, the decoding result information indicating the determination result of whether to perform the prediction processing is decoded, and when performing prediction processing in the second mode, the decoding result information is not decoded. Thereby, the decoding device 200 can improve the processing efficiency.

[0243] Also, for example, the correction process may be a BIO process. As a result, the decoding device 200 can correct the predicted image using the correction value in minute units in the predicted image generated by deriving the motion vector in block units.

[0244] Also, for example, the second mode may be the ATMVP mode. As a result, since the decoding device 200 does not need to perform motion compensation processing in minute units in the ATMVP mode, the processing volume is reduced.

[0245] Further, for example, the second mode may be the STMVP mode. Thereby, since the decoding device 200 does not need to perform motion correction processing in minute units in the STMVP mode, the processing amount is reduced.

[0246] Further, for example, the second mode may be the affine mode. Thereby, since the decoding device 200 does not need to perform motion correction in minute units in the affine mode, the processing amount is reduced.

[0247] Also, the encoding device 100 and the decoding device 200 in the present embodiment may be used as an image encoding device and an image decoding device, respectively, or may be used as a moving image encoding device and a moving image decoding device, respectively.

[0248] Alternatively, each of the encoding device 100 and the decoding device 200 may be used as a prediction device or an inter-prediction device. That is, the encoding device 100 and the decoding device 200 may each correspond only to the inter-prediction unit 126 and the inter-prediction unit 218. And other components such as the entropy encoding unit 110 or the entropy decoding unit 202 may be included in other devices.

[0249] Also, at least a part of the present embodiment may be used as an encoding method, a decoding method, a prediction method, or other methods.

[0250] Also, in the present embodiment, each component may be configured by dedicated hardware or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0251] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuitry and a storage that is electrically connected to the processing circuitry and accessible from the processing circuitry. For example, the processing circuitry corresponds to circuit 160 or 260, and the storage corresponds to memory 162 or 262.

[0252] The processing circuitry includes at least one of dedicated hardware and a program execution unit, and executes processing using the storage. Further, when the processing circuitry includes a program execution unit, the storage stores a software program executed by the program execution unit.

[0253] Here, the software that realizes the encoding device 100 or the decoding device 200 of the present embodiment is the following program.

[0254] That is, this program causes a computer to execute an encoding method for performing prediction processing to encode a moving image, the encoding method including a first mode of performing the prediction processing based on a motion vector in block units in the moving image, and a second mode of performing the prediction processing based on a motion vector in sub-block units obtained by dividing the block. Among a plurality of modes, it is determined which mode to use for performing the prediction processing. When performing the prediction processing in the first mode, it is determined whether to perform correction processing on the predicted image using the spatial gradient of pixel values in the predicted image obtained by performing the prediction processing. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction processing in the second mode, the computer may execute an encoding method in which the correction processing is not performed.

[0255] Alternatively, this program is a decoding method for causing a computer to perform prediction processing to decode a moving image, including a first mode of performing the prediction processing based on a motion vector in block units in the moving image, and a second mode of performing the prediction processing based on a motion vector in sub-block units obtained by dividing the block. It determines which mode among a plurality of modes to use for the prediction processing. When performing the prediction processing in the first mode, it determines whether to perform correction processing on the predicted image using the spatial gradient of pixel values in the predicted image obtained by performing the prediction processing. When it is determined to perform the correction processing, the correction processing is performed. When performing the prediction processing in the second mode, it may execute a decoding method that does not perform the correction processing.

[0256] Also, as described above, each component may be a circuit. These circuits may constitute one circuit as a whole, or may be separate circuits respectively. Also, each component may be realized by a general-purpose processor or a dedicated processor.

[0257] Also, another component may execute the processing executed by a specific component. Also, the order of executing the processing may be changed, or a plurality of processes may be executed in parallel. Also, the encoding / decoding device may include an encoding device 100 and a decoding device 200.

[0258] Also, ordinal numbers such as the first and second used in the description may be appropriately changed. Also, ordinal numbers may be newly given to or removed from components and the like.

[0259] As described above, the aspects of the encoding device 100 and the decoding device 200 have been described based on the embodiments. However, the aspects of the encoding device 100 and the decoding device 200 are not limited to this embodiment. As long as the gist of the present disclosure is not deviated from, various modifications conceived by those skilled in the art applied to this embodiment, or forms constructed by combining components in different embodiments may also be included within the scope of the aspects of the encoding device 100 and the decoding device 200.

[0260] This aspect may be implemented in combination with at least a part of other aspects in the present disclosure. Further, a part of the processing described in the flowchart of this aspect, a part of the configuration of the apparatus, a part of the syntax, etc. may be implemented in combination with other aspects.

[0261] (Embodiment 2) In each of the above embodiments, each of the functional blocks can usually be realized by an MPU, a memory, etc. Further, the processing by each of the functional blocks is usually realized by a program execution unit such as a processor reading and executing software (program) recorded on a recording medium such as a ROM. The software may be distributed by download or the like, or may be recorded on a recording medium such as a semiconductor memory and distributed. Of course, it is also possible to realize each functional block by hardware (a dedicated circuit).

[0262] Also, the processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Further, the processor that executes the above program may be singular or plural. That is, centralized processing may be performed, or distributed processing may be performed.

[0263] Aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these are also included within the scope of the aspects of the present disclosure.

[0264] Furthermore, here, an application example of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments and a system using the same will be described. The system is characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, and an image encoding / decoding device having both. Other configurations in the system can be appropriately changed as the case may be.

[0265] [Usage Example] FIG. 19 is a diagram showing the overall configuration of a content supply system ex100 that realizes a content distribution service. The service area for the communication service is divided into a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed radio stations, are installed in each cell.

[0266] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above elements. The devices may be directly or indirectly connected to each other via a telephone network or short-range wireless communication without passing through the base stations ex106 to ex110, which are fixed radio stations. Also, the streaming server ex103 is connected to devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 via the Internet ex101 or the like. Further, the streaming server ex103 is connected to terminals within a hotspot in an airplane ex117 via a satellite ex116.

[0267] Note that a wireless access point or a hotspot or the like may be used instead of the base stations ex106 to ex110. Also, the streaming server ex103 may be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without passing through the satellite ex116.

[0268] The camera ex113 is a device capable of still image shooting and video shooting such as a digital camera. Also, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handyphone System) etc. that supports the mobile communication system standards generally called 2G, 3G, 3.9G, 4G, and in the future 5G.

[0269] The home appliance ex118 is a device included in a refrigerator or a household fuel cell cogeneration system etc.

[0270] In the content supply system ex100, a terminal having a shooting function is connected to the streaming server ex103 through the base station ex106 etc., enabling live distribution etc. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal etc. in the airplane ex117) performs the encoding process described in each of the above embodiments on the still image or video content shot by the user using the terminal, multiplexes the video data obtained by encoding and the audio data obtained by encoding the sound corresponding to the video, and transmits the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present disclosure.

[0271] On the other hand, the streaming server ex103 stream-distributes the content data transmitted to the requested client. The client is the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, or the terminal etc. in the airplane ex117 that can decode the encoded data described above. Each device that has received the distributed data decodes and plays back the received data. That is, each device functions as an image decoding device according to an aspect of the present disclosure.

[0272] [Distributed Processing] In addition, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, record, and deliver data. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world and the edge servers. In a CDN, a physically closer edge server is dynamically assigned according to the client. Then, by caching and delivering the content to the edge server, the delay can be reduced. Also, when some error occurs or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the delivery entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the delivery, so high-speed and stable delivery can be realized.

[0273] In addition to just the distributed processing of the delivery itself, the encoding process of the captured data may be performed on each terminal, on the server side, or shared between them. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image in units of frames or scenes, or the amount of code is detected. Also, in the second loop, a process to improve the encoding efficiency while maintaining the image quality is performed. For example, by having the terminal perform the first encoding process and the server side that receives the content perform the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode in almost real time, since the data encoded by the terminal in the first time can also be received and played back by other terminals, more flexible real-time delivery becomes possible.

[0274] As another example, cameras such as ex113 perform feature quantity extraction from an image, compress data related to the feature quantity as metadata, and transmit it to a server. The server performs compression according to the meaning of the image, for example, determines the importance of an object from the feature quantity and switches the quantization accuracy. The feature quantity data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.

[0275] As yet another example, in a stadium, a shopping mall, or a factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using the plurality of terminals that have performed shooting and, if necessary, other terminals and a server that have not performed shooting, encoding processes are respectively assigned and distributed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing a picture. Thereby, delay can be reduced and more real-time performance can be realized.

[0276] Also, since the plurality of video data is of substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can refer to each other. Alternatively, the encoded data from each terminal may be received by the server, and the reference relationship may be changed among the plurality of data, or the picture itself may be corrected or replaced and re-encoded. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.

[0277] Also, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based method, or convert H.264 to H.265.

[0278] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the entity performing the process. However, part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.

[0279] [3D, Multi-angle] In recent years, it has also become increasingly common to integrate and utilize different scenes photographed by terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or images or videos of the same scene photographed from different angles. The videos photographed by each terminal are integrated based on the relative positional relationship between the terminals obtained separately, or the regions where the feature points included in the videos match.

[0280] The server may not only encode a two-dimensional moving image, but also automatically or at a time specified by the user, encode a still image based on scene analysis of the moving image, etc., and transmit it to the receiving terminal. When the server can further obtain the relative positional relationship between the photographing terminals, based on not only two-dimensional moving images but also videos of the same scene photographed from different angles, the server can generate the three-dimensional shape of the scene. Note that the server may separately encode three-dimensional data generated by a point cloud or the like, or select or reconstruct the video to be transmitted to the receiving terminal from the videos photographed by a plurality of terminals based on the results of recognizing or tracking a person or an object using the three-dimensional data.

[0281] In this way, the user can arbitrarily select each video corresponding to each photographing terminal to enjoy the scene, or can also enjoy the content obtained by cutting out a video from an arbitrary viewpoint from the three-dimensional data reconstructed using a plurality of images or videos. Furthermore, similar to the video, sound is also collected from a plurality of different angles, and the server may multiplex and transmit the sound from a specific angle or space together with the video according to the video.

[0282] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by means of Multi-View Coding (MVC) or the like, or may perform encoding as separate streams without reference to each other. At the time of decoding the separate streams, they may be played back in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0283] In the case of AR images, the server superimposes virtual object information in the virtual space on the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold the virtual object information and the three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting the virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server, and the server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like and encode it in a state where the portion is transparent. Alternatively, the server may generate data in which the RGB value of a predetermined value is set as the background like chroma key and the portion other than the object is the background color.

[0284] The decoding process of the data delivered in the same way may be performed on each client terminal, on the server side, or they may be shared between each other. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request on another terminal, perform the decoding process, and the decoded signal may be transmitted to a device having a display. By dispersing the process regardless of the performance of the communicable terminals themselves and selecting appropriate content, it is possible to reproduce data with good image quality. As another example, while receiving large-size image data on a TV or the like, only a part of the area such as tiles into which the picture is divided may be decoded and displayed on the viewer's personal terminal. Thereby, while sharing the overall image, it is possible to check the area of one's own field of responsibility or the area to be confirmed in more detail at hand.

[0285] In the future, regardless of indoors or outdoors, in a situation where multiple short-range, medium-range, or long-range wireless communications can be used, by using a delivery system standard such as MPEG-DASH, it is expected to receive content seamlessly while switching appropriate data for the ongoing communication. As a result, the user can switch in real time while freely selecting not only their own terminal but also a decoding device or display device such as a display installed indoors and outdoors. Also, based on the user's location information or the like, it is possible to perform decoding while switching the terminal to be decoded and the terminal to be displayed. Thereby, it becomes possible to move while displaying map information on a part of the wall surface or the ground of the adjacent building in which the displayable device is embedded during the movement to the destination. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or being copied to an edge server in a content delivery service, it is also possible to switch the bit rate of the received data.

[0286] [Scalable Encoding] Regarding content switching, an explanation will be given using a scalable stream compressed and encoded by applying the moving image encoding method shown in each of the above embodiments and shown in FIG. 20. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of a temporally / spatially scalable stream realized by performing encoding by dividing into layers as shown in the figure, it may be configured to switch content. That is, by determining up to which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth on the decoding side, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when you want to watch the continuation of a video that you were watching on your smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.

[0287] Furthermore, as described above, in addition to the configuration that realizes scalability in which pictures are encoded for each layer and an enhancement layer exists above the base layer, the enhancement layer may include meta information based on statistical information of images, etc., and the decoding side may generate high-quality content by super-resolving the pictures of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an increase in resolution. The meta information includes information for specifying linear or non-linear filter coefficients used in super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used in super-resolution processing.

[0288] Alternatively, the picture may be divided into tiles or the like according to the meaning of objects or the like in the image, and the decoding side may be configured to decode only a part of the area by selecting the tile to be decoded. Further, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 21, the meta information is stored using a data storage structure different from the pixel data such as the SEI message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0289] Further, the meta information may be stored in a unit composed of a plurality of pictures such as a stream, a sequence, or a random access unit. Thereby, the decoding side can obtain the time when a specific person appears in the video, etc., and by combining with the information in picture units, can specify the picture in which the object exists and the position of the object in the picture.

[0290] [Optimization of Web Page] FIG. 22 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 23 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 22 and 23, the web page may include a plurality of link images that are links to image contents, and the appearance thereof may be different depending on the device for browsing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) displays a still image or an I picture that each content has as a link image, displays a video like a gif animation with a plurality of still images or I pictures, etc., or receives only the base layer and decodes and displays the video.

[0291] When a user selects a linked image, the display device decodes the base layer with the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Further, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay (the delay from the start of content decoding to the start of display) between the decoding time and the display time of the leading picture by decoding and displaying only forward-reference pictures (I pictures, P pictures, B pictures with only forward reference). Also, the display device may deliberately ignore the reference relationship of the pictures, perform rough decoding by using all B pictures and P pictures as forward references, and perform normal decoding as the received pictures increase over time.

[0292] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for the autonomous driving or driving support of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.

[0293] In this case, since a vehicle, drone, airplane, etc. including the receiving terminal moves, the receiving terminal can realize seamless reception and decoding by transmitting the position information of the receiving terminal at the time of a reception request while switching between base stations ex106 to ex110. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, or the state of the communication bandwidth.

[0294] In the above manner, in the content supply system ex100, the client can receive, decode, and play back the encoded information transmitted by the user in real time.

[0295] [Distribution of Personal Content] In addition, in the content supply system ex100, not only high-quality and long-duration content by video distributors but also unicast or multicast distribution of low-quality and short-duration content by individuals is possible. Also, it is considered that such individual content will increase in the future. In order to make individual content better, the server may perform encoding processing after performing editing processing. This can be realized, for example, in the following configuration.

[0296] During shooting in real time or by accumulation and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection from the original image or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone. The server encodes the edited data based on the editing results. Also, it is known that the viewing rate decreases if the shooting time is too long. The server may automatically clip not only less important scenes but also scenes with little movement, etc., within a specific time range according to the shooting time, based on the image processing results, so that the content is within a specific time range. Or, the server may generate a digest based on the result of semantic analysis of the scene and encode it.

[0297] Note that personal content may contain elements that, as they are, would infringe copyrights, moral rights of the author, or portrait rights, etc., and there may be inconvenient situations for individuals, such as the sharing scope exceeding the intended scope. Therefore, for example, the server may deliberately change the image to make the face of a person in the peripheral part of the screen or the inside of a house out of focus and then encode it. Also, the server may recognize whether a face of a person different from the pre-registered person is reflected in the image to be encoded, and if so, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user designates a person or background area that the user wants to process the image, and the server can perform processing such as replacing the designated area with another video or blurring the focus. In the case of a person, the video of the face part can be replaced while tracking the person in the moving image.

[0298] Also, since the viewing of personal content with a small data volume has a strong requirement for real-time performance, depending on the bandwidth, the decoding device first receives the base layer with the highest priority and decodes and plays it. During this period, the decoding device receives the enhancement layer, and when the playback is looped or played more than twice, it may play a high-quality video including the enhancement layer. For a stream with scalable encoding like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. Besides scalable encoding, a similar experience can be provided even if a rough stream played for the first time and a second stream encoded with reference to the first video are configured as one stream.

[0299] [Other Usage Examples] Also, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSIex500 may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for video encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by a computer ex111 or the like, and encoding or decoding processes may be performed using such software. Further, when the smartphone ex115 has a camera, video data acquired by the camera may be transmitted. The video data at this time is data encoded by the LSIex500 possessed by the smartphone ex115.

[0300] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software and then acquires and plays the content.

[0301] Also, not limited to the content supply system ex100 via the Internet ex101, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of the above embodiments can be incorporated into a digital broadcast system. Since multiplexed data in which video and audio are multiplexed is carried on a broadcast radio wave using a satellite or the like for transmission and reception, there is a difference in that it is more suitable for multicast compared to the unicast-oriented configuration of the content supply system ex100, but similar applications are possible for encoding and decoding processes.

[0302] [Hardware Configuration] FIG. 24 is a diagram showing a smartphone ex115. FIG. 25 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from a base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying data obtained by decoding video captured by the camera unit ex465 and video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with a SIM ex468 for identifying a user and authenticating access to various data including a network. Note that an external memory may be used instead of the memory unit ex467.

[0303] Also, a main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a bus ex470.

[0304] When the power key is turned on by a user's operation, the power supply circuit unit ex461 supplies power to each unit from a battery pack to activate the smartphone ex115 to an operable state.

[0305] The smartphone ex115 performs processes such as calls and data communications based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spectrally spread by the modulation / demodulation unit ex452, and after being subjected to digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, it is transmitted via the antenna ex450. Also, received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, spectrally despread by the modulation / demodulation unit ex452, converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. In the data communication mode, text, still images, or video data is sent to the main control unit ex460 via the operation input control unit ex462 by operating the operation unit ex466 of the main body unit, and the same transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. Also, the voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while a video or still image is being captured by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450.

[0306] When receiving a video attached to an email or chat, or a video linked to a web page or the like, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and the video or still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. Also, the audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. Since real-time streaming has become widespread, there may be a situation where it is not socially appropriate to play the audio depending on the user's situation. Therefore, as an initial value, it is desirable to have a configuration that plays only the video data without playing the audio signal. The audio may be played synchronously only when the user performs an operation such as clicking on the video data.

[0307] Also, although the smartphone ex115 has been described as an example here, as the terminal, in addition to the transceiver type terminal having both an encoder and a decoder, there are three possible implementation forms: a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. Furthermore, in the digital broadcast system, although it has been described as receiving or transmitting multiplexed data in which audio data and the like are multiplexed in video data, the multiplexed data may include character data related to the video in addition to the audio data, or the video data itself may be received or transmitted instead of the multiplexed data.

[0308] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, many terminals are also equipped with a GPU. Therefore, a configuration in which a large area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so that it can be used in common may be adopted. Thereby, the encoding time can be shortened, real-time performance can be ensured, and low latency can be realized. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures using the GPU instead of the CPU.

Industrial Applicability

[0309] The present disclosure can be applied to, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conferencing system, or an electronic mirror.

Description of Signs

[0310] 100 Encoding device 102 Division unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 160, 260 Circuit 162, 262 Memory 200 Decoding device 202 Entropy decoding unit

Claims

Claim 1 An encoding device that performs prediction processing to encode a moving image, comprising: a circuit; a memory, and the circuit uses the memory to determine which mode to use for the prediction processing among a plurality of modes including a first mode of performing the prediction processing based on motion vectors in block units in the moving image and a second mode of performing the prediction processing based on motion vectors in sub-block units obtained by dividing the block; when performing the prediction processing in the first mode, determine whether to perform correction processing on the prediction image using the spatial gradient of pixel values in the prediction image obtained by performing the prediction processing, and perform the correction processing when it is determined to perform the correction processing; when performing the prediction processing in the second mode, do not perform the correction processing; the second mode is included in a merge mode that is a mode of using the motion vector of an adjacent block adjacent to the block as the motion vector, an encoding device. Claim 2 A decoding device that performs prediction processing to decode a moving image, comprising: a circuit; a memory, and the circuit uses the memory to determine which mode to use for the prediction processing among a plurality of modes including a first mode of performing the prediction processing based on motion vectors in block units in the moving image and a second mode of performing the prediction processing based on motion vectors in sub-block units obtained by dividing the block; when performing the prediction processing in the first mode, determine whether to perform correction processing on the prediction image using the spatial gradient of pixel values in the prediction image obtained by performing the prediction processing, and perform the correction processing when it is determined to perform the correction processing; when performing the prediction processing in the second mode, do not perform the correction processing; the second mode is included in a merge mode that is a mode of using the motion vector of an adjacent block adjacent to the block as the motion vector, a decoding device.

Citation Information

Patent Citations

  • Moving image decoding device, moving image coding device, and prediction image generation device

    US20190045214A1

  • Moving image decoding device, moving image encoding device, and prediction image generation device

    WO2017134957A1

Cited By

  • Encoder and decoder

    JP2025137591A