Encoding device, decoding device, and storage medium

By adopting motion vector processing in affine mode in motion image encoding, adjusting the motion vector to control deviation within a prescribed range, the problem of insufficient processing efficiency and circuit scale in the prior art is solved, and the coding efficiency and processing efficiency are improved.

CN120264010APending Publication Date: 2025-07-04PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510466699.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-08-27
Filing Date
2019-08-09
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing moving image encoding methods have shortcomings in processing efficiency, image quality improvement and circuit scale, and it is necessary to improve encoding efficiency and reduce encoding/decoding processing volume and circuit scale.

Method used

The motion vector processing method in the affine mode is adopted, by derive the reference motion vector and the differential motion vector, determine whether the difference exceeds the threshold, adjust the motion vector to control the deviation within the specified range, and use appropriate motion vectors for encoding or decoding.

Benefits of technology

The processing efficiency of encoding and decoding is improved, the memory access in inter-frame prediction processing is reduced, and the encoding efficiency and processing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264010A_ABST
    Figure CN120264010A_ABST
Patent Text Reader

Abstract

The invention relates to an encoding apparatus, a decoding apparatus, and a storage medium. An encoding device of the present disclosure is an encoding device that encodes a moving image, the encoding device including a circuit and a memory connected to the circuit, in which a prediction mode of a current block to be encoded is an affine mode, and in operation, the circuit: derives a reference motion vector for predicting the current block; deriving a first motion vector different from the reference motion vector; deriving a motion vector difference on the basis of the difference between the reference motion vector and the first motion vector; determining whether the motion vector difference is greater than a threshold; a setting unit that sets a first value for a second motion vector, which is different from the reference motion vector and the first motion vector, when it is determined that the motion vector difference is greater than the threshold value, and sets a second value, which is different from the first value, for the second motion vector, when it is determined that the motion vector difference is not greater than the threshold value; and encoding the current block using the second motion vector.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of a patent application for an invention titled "Encoding Device, Decoding Device, Encoding Method, and Decoding Method" with an application date of August 9, 2019, an application number of 201980055826.9. Technical Field

[0002] The present invention relates to an encoding device for encoding moving images, etc. Background Art

[0003] Conventionally, as a standard for encoding moving images, there is H.265 (Non-Patent Document 1) called HEVC (High Efficiency Video Coding).

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] In such encoding and decoding methods, in order to improve processing efficiency, improve image quality, and reduce circuit scale, etc., it is desirable to propose a new method.

[0009] In the embodiments of the present invention or some of them, the structures or methods disclosed can respectively contribute to at least any one of, for example, improvement of encoding efficiency, reduction of encoding / decoding processing volume, reduction of circuit scale, improvement of encoding / decoding speed, and appropriate selection of elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding.

[0010] In addition, the present invention also includes the disclosure of structures or methods that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in processing volume.

[0011] Means for Solving the Problems

[0012] An encoding device according to an aspect of the present invention is an encoding device that encodes a moving image. The encoding device includes a circuit and a memory connected to the circuit. Among them, the prediction mode of the current block to be encoded is the affine mode, and during operation, the circuit: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; in the case where it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and in the case where it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and encodes the current block using the second motion vector.

[0013] A decoding device according to an aspect of the present invention is a decoding device that decodes a moving image. The decoding device includes a circuit and a memory connected to the circuit. Among them, the prediction mode of the current block to be decoded is the affine mode, and during operation, the circuit: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; in the case where it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and in the case where it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and decodes the current block using the second motion vector.

[0014] A computer-readable non-transitory storage medium according to an aspect of the present invention stores a bitstream. Among them, the prediction mode of the current block to be decoded is the affine mode. The bitstream includes an encoded signal and syntax information. The decoding device performs, based on the encoded signal and the syntax information: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; in the case where it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and in the case where it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and decodes the current block using the second motion vector.

[0015] In addition, these inclusive or specific forms can also be implemented by a system, device, method, integrated circuit, computer program, or non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, device, method, integrated circuit, computer program, and recording medium.

[0016] Based on the description of the specification and the drawings, the further benefits and advantages provided by the disclosed embodiments will become clear. These benefits and advantages are sometimes brought about separately according to various embodiments, the description of the specification, and the features of the drawings. In order to obtain one or more benefits and advantages, it is not necessary to provide all of them.

[0017] Advantages of the Invention

[0018] The present invention can provide an encoding device, a decoding device, an encoding method, or a decoding method that can improve processing efficiency. Brief Description of the Drawings

[0019] Figure 1 It is a block diagram showing the functional structure of the encoding device according to the embodiment.

[0020] Figure 2 It is a flowchart showing an example of the overall encoding process performed by the encoding device.

[0021] Figure 3 It is a diagram showing an example of block segmentation.

[0022] Figure 4A It is a diagram showing an example of the structure of a slice.

[0023] Figure 4B It is a diagram showing an example of the structure of a tile.

[0024] Figure 5A It is a table showing transform basis functions corresponding to respective transform types.

[0025] Figure 5B It is a diagram showing SVT (Spatially Varying Transform).

[0026] Figure 6A It is a diagram showing an example of the shape of a filter used in ALF (adaptive loop filter).

[0027] Figure 6B It is a diagram showing another example of the shape of a filter used in ALF.

[0028] Figure 6C It is a diagram showing another example of the shape of a filter used in ALF.

[0029] Figure 7 It is a block diagram showing an example of the detailed structure of a loop filter unit that functions as a DBF.

[0030] Figure 8 It is a diagram showing an example of deblocking filtering with filtering characteristics symmetric with respect to a block boundary.

[0031] Figure 9 It is a diagram for explaining a block boundary where deblocking filtering processing is performed.

[0032] Figure 10 It is a diagram showing an example of a Bs value.

[0033] Figure 11 It is a diagram showing an example of the processing performed by the prediction processing unit of an encoding device.

[0034] Figure 12 It is a diagram showing another example of the processing performed by the prediction processing unit of an encoding device.

[0035] Figure 13 It is a diagram showing another example of the processing performed by the prediction processing unit of an encoding device.

[0036] Figure 14 It is a diagram showing an example of 67 intra prediction modes in intra prediction.

[0037] Figure 15 It is a flowchart showing the flow of the basic processing of inter prediction.

[0038] Figure 16 It is a flowchart showing an example of motion vector derivation.

[0039] Figure 17 It is a flowchart showing another example of motion vector derivation.

[0040] Figure 18 It is a flowchart showing another example of motion vector derivation.

[0041] Figure 19 It is a flowchart showing an example of inter prediction based on a normal inter mode.

[0042] Figure 20 It is a flowchart showing an example of inter prediction based on a merge mode.

[0043] Figure 21 It is a diagram for explaining an example of motion vector derivation processing based on a merge mode.

[0044] Figure 22 It is a flowchart showing an example of FRUC (frame rate up conversion).

[0045] Figure 23 This is a diagram illustrating an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0046] Figure 24 This is a diagram illustrating an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture.

[0047] Figure 25A This is a diagram illustrating an example of the derivation of a motion vector in sub - block units based on motion vectors of multiple adjacent blocks.

[0048] Figure 25B This is a diagram illustrating an example of the derivation of a motion vector in sub - block units in an affine mode with three control points.

[0049] Figure 26A This is a conceptual diagram illustrating the affine merge mode.

[0050] Figure 26B This is a conceptual diagram illustrating the affine merge mode with two control points.

[0051] Figure 26C This is a conceptual diagram illustrating the affine merge mode with three control points.

[0052] Figure 27 This is a flowchart showing an example of the process of the affine merge mode.

[0053] Figure 28A This is a diagram illustrating the affine inter - frame mode with two control points.

[0054] Figure 28B This is a diagram illustrating the affine inter - frame mode with three control points.

[0055] Figure 29 This is a flowchart showing an example of the process of the affine inter - frame mode.

[0056] Figure 30A This is a diagram illustrating the affine inter - frame mode where the current block has three control points and the adjacent block has two control points.

[0057] Figure 30B This is a diagram illustrating the affine inter - frame mode where the current block has two control points and the adjacent block has three control points.

[0058] Figure 31A This is a diagram showing the relationship between the merge mode and DMVR (dynamic motion vector refreshing).

[0059] Figure 31B It is a conceptual diagram for explaining an example of DMVR processing.

[0060] Figure 32 It is a flowchart showing an example of the generation of a predicted image.

[0061] Figure 33 It is a flowchart showing another example of the generation of a predicted image.

[0062] Figure 34 It is a flowchart showing still another example of the generation of a predicted image.

[0063] Figure 35 It is a flowchart for explaining an example of a predicted image correction process based on OBMC (overlapped block motion compensation) processing.

[0064] Figure 36 It is a conceptual diagram for explaining an example of a predicted image correction process based on OBMC processing.

[0065] Figure 37 It is a diagram for explaining the generation of predicted images of two triangles.

[0066] Figure 38 It is a diagram for explaining a model assuming uniform linear motion.

[0067] Figure 39 It is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC (local illumination compensation) processing.

[0068] Figure 40 It is a block diagram showing an installation example of an encoding device.

[0069] Figure 41 It is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0070] Figure 42 It is a flowchart showing an example of the overall decoding process performed by the decoding device.

[0071] Figure 43 It is a diagram showing an example of the process performed by the prediction processing unit of the decoding device.

[0072] Figure 44 It is a diagram showing another example of the process performed by the prediction processing unit of the decoding device.

[0073] Figure 45 It is a flowchart showing an example of inter-frame prediction based on a normal inter-frame mode in the decoding device.

[0074] Figure 46 It is a block diagram showing an installation example of a decoding device.

[0075] Figure 47 It is a flowchart showing an example of an inter-frame prediction process in the first form.

[0076] Figure 48 It is a diagram showing an example of a processing target block.

[0077] Figure 49 It is a diagram showing an example of a calculated threshold value.

[0078] Figure 50 It is an overall structure diagram of a content supply system that implements a content distribution service.

[0079] Figure 51 It is a diagram showing an example of an encoding structure in scalable coding.

[0080] Figure 52 It is a diagram showing an example of an encoding structure in scalable coding.

[0081] Figure 53 It is a diagram showing an example of a display screen of a web page.

[0082] Figure 54 It is a diagram showing an example of a display screen of a web page.

[0083] Figure 55 It is a diagram showing an example of a smart phone.

[0084] Figure 56 It is a block diagram showing an example of the structure of a smart phone. Detailed implementation mode

[0085] (Insight underlying the present invention)

[0086] For example, an encoding device that encodes a moving image suppresses an increase in the processing amount during encoding of the moving image and the like, and when encoding a moving image that undergoes a more fragmented prediction process, subtracts a predicted image from the images that make up the moving image to derive a prediction error. Then, the encoding device performs frequency conversion and quantization on the prediction error and encodes the result as image data. At this time, when performing motion prediction processing on the motion of encoding target units such as blocks included in a moving image in units of blocks or sub-blocks constituting the blocks, there is a possibility of improving the encoding efficiency by controlling the deviation of the motion vector.

[0087] However, in the encoding of blocks and the like included in a moving image, if the deviation of the motion vector is not controlled, it leads to an increase in the processing amount and a decrease in the encoding efficiency.

[0088] Therefore, an encoding device according to an aspect of the present invention is an encoding device that encodes a moving image, and includes a circuit and a memory connected to the circuit. During operation, the circuit derives a reference motion vector used in prediction of a processing target block, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and encodes the processing target block using the changed first motion vector or the unchanged first motion vector.

[0089] Thereby, the encoding device can adjust the motion vector so that the deviation of a plurality of motion vectors within a processing target block divided into a plurality of sub-blocks converges within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referred to becomes smaller, and thus the amount of data transferred from the memory can be suppressed within a specified range. Therefore, the encoding device can reduce the memory access amount in the inter-frame prediction process, and thus the encoding efficiency is improved.

[0090] For example, it may be that the reference motion vector corresponds to a first pixel set within the processing target block, and the first motion vector corresponds to a second pixel set within the processing target block that is different from the first pixel set.

[0091] Thereby, the encoding device can appropriately control the deviation of a plurality of motion vectors within the processing target block using the reference motion vector determined based on the first pixel set and the first motion vector determined based on the second pixel set.

[0092] For example, it may be that when it is determined that the differential motion vector is greater than the threshold value, the circuit changes the first motion vector using a value obtained by clipping the differential motion vector.

[0093] Thereby, the encoding device can change the first motion vector so that the deviation of the first motion vector converges within a specified range. Therefore, the encoding device can reduce the memory access amount in the inter-frame prediction process, and thus the encoding efficiency is improved.

[0094] For example, it may be that the prediction mode of the processing target block is an affine mode.

[0095] Thereby, the encoding device can predict the processing target block in units of sub-blocks, and thus the prediction accuracy is improved.

[0096] For example, it may also be that the threshold is determined such that the worst-case memory access amount when the processing target block is predicted in the prediction mode is equal to or less than the worst-case memory access amount when the prediction process is performed in a prediction mode other than the prediction mode.

[0097] Accordingly, by limiting the memory access amount for transferring data from the storage area within a specified range, the encoding device can smoothly perform the transfer process, and thus the processing efficiency is improved.

[0098] For example, it may also be that the worst-case memory access amount when the prediction process is performed in a prediction mode other than the prediction mode is the memory access amount when the bidirectional prediction process is performed on the processing target block in units of 8×8 pixels in a prediction mode other than the prediction mode.

[0099] At this time, for example, when the encoding device performs the prediction process without using the affine mode, it generates a prediction image from past pictures in block units. When using a prediction mode other than the affine mode and performing bidirectional prediction in units of 8×8 pixel blocks, the amount of memory accessed is the largest. Therefore, in the affine mode, the threshold of the affine mode is also determined so as to converge within this memory access amount. In this way, by limiting the memory access amount for transferring data from the storage area within a specified range, the encoding device can smoothly perform the transfer process, and thus the processing efficiency is improved.

[0100] For example, it may also be that the threshold is different when the processing target block is unidirectionally predicted and when the processing target block is bidirectionally predicted.

[0101] Accordingly, compared with the case where the processing target block is unidirectionally predicted, the number of accesses in the case where the encoding device bidirectionally predicts the processing target block increases. Therefore, the threshold can be set, for example, to be smaller in the case of bidirectional prediction than in the case of unidirectional prediction. In this way, by setting the threshold to be smaller as the number of accesses increases, the encoding efficiency is improved.

[0102] For example, it may also be that the circuit determines whether the differential motion vector is greater than the threshold. When the absolute value of the horizontal component of the differential motion vector is greater than the first value of the threshold or the absolute value of the vertical component of the differential motion vector is greater than the second value of the threshold, it is determined that the differential motion vector is greater than the threshold. When the absolute value of the horizontal component of the differential motion vector is not greater than the first value of the threshold and the absolute value of the vertical component of the differential motion vector is not greater than the second value of the threshold, it is determined that the differential motion vector is not greater than the threshold.

[0103] Thus, the threshold is represented by a first value as a horizontal component (hereinafter, also referred to as the first threshold) and a second value as a vertical component (hereinafter, also referred to as the second threshold). Therefore, the encoding device can determine the deviation of the first motion vector in two dimensions. Furthermore, the encoding device determines whether the deviation of the first motion vector in the processing target block is within a specified range, and thus can adjust the first motion vector in the processing target block to an appropriate value. Therefore, the encoding efficiency is improved.

[0104] For example, it may also be that when the circuit determines that the differential motion vector is greater than the threshold, the circuit uses the value obtained by limiting the differential motion vector and the reference motion vector to change the first motion vector, and the differential motion vector between the changed first motion vector and the reference motion vector is not greater than the threshold.

[0105] At this time, when the deviation of the first motion vector exceeds the specified range, the encoding device, for example, limits the absolute value of the component exceeding the threshold among the horizontal and vertical components of the differential motion vector to the value of the threshold. And the encoding device may, for example, also use the value obtained by adding the reference motion vector and the value obtained by limiting the differential motion vector as the changed first motion vector. Thus, the deviation of the first motion vector in the processing target block is converged within the specified range. Therefore, the encoding efficiency is improved.

[0106] For example, it may also be that the reference motion vector is the average of the plurality of first motion vectors in the processing target block.

[0107] Thus, the encoding device can derive the difference from the reference for each of the plurality of first motion vectors in the processing target block, with the average of all the first motion vectors in the processing target block as the reference.

[0108] For example, it may also be that the reference motion vector is one of the plurality of first motion vectors in the processing target block.

[0109] Thus, the encoding device can derive the difference from the reference for each of the plurality of first motion vectors in the processing target block, with one first motion vector in the processing target block as the reference.

[0110] For example, it may also be that the threshold corresponds to the size of the processing target block.

[0111] Thus, since the memory access amount varies depending on the size of the processing target block for the encoding device, the threshold may be smaller as the memory access amount increases. For example, the larger the size of the processing target block, the smaller the threshold. Therefore, even when the size of the processing target block becomes larger, the encoding device can converge the memory access amount within the specified range, and thus the encoding efficiency is improved.

[0112] For example, it may also be that the threshold value is predetermined and decoding is not performed on the stream.

[0113] Accordingly, the encoding device can use a predetermined threshold value, so there is no need to encode the threshold value every time prediction processing is performed. Therefore, the encoding efficiency is improved.

[0114] In addition, a decoding device according to one aspect of the present invention is a decoding device that decodes a moving image, includes a circuit and a memory connected to the circuit. The circuit derives a reference motion vector used in the prediction of a processing target block during operation, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on the difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and decodes the processing target block using the changed first motion vector or the unchanged first motion vector.

[0115] Accordingly, the decoding device can adjust the motion vector so that the deviation of a plurality of motion vectors within a processing target block divided into a plurality of sub-blocks converges within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referred to is smaller, so the amount of data transferred from the memory can be suppressed within a specified range. Therefore, since the decoding device can reduce the memory access amount in the inter-frame prediction processing, the processing efficiency is improved.

[0116] For example, it may also be that the reference motion vector corresponds to a first pixel set within the processing target block, and the first motion vector corresponds to a second pixel set within the processing target block that is different from the first pixel set.

[0117] Accordingly, the decoding device can appropriately control the deviation of a plurality of motion vectors within a processing target block using a reference motion vector determined based on the first pixel set and a first motion vector determined based on the second pixel set.

[0118] For example, it may also be that when it is determined that the differential motion vector is greater than the threshold value, the circuit changes the first motion vector using a value obtained by limiting the differential motion vector.

[0119] Accordingly, the decoding device can change the first motion vector so that the deviation of the first motion vector converges within a specified range. Therefore, since the decoding device can reduce the memory access amount in the inter-frame prediction processing, the processing efficiency is improved.

[0120] For example, it may also be that the prediction mode of the processing target block is an affine mode.

[0121] Thus, the decoding device can perform prediction processing on the processing target block in units of sub-blocks, thereby improving the prediction accuracy.

[0122] For example, it may also be that the threshold is determined such that the memory access amount in the worst case when the prediction processing is performed on the processing target block in the prediction mode becomes equal to or less than the memory access amount in the worst case when the prediction processing is performed in a prediction mode other than the prediction mode.

[0123] Thus, by limiting the memory access amount for transferring data from the storage area within a specified range, the decoding device can smoothly perform the transfer processing, thereby improving the processing efficiency.

[0124] For example, it may also be that the memory access amount in the worst case when the prediction processing is performed in a prediction mode other than the prediction mode is the memory access amount when the bidirectional prediction processing is performed on the processing target block in units of 8×8 pixels in a prediction mode other than the prediction mode.

[0125] At this time, for example, when the prediction processing is not performed using the affine mode, the decoding device generates a prediction image from past pictures in units of blocks. When performing bidirectional prediction in units of 8×8 pixel blocks using a prediction mode other than the affine mode, the amount of memory accessed is the largest. Therefore, in the affine mode, the threshold of the affine mode is also determined so as to converge within this memory access amount. Thus, by limiting the memory access amount for transferring data from the storage area within a specified range, the decoding device can smoothly perform the transfer processing, thereby improving the processing efficiency.

[0126] For example, it may also be that the threshold is different when the processing target block is unidirectionally predicted and when the processing target block is bidirectionally predicted.

[0127] Thus, compared with the case where the processing target block is unidirectionally predicted, the number of accesses when the encoding device bidirectionally predicts the processing target block increases. Therefore, the threshold can be set, for example, to be smaller in the case of bidirectional prediction than in the case of unidirectional prediction. In this way, by setting the threshold to be smaller as the number of accesses increases, the processing efficiency is improved.

[0128] For example, it may also be that the circuit determines whether the differential motion vector is greater than the threshold value. When the absolute value of the horizontal component of the differential motion vector is greater than the first value of the threshold value, or the absolute value of the vertical component of the differential motion vector is greater than the second value of the threshold value, it is determined that the differential motion vector is greater than the threshold value. When the absolute value of the horizontal component of the differential motion vector is not greater than the first value of the threshold value and the absolute value of the vertical component of the differential motion vector is not greater than the second value of the threshold value, it is determined that the differential motion vector is not greater than the threshold value.

[0129] Thus, the threshold value is represented by the first value as the horizontal component (hereinafter, also referred to as the first threshold value) and the second value as the vertical component (hereinafter, also referred to as the second threshold value). Therefore, the decoding device can determine the deviation of the first motion vector in two dimensions. Furthermore, the decoding device determines whether the deviation of the first motion vector in the processing target block is within a specified range, so that the first motion vector in the processing target block can be adjusted to an appropriate value. Therefore, the processing efficiency is improved.

[0130] For example, it may also be that when the circuit determines that the differential motion vector is greater than the threshold value, it uses the value after limiting the differential motion vector and the reference motion vector to change the first motion vector, and the differential motion vector between the changed first motion vector and the reference motion vector is not greater than the threshold value.

[0131] At this time, when the deviation of the first motion vector exceeds the specified range, the decoding device, for example, limits the absolute value of the component exceeding the threshold value in the horizontal and vertical components of the differential motion vector to the threshold value. Moreover, the decoding device may, for example, use the value obtained by adding the reference motion vector and the value after limiting the differential motion vector as the changed first motion vector. Thus, the deviation of the first motion vector in the processing target block converges within the specified range. Therefore, the processing efficiency is improved.

[0132] For example, it may also be that the reference motion vector is the average of the multiple first motion vectors of the processing target block.

[0133] Thus, the decoding device can derive the difference from the reference for each of the multiple first motion vectors in the processing target block based on the average of all the first motion vectors in the processing target block.

[0134] For example, it may also be that the reference motion vector is one of the multiple first motion vectors of the processing target block.

[0135] Accordingly, the decoding device can derive a difference from one first motion vector in the processing target block as a reference for each of the multiple first motion vectors in the processing target block.

[0136] For example, it may also be that the threshold value corresponds to the size of the processing target block.

[0137] Accordingly, it may also be that, since the memory access amount varies depending on the size of the processing target block in the decoding device, the threshold value becomes smaller as the memory access amount increases. For example, the threshold value becomes smaller as the size of the processing target block increases. Therefore, even if the size of the processing target block becomes larger, the decoding device can converge the memory access amount within a specified range, thereby improving the processing efficiency.

[0138] For example, it may also be that the threshold value is predetermined and not decoded from the stream.

[0139] Accordingly, since the decoding device can use a predetermined threshold value, it is not necessary to decode the threshold value every time prediction processing is performed. Therefore, the processing efficiency is improved.

[0140] In addition, an encoding method according to an aspect of the present invention is an encoding method for encoding a moving image, which derives a reference motion vector used in prediction of a processing target block, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and encodes the processing target block using the changed first motion vector or the unchanged first motion vector.

[0141] Accordingly, the encoding method can adjust the motion vector so that the deviation of multiple motion vectors in the processing target block divided into multiple sub-blocks converges within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referred to is smaller, so that the amount of data transferred from the memory can be suppressed within a specified range. Therefore, the encoding method can reduce the memory access amount in the inter-frame prediction processing, thereby improving the encoding efficiency.

[0142] In addition, a decoding method according to one aspect of the present invention is a decoding method for decoding a moving image, which derives a reference motion vector used in prediction of a processing target block, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on a difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and decodes the processing target block using the changed first motion vector or the unchanged first motion vector.

[0143] Thereby, the decoding method can adjust the motion vector so that the deviation of a plurality of motion vectors within a processing target block divided into a plurality of sub-blocks converges within a specified range. Therefore, compared with the case where the deviation of the motion vector is larger than the specified range, the number of pixels to be referred to is smaller, so that the amount of data transferred from the memory can be suppressed within a specified range. Therefore, since the decoding method can reduce the memory access amount in the inter-frame prediction process, the processing efficiency is improved.

[0144] Moreover, these inclusive or specific aspects can also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.

[0145] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and are not intended to limit the claims.

[0146] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processing and / or structure described in each aspect of the present invention. The processing and / or structure can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processing and / or structure applied to the embodiments, for example, one of the following can be performed.

[0147] (1) Some of the plurality of constituent elements of the encoding device or decoding device of the embodiment described in each aspect of the present invention can be replaced with other constituent elements described in a certain aspect of the present invention, or they can be combined;

[0148] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can also be made to the functions or processes performed by a part of the multiple components of the encoding device or decoding device. For example, any function or process can be replaced with other functions or processes described in a certain one of the various forms of the present invention, or they can be combined;

[0149] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can also be made to a part of the multiple processes included in the method. For example, any process in the method can be replaced with other processes described in a certain one of the various forms of the present invention, or they can be combined;

[0150] (4) A part of the multiple components constituting the encoding device or decoding device of the embodiment can be combined with the components described in a certain one of the various forms of the present invention, or can be combined with the components having a part of the functions described in a certain one of the various forms of the present invention, or can be combined with the components implementing a part of the processes implemented by the components described in the various forms of the present invention;

[0151] (5) The components having a part of the functions of the encoding device or decoding device of the embodiment, or the components implementing a part of the processes of the encoding device or decoding device of the embodiment, are combined or replaced with the components described in a certain one of the various forms of the present invention, the components having a part of the functions described in a certain one of the various forms of the present invention, or the components implementing a part of the processes described in a certain one of the various forms of the present invention;

[0152] (6) In the method implemented by the encoding device or decoding device of the embodiment, a certain one of the multiple processes included in the method is replaced with the process described in a certain one of the various forms of the present invention or the same certain process, or they are combined;

[0153] (7) A part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various forms of the present invention.

[0154] (8) The implementation manners of the processes and / or structures described in the various forms of the present invention are not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for purposes different from the motion image encoding or motion image decoding disclosed in the embodiment.

[0155] (Embodiment 1)

[0156] [Encoding Device]

[0157] First, the encoding device according to this embodiment will be described. Figure 1 FIG. is a block diagram showing the functional configuration of the encoding device 100 according to this embodiment. The encoding device 100 is a moving image encoding device that encodes moving images in units of blocks.

[0158] As Figure 1 shown, the encoding device 100 is a device that encodes images in units of blocks, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0159] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Further, the encoding device 100 may also be implemented as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0160] Hereinafter, after explaining the overall processing flow of the encoding device 100, each component included in the encoding device 100 will be described.

[0161] [Overall Flow of Encoding Processing]

[0162] Figure 2 FIG. is a flowchart showing an example of the overall encoding processing performed by the encoding device 100.

[0163] First, the segmentation unit 102 of the encoding device 100 divides each picture included in the input image, which is a moving image, into a plurality of blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern (also referred to as a block shape) for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block with the fixed size into a plurality of blocks that constitute the selected segmentation pattern. Then, for each of the plurality of blocks, the encoding device 100 performs the processing of steps Sa_3 to Sa_9 on the block (i.e., the encoding target block).

[0164] That is, the prediction processing unit constituted by all or a part of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 generates a prediction signal (also referred to as a prediction block) of the encoding target block (also referred to as the current block) (step Sa_3).

[0165] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).

[0166] Next, the transform unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing a transform and quantization on the difference block (step Sa_5). In addition, the block constituted by the plurality of quantization coefficients is also referred to as a coefficient block.

[0167] Next, the entropy encoding unit 110 generates an encoded signal by encoding the coefficient block and the prediction parameters related to the generation of the prediction signal (specifically, entropy encoding) (step Sa_6). In addition, the encoded signal is also referred to as an encoded bitstream, a compressed bitstream, or a stream.

[0168] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore a plurality of prediction residuals (i.e., difference blocks) by performing an inverse quantization and an inverse transform on the coefficient block (step Sa_7).

[0169] Next, the addition unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the restored difference block to the prediction block (step Sa_8). Thereby, a reconstructed image is generated.

[0170] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).

[0171] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.

[0172] In addition, in the above example, the encoding device 100 selects one splitting pattern for blocks of a fixed size and encodes each block according to the splitting pattern. However, each block may also be encoded according to each of a plurality of splitting patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of splitting patterns, and for example, may select the encoded signal obtained by encoding according to the splitting pattern with the minimum cost as the finally output encoded signal.

[0173] In addition, the processes of these steps Sa_1 to Sa_10 may be sequentially performed by the encoding device 100, and a plurality of these processes may be performed in parallel, or the order may be exchanged.

[0174] [Splitting Unit]

[0175] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128). Such blocks of a fixed size are sometimes referred to as coding tree units (CTUs). And, the splitting unit 102 splits each of the blocks of a fixed size into blocks of a variable size (e.g., 64×64 or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects a splitting pattern. Such blocks of a variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture may be used as the processing units for CUs, PUs, and TUs.

[0176] Figure 3 is a diagram showing an example of block splitting according to the present embodiment. In Figure 3 the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.

[0177] Here, the block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into four square 64×64 blocks (quadtree block splitting).

[0178] The upper left 64×64 block is further vertically split into two rectangular 32×64 blocks, and the left 32×64 block is further vertically split into two rectangular 16×64 blocks (binary tree block splitting). As a result, the upper left 64×64 block is split into two 16×64 blocks 11, 12 and a 32×64 block 13.

[0179] The 64×64 block in the upper right is horizontally divided into two rectangular 64×32 blocks 14 and 15 (binary tree block division).

[0180] The 64×64 block in the lower left is divided into four square 32×32 blocks (quad-tree block division). The upper left and lower right blocks among the four 32×32 blocks are further divided. The upper left 32×32 block is vertically divided into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally divided into two 16×16 blocks (binary tree block division). The lower right 32×32 block is horizontally divided into two 32×16 blocks (binary tree block division). As a result, the 64×64 block in the lower left is divided into 16×32 blocks 16, two 16×16 blocks 17 and 18, two 32×32 blocks 19 and 20, and two 32×16 blocks 21 and 22.

[0181] The 64×64 block 23 in the lower right is not divided.

[0182] As described above, in Figure 3 , block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0183] In addition, in Figure 3 , one block is divided into four or two blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, one block can also be divided into three blocks (ternary tree division). The division including such ternary tree division is sometimes called MBT (multi type tree) division.

[0184] [Structural slices / tile of the picture]

[0185] In order to decode the picture in parallel, the picture is sometimes composed of slices or tiles. The picture composed of slices or tiles can be composed of the dividing unit 102.

[0186] A slice is the basic encoding unit that constitutes the picture. The picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs (Coding Tree Units).

[0187] Figure 4AIt is a diagram showing an example of the structure of slices. For example, the picture includes 11×8 CTUs and is divided into 4 slices (Slice 1 to Slice 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any one of the slices. The shape of the slice is the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can be at any position among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. In addition, the slice contains header information and encoded data. In the header information, features of the slice such as the CTU address at the start of the slice and the slice type can also be described.

[0188] A tile is a unit of a rectangular area that makes up a picture. Numbers called TileIds can also be assigned to each tile in the raster scan order.

[0189] Figure 4B It is a diagram showing an example of the structure of tiles. For example, the picture includes 11×8 CTUs and is divided into 4 rectangular area tiles (Tile 1 to Tile 4). When using tiles, the processing order of the CTUs is changed compared to the case of not using tiles. When not using tiles, multiple CTUs in the picture are processed in the raster scan order. When using tiles, in each of the multiple tiles, at least 1 CTU is processed in the raster scan order. For example, as Figure 4B shown, the processing order of the multiple CTUs included in Tile 1 is the order from the left end of the first column of Tile 1 to the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 to the right end of the second column of Tile 1.

[0190] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.

[0191] [Subtraction unit]

[0192] The subtraction unit 104 subtracts the prediction signal (the prediction sample input from the prediction control unit 128 as shown below) from the original signal (original sample) in units of blocks input from and divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction unit 104 outputs the calculated prediction error (residual) to the transformation unit 106.

[0193] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, there are also cases where a signal representing an image is called a sample.

[0194] [Transform section]

[0195] The transform section 106 transforms the prediction error in the spatial domain into transform coefficients in the frequency domain and outputs them to the transform coefficient vectorization section 108. Specifically, the transform section 106 performs a preset discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.

[0196] In addition, the transform section 106 can also adaptively select a transform type from multiple transform types and use a transform basis function corresponding to the selected transform type to transform the prediction error into transform coefficients. Such a transform is called EMT (explicit multiple core transform) or AMT (adaptive multiple transform) in some cases.

[0197] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing the transform basis functions corresponding to each transform type. In Figure 5A where N represents the number of input pixels. The selection of the transform type from these multiple transform types can depend on, for example, the type of prediction (intra prediction and inter prediction) or the intra prediction mode.

[0198] The information indicating whether to apply such EMT or AMT (for example, called the EMT flag or AMT flag) and the information indicating the selected transform type are usually signaled at the CU level. In addition, the signaling of these information does not need to be limited to the CU level and can also be other levels (for example, bit sequence level, picture level, slice level, tile level, or CTU level).

[0199] In addition, the transform section 106 can also perform a re-transformation on the transform coefficients (transformation results). Such a re-transformation is called AST (adaptive secondary transform) or NSST (non-separable secondary transform) in some cases. For example, the transform section 106 performs a re-transformation on each sub-block (for example, 4×4 sub-block) included in the block of transform coefficients corresponding to the intra prediction error. The information indicating whether to apply NSST and the information related to the transform matrix used in NSST are usually signaled at the CU level. In addition, the signaling of these information does not need to be limited to the CU level and can also be other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0200] In the transformation unit 106, separable transformation and non-separable transformation can also be applied. Separable transformation refers to a method of performing multiple transformations by separating in each direction according to the number of dimensions of the input. Non-separable transformation refers to a method of treating two or more dimensions as one dimension and performing transformation together when the input is multi-dimensional.

[0201] For example, as an example of non-separable transformation, when the input is a 4×4 block, it can be regarded as a permutation with 16 elements, and a transformation process is performed on this permutation with a 16×16 transformation matrix.

[0202] In addition, in a further example of non-separable transformation, after regarding a 4×4 input block as a permutation with 16 elements, a transformation (Hypercube Givens Transform) of performing Givens rotation on this permutation multiple times can also be performed.

[0203] In the transformation in the transformation unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, as Figure 5B shown, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The type of transformation basis can be set for each region. For example, DST7 and DCT8 are used. In this example, only one of the two regions within the CU is transformed, and the other is not transformed, but both regions can also be transformed. In addition, the splitting method is not limited to bisection, and can be more flexible, such as quartering or encoding the information indicating the splitting separately and performing signaling in the same way as CU splitting. Sometimes, SVT is also called SBT (Sub-block Transform).

[0204] [Quantization Unit]

[0205] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a prescribed scanning order, and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. And the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0206] The specified scanning order is the order for quantization / inverse quantization of transform coefficients. For example, the specified scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).

[0207] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0208] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing the values sampled at a predetermined interval by associating them with a predetermined level, and in this technical field, expressions such as rounding, truncation, and scaling are sometimes used.

[0209] As a method of using the quantization matrix, there are a method of using the quantization matrix directly set on the encoding device side and a method of using the default quantization matrix (default matrix). On the encoding device side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of coding increases due to the coding of the quantization matrix.

[0210] On the other hand, there is also a method of quantizing in such a way that the coefficients of the high-frequency components and the coefficients of the low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to the method of using a quantization matrix (flat matrix) in which all coefficients are the same value.

[0211] The quantization matrix can be specified by, for example, SPS (Sequence Parameter Set) or PPS (Picture Parameter Set). SPS contains parameters used for a sequence, and PPS contains parameters used for a picture. SPS and PPS are sometimes simply referred to as parameter sets.

[0212] [Entropy Encoding Unit]

[0213] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) based on the quantized coefficients input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantized coefficients, performs arithmetic coding on the binary signal, and outputs a compressed bitstream or sequence.

[0214] [Inverse Quantization Unit]

[0215] The inverse quantization unit 112 inverse quantizes the quantization coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantization coefficients of the current block in a prescribed scan order. Further, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0216] [Inverse transform unit]

[0217] The inverse transform unit 114 restores the prediction error (residual) by inverse-transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform of the transform unit 106 on the transform coefficients. Further, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0218] In addition, the restored prediction error usually does not match the prediction error calculated by the subtraction unit 104 because information is lost through quantization. That is, the restored prediction error usually includes quantization error.

[0219] [Addition unit]

[0220] The addition unit 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. Further, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may be referred to as a local decoded block.

[0221] [Block memory]

[0222] The block memory 118 is, for example, a storage unit that stores blocks within the coded object picture (referred to as the current picture) that are referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.

[0223] [Frame memory]

[0224] The frame memory 122 is, for example, a storage unit that stores reference pictures used in inter prediction, and may also be referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0225] [Loop filter unit]

[0226] The loop filter unit 120 applies loop filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0227] In the ALF, a least-squares error filter used to remove coding distortion is adopted. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters is adopted based on the direction and activity of the gradient based on locality.

[0228] Specifically, first, sub-blocks (such as 2×2 sub-blocks) are classified into multiple classes (such as 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (such as 0 to 2 or 0 to 4) and the activity value A of the gradient (such as 0 to 4), the classification value C (such as C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes.

[0229] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (such as horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantifying the addition result.

[0230] Based on the result of such classification, the filter to be used for the sub-block is determined from multiple filters.

[0231] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 6A - 6C It is a diagram showing multiple examples of the shape of the filter used in the ALF. Figure 6A Represents a 5×5 diamond-shaped filter, Figure 6B Represents a 7×7 diamond-shaped filter, Figure 6C Represents a 9×9 diamond-shaped filter. The information representing the shape of the filter is usually signaled at the picture level. In addition, the signaling of the information representing the shape of the filter does not need to be limited to the picture level and can also be other levels (such as sequence level, slice level, tile level, CTU level, or CU level).

[0232] The on / off of the ALF can also be determined, for example, at the picture level or CU level. For example, regarding luminance, it can be determined at the CU level whether to adopt the ALF, and regarding chrominance difference, it can be determined at the picture level whether to adopt the ALF. The information representing the on / off of the ALF is usually signaled at the picture level or CU level. In addition, the signaling of the information representing the on / off of the ALF does not need to be limited to the picture level or CU level and can also be other levels (such as sequence level, slice level, tile level, or CTU level).

[0233] The coefficient sets of multiple selectable filters (such as up to 15 or 25 filters) are usually signaled at the picture level. In addition, the signaling of the coefficient sets does not need to be limited to the picture level and can also be other levels (such as sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0234] [Loop Filter Section > Deblocking Filter]

[0235] In the deblocking filter, the loop filter section 120 reduces the distortion generated at the block boundary by performing a filtering process on the block boundary of the reconstructed image.

[0236] Figure 7 is a block diagram showing an example of the detailed structure of the loop filter section 120 that functions as a deblocking filter.

[0237] The loop filter section 120 includes a boundary determination section 1201, a filtering determination section 1203, a filtering process section 1205, a process determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0238] The boundary determination section 1201 determines whether there are pixels (i.e., target pixels) for which deblocking filtering is to be performed near the block boundary. Then, the boundary determination section 1201 outputs the determination result to the switch 1202 and the process determination section 1208.

[0239] When it is determined by the boundary determination section 1201 that target pixels exist near the block boundary, the switch 1202 outputs the image before the filtering process to the switch 1204. On the contrary, when it is determined by the boundary determination section 1201 that target pixels do not exist near the block boundary, the switch 1202 outputs the image before the filtering process to the switch 1206.

[0240] The filtering determination section 1203 determines whether to perform deblocking filtering on the target pixels based on the pixel values of at least one neighboring pixel located around the target pixels. Then, the filtering determination section 1203 outputs the determination result to the switch 1204 and the process determination section 1208.

[0241] When it is determined by the filtering determination section 1203 that deblocking filtering is to be performed on the target pixels, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the filtering process section 1205. On the contrary, when it is determined by the filtering determination section 1203 that deblocking filtering is not to be performed on the target pixels, the switch 1204 outputs the image before the filtering process obtained via the switch 1202 to the switch 1206.

[0242] When the image before the filtering process is obtained via the switches 1202 and 1204, the filtering process section 1205 performs deblocking filtering on the target pixels with the filtering characteristics determined by the filtering characteristic determination section 1207. Then, the filtering process section 1205 outputs the pixels after the filtering process to the switch 1206.

[0243] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs pixels that have not been deblocking-filtered and pixels that have been deblocking-filtered by the filtering processing unit 1205.

[0244] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processing determination unit 1208 outputs the deblocking-filtered pixels from the switch 1206. In addition, in cases other than the above, the processing determination unit 1208 outputs the pixels that have not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206.

[0245] Figure 8 It is a diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to a block boundary.

[0246] In deblocking filtering processing, for example, using the pixel value and the quantization parameter, either one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Figure 8 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.

[0247] q'0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0248] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0249] q'2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0250] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering processing are filtering coefficients.

[0251] Furthermore, in the deblocking filtering processing, clipping processing may also be performed in such a way that the pixel value after the operation does not exceed the threshold value. In this clipping processing, using the threshold value determined according to the quantization parameter, the pixel value after the operation based on the above equation is clipped to "the pixel value before the operation ± 2×threshold value". Thereby, excessive smoothing can be prevented.

[0252] Figure 9 It is a diagram for explaining the block boundary where deblocking filtering is performed. Figure 10 It is a diagram showing an example of the Bs value.

[0253] The block boundary where deblocking filtering is performed is, for example, Figure 9 the boundary of the PU (Prediction Unit) or TU (Transform Unit) of the 8×8 pixel block shown. The deblocking filtering is performed in units of 4 rows or 4 columns. First, for Figure 9 the blocks P and Q shown, the Bs (Boundary Strength) value is determined as Figure 10 shown.

[0254] Based on Figure 10 the Bs value, it is determined whether to perform deblocking filtering with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering for the chrominance signal is performed. When the Bs value is 1 or more and satisfies a specified condition, deblocking filtering for the luminance signal is performed. In addition, the determination condition of the Bs value is not limited to Figure 10 the conditions shown, and it can also be determined based on other parameters.

[0255] [Prediction processing unit (intra prediction unit / inter prediction unit / prediction control unit)]

[0256] Figure 11 It is a diagram showing an example of the processing performed by the prediction processing unit of the encoding device 100. In addition, the prediction processing unit is composed of all or part of the constituent elements of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0257] The prediction processing unit generates a prediction image of the current block (step Sb_1). This prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, there are, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit uses the reconstructed image that has already been obtained by generating a prediction block, a differential block, a coefficient block, restoring the differential block, and generating a decoded image block, and generates a prediction image of the current block.

[0258] The reconstructed image can be, for example, an image of a reference picture, or an image of an encoded block within the current picture including the current block. The encoded block within the current picture is, for example, an adjacent block of the current block.

[0259] Figure 12 It is a diagram showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0260] The prediction processing unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such a prediction method, the above-mentioned reconstructed image can also be used.

[0261] Next, the prediction processing unit selects any one of the plurality of prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). The selection of this prediction image, that is, the selection of the method or mode for obtaining the final prediction image, can also calculate the cost for each generated prediction image and be based on this cost. In addition, the selection of this prediction image can be performed based on the parameters for the encoding process. The encoding device 100 can signal the information for determining the selected prediction image, method, or mode as an encoding signal (also referred to as an encoding bitstream). This information can be, for example, a flag or the like. Thus, the decoding device can generate a prediction image based on this information in the manner or mode selected in the encoding device 100. In addition, in Figure 12 the example shown, the prediction processing unit selects any one of the prediction images after generating the prediction images by each method. However, before generating these prediction images, the prediction processing unit can select the method or mode based on the parameters for the above-mentioned encoding process, and can generate the prediction image according to this method or mode.

[0262] For example, the first method and the second method are intra-frame prediction and inter-frame prediction, respectively, and the prediction processing unit can select the final prediction image for the current block from the prediction images generated according to these prediction methods.

[0263] Figure 13 FIG. is another example of the processing performed by the prediction processing unit of the encoding device 100.

[0264] First, the prediction processing unit generates a prediction image by intra-frame prediction (step Sd_1a), and generates a prediction image by inter-frame prediction (step Sd_1b). In addition, the prediction image generated by intra-frame prediction is also referred to as an intra-frame prediction image, and the prediction image generated by inter-frame prediction is also referred to as an inter-frame prediction image.

[0265] Next, the prediction processing unit evaluates each of the intra-predicted image and the inter-predicted image (step Sd_2). Cost can also be used in this evaluation. That is, the prediction processing unit calculates the cost C of each of the intra-predicted image and the inter-predicted image. This cost C is calculated by an equation of the R-D optimization model such as C = D + λ × R. In this equation, D is the coding distortion of the predicted image and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. In addition, R is the generated coding amount of the predicted image. Specifically, it is the coding amount required for coding motion information, etc. used to generate the predicted image. In addition, λ is, for example, the undetermined multiplier of Lagrange.

[0266] Then, the prediction processing unit selects, as the final predicted image of the current block, the predicted image that calculates the minimum cost C from the intra-predicted image and the inter-predicted image (step Sd_3). That is, the prediction method or mode used to generate the predicted image of the current block is selected.

[0267] [Intra-Prediction Unit]

[0268] The intra-prediction unit 124 performs intra-prediction (also called intra-frame prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra-prediction unit 124 generates an intra-prediction signal by performing intra-prediction with reference to the samples (such as luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 128.

[0269] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of predefined intra-prediction modes. The plurality of intra-prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0270] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC standard.

[0271] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32-direction prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 14 is a diagram showing all 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra-prediction. The solid arrows represent 33 directions defined by the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the 2 non-directional prediction modes are not Figure 14 shown in the figure).

[0272] In various installation examples, in the intra prediction of a chrominance block, a luminance block may also be referred to. That is, the chrominance component of the current block may also be predicted based on the luminance component of the current block. Such intra prediction is sometimes referred to as CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block in this way (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0273] The intra prediction unit 124 may also correct the intra-predicted pixel value based on the gradient of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intraprediction combination). Information indicating whether PDPC is used (for example, called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and may also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0274] [Inter prediction unit]

[0275] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or the current sub-block (for example, 4×4 block) within the current block. For example, for the current block or the current sub-block, the inter prediction unit 126 performs motion search (motion estimation) within the reference picture to find the reference block or sub-block that most matches the current block or the current sub-block. And the inter prediction unit 126 obtains motion information (for example, a motion vector) that compensates for the motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction unit 126 performs motion compensation (or motion prediction) based on this motion information, thereby generating an inter prediction signal for the current block or sub-block. And the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0276] The motion information used in motion compensation is signaled as an inter prediction signal in various forms. For example, the motion vector may also be signaled. As another example, the difference between the motion vector and the predicted motion vector (motion vector predictor) may also be signaled.

[0277] [Basic process of inter prediction]

[0278] Figure 15It is a flowchart showing the basic process of inter-frame prediction.

[0279] First, the inter-frame prediction unit 126 generates a prediction image (Steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the prediction image as the prediction residual (Step Se_4).

[0280] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by determining the motion vector (MV) of the current block (Steps Se_1 and Se_2) and performing motion compensation (Step Se_3). In addition, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by selecting candidate motion vectors (candidate MVs) (Step Se_1) and deriving the MV (Step Se_2). The selection of candidate MVs is performed, for example, by selecting at least one candidate MV from a candidate MV list. Also, in the derivation of the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from at least one candidate MV and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching for the region of the reference picture indicated by each of the selected at least one candidate MVs. Additionally, the action of searching for the region of the reference picture may also be referred to as motion estimation.

[0281] Furthermore, in the above example, Steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126, but for example, the processing such as Step Se_1 or Step Se_2 may also be performed by other components included in the encoding device 100.

[0282] [Flow of Derivation of Motion Vector]

[0283] Figure 16 It is a flowchart showing an example of the derivation of a motion vector.

[0284] The inter-frame prediction unit 126 derives the MV of the current block in a mode where motion information (such as MV) is encoded. In this case, for example, the motion information is encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the encoded signal (also referred to as the encoded bitstream).

[0285] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode where motion information is not encoded. In this case, the motion information is not included in the encoded signal.

[0286] Here, the modes for MV derivation include the following ordinary inter-frame mode, merge mode, FRUC mode, and affine mode. Among these modes, the modes for encoding motion information include the ordinary inter-frame mode, merge mode, and affine mode (specifically, affine inter-frame mode and affine merge mode). In addition, the motion information can include not only MVs but also the predicted motion vector selection information described later. In addition, the modes that do not encode motion information include the FRUC mode. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0287] Figure 17 It is a flowchart showing another example of motion vector derivation.

[0288] The inter-frame prediction unit 126 derives the MV of the current block in the mode for encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0289] Alternatively, the inter-frame prediction unit 126 derives the MV in the mode that does not encode the differential MV. In this case, the encoded differential MV is not included in the encoded signal.

[0290] Here, as described above, the modes for MV derivation include the following ordinary inter-frame mode, merge mode, FRUC mode, and affine mode. Among these modes, the modes for encoding the differential MV include the ordinary inter-frame mode and affine mode (specifically, affine inter-frame mode). In addition, the modes that do not encode the differential MV include the FRUC mode, merge mode, and affine mode (specifically, affine merge mode). The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0291] [Flow of Motion Vector Derivation]

[0292] Figure 18It is a flowchart showing another example of motion vector derivation. There are multiple modes for MV derivation, i.e., inter-frame prediction modes, which are roughly classified into a mode for encoding differential MVs and a mode for not encoding differential motion vectors. The modes for not encoding differential MVs include the merge mode, the FRUC mode, and the affine mode (specifically, the affine merge mode). The detailed content of these modes will be described later. Briefly speaking, the merge mode is a mode for deriving the MV of the current block by selecting a motion vector from the surrounding encoded blocks, and the FRUC mode is a mode for deriving the MV of the current block by searching between the encoded regions. In addition, the affine mode is a mode for assuming an affine transformation and deriving the motion vectors of the respective sub-blocks constituting the current block as the MV of the current block.

[0293] Specifically, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector based on the merge mode (Sf_2). In addition, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the FRUC mode (Sf_3). In addition, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the affine mode (specifically, the affine merge mode) (Sf_4). In addition, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the mode for encoding differential MVs (e.g., the normal inter-frame mode) (Sf_5).

[0294] [MV Derivation > Normal Inter-Frame Mode]

[0295] The normal inter-frame mode is an inter-frame prediction mode for deriving the MV of the current block by searching for a block similar to the image of the current block in the region of the reference picture represented by the candidate MVs. In addition, in this normal inter-frame mode, the differential MV is encoded.

[0296] Figure 19 It is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.

[0297] First, the inter-frame prediction unit 126 obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks located around the current block in time or space (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0298] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined order of priority (step Sg_2). In addition, this order of priority is predetermined for each of the N candidate MVs.

[0299] Next, the inter-frame prediction unit 126 selects one prediction motion vector candidate from the N prediction motion vector candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction motion vector selection information for identifying the selected prediction motion vector into the stream. In addition, the stream is the above-mentioned encoded signal or encoded bitstream.

[0300] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction motion vector into the stream as the differential MV. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0301] Finally, the inter-frame prediction unit 126 generates the predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). In addition, the predicted image is the above-mentioned inter-frame prediction signal.

[0302] In addition, the information indicating the inter-frame prediction mode (in the above example, the normal inter-frame mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.

[0303] In addition, the candidate MV list may also be used in common with the list used in other modes. In addition, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to this candidate MV list is, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs, etc.

[0304] [MV Derivation> Merge Mode]

[0305] The merge mode is an inter-frame prediction mode in which the MV of the current block is derived by selecting a candidate MV from the candidate MV list.

[0306] Figure 20 It is a flowchart showing an example of inter-frame prediction based on the merge mode.

[0307] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0308] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream.

[0309] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).

[0310] In addition, the information indicating the inter-frame prediction mode (in the above example, the merge mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.

[0311] Figure 21 It is a diagram for explaining an example of the motion vector derivation process of the current picture based on the merge mode.

[0312] First, a predicted MV list registering candidates for predicted MVs is generated. As candidates for predicted MVs, there are: a spatially adjacent predicted MV, which is the MV possessed by a plurality of encoded blocks located in the spatial vicinity of the object block; a temporally adjacent predicted MV, which is the MV possessed by a nearby block obtained by projecting the position of the object block in the encoded reference picture; a combined predicted MV, which is an MV generated by combining the MV values of the spatially adjacent predicted MV and the temporally adjacent predicted MV; and a zero predicted MV, which is an MV with a value of zero, etc.

[0313] Next, by selecting one predicted MV from the plurality of predicted MVs registered in the predicted MV list, the MV for the object block is determined.

[0314] Moreover, in the variable length coding unit, the signal indicating which predicted MV is selected, i.e., merge_idx, is described in the stream and encoded.

[0315] In addition, in Figure 21 the predicted MVs registered in the predicted MV list described are an example, and the number may be different from that in the figure, or the structure may not include some of the predicted MVs in the figure, or the structure may be appended with predicted MVs other than the types of predicted MVs in the figure.

[0316] It is also possible to use the MV of the object block exported in the merge mode and determine the final MV by performing the subsequent DMVR (dynamic motion vector refreshing) process.

[0317] In addition, the candidate for the predicted MV is the above-mentioned candidate MV, and the predicted MV list is the above-mentioned candidate MV list. In addition, the candidate MV list may also be referred to as the candidate list. In addition, merge_idx is MV selection information.

[0318] [MV Export>FRUC Mode]

[0319] Motion information may also be derived on the decoding device side instead of being signaled from the encoding device side. In addition, as described above, the merge mode defined by the H.265 / HEVC standard may also be used. In addition, for example, motion information may be derived by performing a motion search on the decoding device side. In this case, the pixel values of the current block are not used for the motion search on the decoding device side.

[0320] Here, the mode of performing motion estimation on the decoding device side will be described. The mode of performing motion estimation on the decoding device side may be a mode called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0321] In Figure 22This represents an example of FRUC processing. First, with reference to the motion vectors of coded blocks adjacent to the current block in space or time, a plurality of candidate lists each having a predicted motion vector (MV) are generated (i.e., it is a candidate MV list and can also be shared with the merge list) (step Si_1). Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected based on the evaluation value. And, based on the motion vector of the selected candidate, a motion vector for the current block is derived (step Si_4). Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, the motion vector for the current block can also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV can be searched by using pattern matching and evaluation values in the reference picture. In the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and it is used as the final MV of the current block. It is also possible to configure it such that the process of updating to an MV with a better evaluation value is not performed.

[0322] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block by using the derived MV and the coded reference picture (step Si_5).

[0323] The same process can also be performed in the case of processing in sub-block units.

[0324] The evaluation value can also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the motion vector is compared with the reconstructed image of a specified region (for example, as shown below, this region can be a region of another reference picture or a neighboring block of the current picture). Then, the difference in pixel values of the two reconstructed images can also be calculated and used as the evaluation value for the motion vector. Additionally, it can also be that other information is used in addition to the difference value to calculate the evaluation value.

[0325] Next, the pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (for example, the merge list) is selected as the starting point for the search based on pattern matching. As the pattern matching, the first pattern matching or the second pattern matching is used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.

[0326] [MV Derivation>FRUC>Bilateral Matching]

[0327] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block within two different reference pictures. Therefore, in the first pattern matching, as the specified area for calculating the evaluation value for candidates as described above, an area within another reference picture along the motion trajectory of the current block is used.

[0328] Figure 23 This is a diagram for explaining an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along the motion trajectory. As Figure 23 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks within two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two motion vectors (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position within the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position within the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the final MV.

[0329] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0330] [MV Derivation>FRUC>Template Matching]

[0331] In the second pattern matching (template matching), pattern matching is performed between a template within the current picture (a block adjacent to the current block within the current picture, such as an upper and / or left adjacent block) and a block within the reference picture. Therefore, in the second pattern matching, as the specified area for calculating the evaluation value for candidates as described above, a block adjacent to the current block within the current picture is used.

[0332] Figure 24 This is a diagram for explaining an example of the pattern matching (template matching) between a template within the current picture and a block within the reference picture. As Figure 24As shown, in the second pattern matching, the motion vector of the current block is derived by searching for the block in the reference picture (Ref0) that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or one of the left and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. Among the multiple candidate MVs, the candidate MV with the best evaluation value is selected as the best candidate MV.

[0333] Information indicating whether to adopt the FRUC mode (e.g., referred to as the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the pattern matching method that can be adopted (the first pattern matching or the second pattern matching) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0334] [MV Derivation > Affine Mode]

[0335] Next, the affine mode of deriving the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks will be described. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0336] Figure 25A is a diagram for explaining an example of the derivation of the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks. In Figure 25A , the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, according to the following equation (1A), the two motion vectors v0 and v1 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.

[0337] [Equation 1]

[0338]

[0339] Here, x and y represent the horizontal and vertical positions of the sub-block respectively, and w represents a pre-specified weight coefficient.

[0340] Information indicating such an affine mode (e.g., an affine flag) can be signaled as a signal at the CU level. In addition, the signaling of the information indicating the affine mode need not be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0341] In addition, in such an affine mode, several modes in which the derivation methods of the motion vectors of the upper left and upper right control points are different can also be included. For example, in the affine mode, there are two modes: affine inter-frame (also referred to as affine ordinary inter-frame) mode and affine merge mode.

[0342] [MV Derivation > Affine Mode]

[0343] Figure 25B is a diagram for explaining an example of the derivation of the motion vectors of sub-block units in an affine mode with three control points. In Figure 25B , the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the motion vectors of adjacent blocks, and the motion vector v2 of the lower left control point of the current block is derived based on the motion vectors of adjacent blocks. Then, according to the following equation (1B), the three motion vectors v0, v1, and v2 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.

[0344]

Equation 2

[0345]

[0346] Here, x and y represent the horizontal position and vertical position of the center of the sub-block, w represents the width of the current block, and h represents the height of the current block.

[0347] Affine modes with different numbers of control points (e.g., two and three) can also be switched and signaled at the CU level. In addition, the information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0348] In addition, in such an affine mode with three control points, several modes in which the derivation methods of the motion vectors of the upper left, upper right, and lower left control points are different can also be included. For example, in the affine mode, there are two modes: affine inter-frame (also referred to as affine ordinary inter-frame) mode and affine merge mode.

[0349] [MV Derivation > Affine Merge Mode]

[0350] Figure 26A , Figure 26B and Figure 26C are conceptual diagrams for explaining the affine merge mode.

[0351] In the affine merge mode, as Figure 26A shown, for example, based on a plurality of motion vectors corresponding to blocks encoded in an affine mode in the encoded blocks A (left), B (top), C (top - right), D (bottom - left), and E (top - left) adjacent to the current block, a predicted motion vector for each of the control points of the current block is calculated. Specifically, these blocks are examined in the order of the encoded blocks A (left), B (top), C (top - right), D (bottom - left), and E (top - left), and the first valid block encoded in the affine mode is determined. The predicted motion vector of the control point of the current block is calculated based on the plurality of motion vectors corresponding to the determined block.

[0352] For example, as Figure 26B shown, in the case where the block A adjacent to the left side of the current block is encoded in an affine mode with 2 control points, motion vectors v3 and v4 projected onto positions of the upper - left corner and the upper - right corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3 and v4, a predicted motion vector v0 for the upper - left corner control point of the current block and a predicted motion vector v1 for the upper - right corner control point of the current block are calculated.

[0353] For example, as Figure 26C shown, when the block A adjacent to the left side of the current block is encoded in an affine mode with 3 control points, motion vectors v3, v4, and v5 projected onto positions of the upper - left corner, the upper - right corner, and the lower - left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, a predicted motion vector v0 for the upper - left corner control point of the current block, a predicted motion vector v1 for the upper - right corner control point of the current block, and a predicted motion vector v2 for the lower - left corner control point of the current block are calculated.

[0354] In addition, in the derivation of the predicted motion vector for each of the control points of the current block in step Sj_1 described later Figure 29 , this predicted motion vector derivation method can also be used.

[0355] Figure 27 is a flowchart showing an example of the affine merge mode.

[0356] In the affine merge mode, first, the inter - frame prediction unit 126 derives a predicted MV for each of the control points of the current block (step Sk_1). The control points are, as Figure 25A shown, the points of the upper - left corner and the upper - right corner of the current block, or as Figure 25B shown, the points of the upper - left corner, the upper - right corner, and the lower - left corner of the current block.

[0357] That is, as Figure 26A shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left), and determines the initial valid block encoded in the affine mode.

[0358] Then, when block A is determined and block A has two control points, as Figure 26B shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner and the motion vector v1 of the control point at the upper right corner of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block including block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner and the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0359] Alternatively, when block A is determined and block A has three control points, as Figure 26C shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner, the motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block including block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner, the predicted motion vector v1 of the control point at the upper right corner, and the motion vector v2 of the control point at the lower left corner of the current block.

[0360] Next, the inter-frame prediction unit 126 performs motion compensation for each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 126 calculates the motion vector of the sub-block as an affine MV using two predicted motion vectors v0 and v1 and the above formula (1A), or three predicted motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter-frame prediction unit 126 performs motion compensation for the sub-block using the affine MV and the encoded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0361] [MV Derivation>Affine Inter-frame Mode]

[0362] Figure 28A is a diagram for explaining the affine inter-frame mode with two control points.

[0363] In this affine inter-frame mode, as Figure 28AAs shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0364] Figure 28B FIG. is for explaining the affine inter-frame mode with three control points.

[0365] In this affine inter-frame mode, as Figure 28B shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block. In addition, the motion vector selected from the motion vectors of the encoded blocks F and G adjacent to the current block is used as the predicted motion vector v2 of the control point at the lower left corner of the current block.

[0366] Figure 29 FIG. is a flowchart showing an example of the affine inter-frame mode.

[0367] In the affine inter-frame mode, first, the inter-frame prediction unit 126 derives the predicted MV (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). As Figure 25A or Figure 25B shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0368] That is, the inter-frame prediction unit 126 derives the predicted motion vector (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the motion vector of a certain block in the encoded blocks near each control point of the current block as shown in Figure 28A or Figure 28B At this time, the inter-frame prediction unit 126 encodes the prediction motion vector selection information for identifying the two selected motion vectors into the stream.

[0369] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select from the encoded blocks adjacent to the current block as the predicted motion vector of the control point by using cost evaluation, etc., and can describe a flag indicating which predicted motion vector is selected in the bitstream.

[0370] Next, while using the prediction motion vector selected or derived in the update step Sj_1, the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4) (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vectors of the respective sub-blocks corresponding to the prediction motion vector to be updated as the affine MVs, and calculates them using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference picture (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the prediction motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the prediction motion vector as a differential MV into the stream.

[0371] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).

[0372] [MV Derivation>Affine Inter-frame Mode]

[0373] In the case where different numbers of control points (for example, 2 and 3) are signaled in the affine mode at the CU level, sometimes the number of control points in the encoded block and the current block is different. Figure 30A And Figure 30B is a conceptual diagram for explaining a method for deriving a prediction vector of a control point in the case where the number of control points in the encoded block and the current block is different.

[0374] For example, as Figure 30A shown, when the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points, motion vectors v3 and v4 projected onto the positions of the upper left corner and upper right corner of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the upper left corner control point and the predicted motion vector v1 of the upper right corner control point of the current block are calculated. In addition, the predicted motion vector v2 of the lower left corner control point is calculated based on the derived motion vectors v0 and v1.

[0375] For example, as Figure 30BAs shown, when the current block has two control points at the upper left corner and the upper right corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 projected onto the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner of the current block are calculated.

[0376] In Figure 29 the derivation of each predicted motion vector of the control points of the current block in step Sj_1, this predicted motion vector derivation method can also be used.

[0377] [MV Derivation > DMVR]

[0378] Figure 31A is a diagram showing the relationship between the merge mode and DMVR.

[0379] The inter-frame prediction unit 126 derives the motion vector of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, that is, a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the motion vector derived in step Sl_1 as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is determined in the merge mode.

[0380] On the other hand, when it is determined in step Sl_1 to perform a motion search (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final motion vector for the current block by searching the peripheral area of the reference picture represented by the motion vector derived in step Sl_1 (step Sl_3). That is, in this case, the motion vector of the current block is determined by DMVR.

[0381] Figure 31B is a conceptual diagram for explaining an example of the DMVR process for determining the MV.

[0382] First, the best MVP set for the current block (for example, in the merge mode) is set as the candidate MV. Then, according to the candidate MV (L0), the reference pixels are determined based on the encoded picture in the L0 direction, that is, the first reference picture (L0). Similarly, according to the candidate MV (L1), the reference pixels are determined based on the encoded picture in the L1 direction, that is, the second reference picture (L1). A template is generated by taking the average of these reference pixels.

[0383] Next, using the above template, the peripheral regions of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV. In addition, the cost value can also be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search region, and the candidate MV values, etc.

[0384] In addition, in the encoding device and the decoding device described later, the structures and operations of the processes described here are basically common.

[0385] Even if it is not the process itself described here, as long as it is a process that can search the periphery of the candidate MV and derive the final MV, any process can be used.

[0386] [Motion Compensation > BIO / OBMC]

[0387] In motion compensation, there is a mode of generating a predicted image and correcting the predicted image. This mode is, for example, BIO and OBMC described later.

[0388] Figure 32 It is a flowchart showing an example of the generation of a predicted image.

[0389] The inter-frame prediction unit 126 generates a predicted image (step Sm_1), and corrects the predicted image by any of the above modes (step Sm_2).

[0390] Figure 33 It is a flowchart showing another example of the generation of a predicted image.

[0391] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting the predicted image (step Sn_4). On the other hand, when it is determined not to perform the correction process (No in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).

[0392] In addition, in motion compensation, there is a mode of correcting the luminance when generating a predicted image. This mode is, for example, LIC described later.

[0393] Figure 34 It is a flowchart showing yet another example of the generation of a predicted image.

[0394] The inter-frame prediction unit 126 derives the motion vector of the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform the luminance correction process (step So_2). Here, when it is determined to perform the luminance correction process (Yes in step So_2), the inter-frame prediction unit 126 generates a prediction image while performing luminance correction (step So_3). That is, the prediction image is generated by LIC. On the other hand, when it is determined not to perform the luminance correction process (No in step So_2), the inter-frame prediction unit 126 generates a prediction image by normal motion compensation without performing luminance correction (step So_4).

[0395] [Motion Compensation > OBMC]

[0396] Not only the motion information of the current block obtained through motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame prediction signal. Specifically, an inter-frame prediction signal can also be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained through motion search (within the reference picture) and a prediction signal based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0397] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. And information indicating whether to apply the OBMC mode (e.g., referred to as the OBMC flag) can also be signaled at the CU level. Additionally, the level at which these information are signaled is not limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0398] A more specific description of the OBMC mode will be given. Figure 35 and Figure 36 are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process based on OBMC processing.

[0399] First, as Figure 36 shown, using the motion vector (MV) assigned to the processing target (current) block, a prediction image (Pred) based on normal motion compensation is obtained. In Figure 36 , the arrow "MV" points to the reference picture and indicates which block in the reference picture the current block in the current picture refers to for obtaining the prediction image.

[0400] Next, the motion vector (MV_L) that has been derived for the already-encoded left adjacent block is applied (reused) to the block to be encoded, obtaining a predicted image (Pred_L). The motion vector (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0401] Similarly, the motion vector (MV_U) that has been derived for the already-encoded upper adjacent block is applied (reused) to the block to be encoded, obtaining a predicted image (Pred_U). The motion vector (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted image that has undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block whose boundaries with adjacent blocks are blended (smoothed).

[0402] Furthermore, the above example is a two-path correction method that uses the left adjacent and upper adjacent blocks, but this correction method can also be a three-path or more-path correction method that also uses the right adjacent and / or lower adjacent blocks.

[0403] In addition, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.

[0404] In addition, the prediction image correction process of OBMC has been described here, where the prediction image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can also be applied to each of the multiple reference pictures. In this case, through the OBMC image correction based on multiple reference pictures, after obtaining the corrected predicted images from each reference picture, the final predicted image is obtained by further overlapping the multiple obtained corrected predicted images.

[0405] In addition, in OBMC, the unit of the object block can be the prediction block unit or the sub-block unit obtained by further dividing the prediction block.

[0406] As a method for determining whether to apply OBMC processing, for example, there is a method of using a signal indicating whether to apply OBMC processing, namely obmc_flag. As a specific example, the encoding device can also determine whether the target block belongs to a region with complex motion. When the encoding device belongs to a region with complex motion, it sets the obmc_flag value to 1 and applies OBMC processing for encoding. When it does not belong to a region with complex motion, it sets the obmc_flag value to 0 and encodes the block without applying OBMC processing. On the other hand, in the decoding device, by decoding the obmc_flag described in the stream (such as the compressed sequence), it switches whether to apply OBMC processing for decoding according to this value.

[0407] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter-frame prediction unit 126 can generate multiple prediction images with shapes different from the rectangle for the rectangular current block, and can generate the final rectangular prediction image by combining these multiple prediction images. Shapes different from the rectangle can also be triangles, for example.

[0408] Figure 37 It is a diagram for explaining the generation of two triangular prediction images.

[0409] The inter-frame prediction unit 126 generates a triangular prediction image by performing motion compensation on the first triangular partition within the current block using the first MV of the first partition. Similarly, the inter-frame prediction unit 126 generates a triangular prediction image by performing motion compensation on the second triangular partition in the current block using the second MV of the second partition. Then, the inter-frame prediction unit 126 generates a rectangular prediction image identical to the current block by combining these prediction images.

[0410] In addition, in Figure 37 the example shown, the first partition and the second partition are triangles respectively, but they can also be trapezoids, or can be of different shapes respectively. Moreover, in Figure 37 the example shown, the current block is composed of two partitions, but it can also be composed of three or more partitions.

[0411] In addition, the first partition and the second partition can also be repeated. That is, the first partition and the second partition can also include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition can be used to generate the prediction image of the current block.

[0412] In addition, in this example, an example where the prediction images of both partitions are generated by inter-frame prediction is shown, but the prediction image for at least one partition can also be generated by intra-frame prediction.

[0413] [Motion Compensation > BIO]

[0414] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.

[0415] Figure 38 is a diagram for explaining a model assuming uniform linear motion. In Figure 38 , (v x , v y ) represents a velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.

[0416] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) holds.

[0417] [Equation 3]

[0418]

[0419] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from a merge list or the like can be corrected in pixel units.

[0420] In addition, a motion vector can also be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, a motion vector can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0421] [Motion Compensation > LIC]

[0422] Next, an example of a mode for generating a predicted image (prediction) by using LIC (local illumination compensation) processing will be described.

[0423] Figure 39 This is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC processing.

[0424] First, an MV is derived from the encoded reference picture, and a reference image corresponding to the current block is obtained.

[0425] Next, information indicating how the luminance values change in the reference picture and the current picture is extracted for the current block. This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions within the reference picture specified by the derived MV. Then, using the information indicating how the luminance values change, a luminance correction parameter is calculated.

[0426] By applying the above luminance correction parameter to the reference image within the reference picture specified by the MV, a predicted image for the current block is generated.

[0427] In addition, Figure 39 the shape of the above peripheral reference region in Figure 39 is an example, and shapes other than this can also be used.

[0428] Furthermore, the processing for generating a predicted image based on one reference picture has been described here, but the same applies when generating a predicted image based on multiple reference pictures. It is also possible to generate a predicted image after performing luminance correction processing on the reference images obtained from each reference picture in the same manner as described above.

[0429] As a method for determining whether to adopt LIC processing, for example, there is a method of using lic_flag, which is a signal indicating whether to adopt LIC processing. As a specific example, in an encoding device, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as lic_flag, and encoding is performed using LIC processing. If it does not belong to a region where a luminance change has occurred, the value 0 is set as lic_flag, and encoding is performed without using LIC processing. On the other hand, in a decoding device, it is also possible to decode the lic_flag described in the stream and switch whether to adopt LIC processing according to its value for decoding.

[0430] As another method for determining whether to use the LIC process, there is a method of determining whether the LIC process is used in the surrounding blocks. As a specific example, when the current block is in the merge mode, it is determined whether the surrounding coded blocks selected when deriving the MV in the merge mode process are coded using the LIC process, and based on the result, it is switched whether to use the LIC process for coding. In addition, in the case of this example, the same process is also applied to the decoding device side.

[0431] use Figure 39 The LIC process (luminance correction process) has been described above, and its details will be described below.

[0432] First, the inter prediction unit 126 derives a motion vector for acquiring a reference image corresponding to a current block to be encoded from a reference picture that is an already encoded picture.

[0433] Next, the inter-frame prediction unit 126 uses the brightness pixel values ​​of the encoded surrounding reference areas adjacent to the left and above and the brightness pixel values ​​at the same position in the reference picture specified by the motion vector to extract information indicating how the brightness values ​​in the reference picture and the encoding target picture change, and calculates the brightness correction parameter. For example, the brightness pixel value of a certain pixel in the surrounding reference area in the encoding target picture is set to p0, and the brightness pixel value of the pixel in the surrounding reference area in the reference picture at the same position as the pixel is set to p1. The inter-frame prediction unit 126 calculates the coefficients A and B for optimizing A×p1+B=p0 as brightness correction parameters for multiple pixels in the surrounding reference area.

[0434] Next, the inter-frame prediction unit 126 generates a predicted image for the encoding target block by performing a brightness correction process on the reference image in the reference picture specified by the motion vector using the brightness correction parameter. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after the brightness correction process is set to p3. The inter-frame prediction unit 126 generates a predicted image after the brightness correction process by calculating A×p2+B=p3 for each pixel in the reference image.

[0435] also, Figure 39 The shape of the peripheral reference area in is an example, and other shapes may be used. Figure 39 For example, a region including a predetermined number of pixels thinned out from the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference region. In addition, the peripheral reference region is not limited to the region adjacent to the encoding target block, and may also be a region not adjacent to the encoding target block. Figure 39In the example shown, the peripheral reference region in the reference picture is the region specified by the motion vector of the coded object picture from among the peripheral reference regions in the coded object picture, but it may also be a region specified by another motion vector. For example, this other motion vector may also be the motion vector of the peripheral reference region in the coded object picture.

[0436] In addition, the operation in the encoding device 100 has been described here, but the operation in the decoding device 200 is the same.

[0437] Furthermore, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0438] Furthermore, the LIC process can also be applied in units of sub-blocks. For example, correction parameters can be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0439] [Prediction control unit]

[0440] The prediction control unit 128 selects one of the intra prediction signal (the signal output from the intra prediction unit 124) and the inter prediction signal (the signal output from the inter prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0441] As Figure 1 shown, in various installation examples, the prediction control unit 128 can also output prediction parameters input to the entropy encoding unit 110. The entropy encoding unit 110 can generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device. The decoding device can also receive and decode the encoded bitstream and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selection of the prediction signal (e.g., the motion vector, the prediction type, or the prediction mode used by the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0442] [Installation example of encoding device]

[0443] Figure 40 is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 1 shown, the multiple components of the encoding device 100 are composed ofFigure 40 It is implemented by installing the shown processor a1 and memory a2.

[0444] The processor a1 is a circuit for information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding moving images. The processor a1 can also be a processor such as a CPU. In addition, the processor a1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 can also play the role of Figure 1 multiple components of the encoding device 100 shown in etc., except for the components for storing information.

[0445] The memory a2 is a dedicated or general-purpose memory for storing information for the processor a1 to encode moving images. The memory a2 can be either an electronic circuit or connected to the processor a1. In addition, the memory a2 can also be included in the processor a1. In addition, the memory a2 can also be an aggregate of multiple electronic circuits. In addition, the memory a2 can be a magnetic disk or an optical disk, etc., or can be represented as a storage or a recording medium, etc. In addition, the memory a2 can be either a non-volatile memory or a volatile memory.

[0446] For example, the memory a2 can store the encoded moving image or the bit string corresponding to the encoded moving image. In addition, a program for the processor a1 to encode moving images can also be stored in the memory a2.

[0447] In addition, for example, the memory a2 can also play the role of Figure 1 the components for storing information among the multiple components of the encoding device 100 shown in etc. Specifically, the memory a2 can play the role of Figure 1 the block memory 118 and the frame memory 122 shown in. More specifically, the memory a2 can store the reconstructed blocks and the reconstructed pictures, etc.

[0448] In addition, in the encoding device 100, all of the multiple components shown in etc. may not be installed, and all of the above-mentioned multiple processes may not be performed. Figure 1 A part of the multiple components shown in etc. can be included in other devices, or a part of the above-mentioned multiple processes can be performed by other devices. Figure 1

[0449] [Decoding Device]

[0450] Next, a decoding device that can decode the encoded signal (encoded bit stream) output from the above encoding device 100 will be described. Figure 41 Figure 41FIG. 0 is a block diagram showing the functional configuration of the decoding apparatus 200 according to the present embodiment. The decoding apparatus 200 is a moving image decoding apparatus that decodes moving images in units of blocks.

[0451] As Figure 41 shown in FIG. 1, the decoding apparatus 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0452] The decoding apparatus 200 is implemented by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoding apparatus 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0453] Hereinafter, after explaining the overall processing flow of the decoding apparatus 200, each component included in the decoding apparatus 200 will be described.

[0454] [Overall Flow of Decoding Process]

[0455] Figure 42 FIG. 2 is a flowchart showing an example of the overall decoding process performed by the decoding apparatus 200.

[0456] First, the entropy decoding unit 202 of the decoding apparatus 200 determines a segmentation pattern (step Sp_1) of a block (128×128 pixels) having a fixed size. This segmentation pattern is the segmentation pattern selected by the encoding apparatus 100. Then, the decoding apparatus 200 performs the processes of steps Sp_2 to Sp_6 on each of the plurality of blocks constituting the segmentation pattern.

[0457] That is, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the block to be decoded (also referred to as the current block) (step Sp_2).

[0458] Next, the inverse quantization unit 204 and the inverse transform unit 206 restore a plurality of prediction residuals (i.e., differential blocks) by performing inverse quantization and inverse transform on the plurality of quantization coefficients (step Sp_3).

[0459] Next, a prediction processing unit, which is composed of all or part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220, generates a prediction signal (also referred to as a prediction block) for the current block (step Sp_4).

[0460] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the differential block (step Sp_5).

[0461] Moreover, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).

[0462] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If the determination is that it has not been completed (No in step Sp_7), the processing from step Sp_1 is repeated.

[0463] In addition, the processing of these steps Sp_1 to Sp_7 can be sequentially performed by the decoding device 200. Multiple processes of a part of these processes can be performed in parallel, or the order can be changed.

[0464] [Entropy Decoding Unit]

[0465] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202 arithmetic decodes the encoded bitstream into a binary signal, for example. Next, the entropy decoding unit 202 de-binarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the encoded bitstream to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 (refer to Figure 1 ) the prediction parameters included therein. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.

[0466] [Inverse Quantization Unit]

[0467] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoding target block (hereinafter referred to as the current block) that is the input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. And the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0468] [Inverse Transform Unit]

[0469] The inverse transform unit 206 restores the prediction error by performing an inverse transform on the transform coefficients that are input from the inverse quantization unit 204.

[0470] For example, when the information read from the coded bitstream indicates the use of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type that has been read.

[0471] In addition, for example, when the information read from the coded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re - transform to the transform coefficients.

[0472] [Addition unit]

[0473] The addition unit 208 reconstructs the current block by adding the prediction error that is input from the inverse transform unit 206 and the prediction sample that is input from the prediction control unit 220. Further, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0474] [Block memory]

[0475] The block memory 210 is a storage unit that stores blocks within the decoded picture (hereinafter referred to as the current picture) that are referred to in intra - prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.

[0476] [Loop filter unit]

[0477] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.

[0478] When the information indicating the ON / OFF of ALF read from the coded bitstream indicates that ALF is ON, one filter is selected from among a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.

[0479] [Frame memory]

[0480] The frame memory 214 is a storage unit that stores reference pictures used in inter - prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0481] [Prediction processing unit (intra - prediction unit / inter - prediction unit / prediction control unit)]

[0482] Figure 43This is a diagram showing an example of the processing performed by the prediction processing unit of the decoding device 200. In addition, the prediction processing unit is composed of all or some of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0483] The prediction processing unit generates a predicted image of the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Additionally, in the prediction signal, there are, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit uses the reconstructed image that has been obtained by generating a prediction block, a differential block, a coefficient block, restoring the differential block, and generating a decoded image block, to generate a predicted image of the current block.

[0484] The reconstructed image can be, for example, an image of a reference picture, or an image of the decoded blocks within the current picture that includes the current block. The decoded blocks within the current picture are, for example, adjacent blocks of the current block.

[0485] Figure 44 This is a diagram showing another example of the processing performed by the prediction processing unit of the decoding device 200.

[0486] The prediction processing unit determines the method or mode for generating the predicted image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.

[0487] When it is determined that the first method is the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the first method (step Sr_2a). In addition, when it is determined that the second method is the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the second method (step Sr_2b). In addition, when it is determined that the third method is the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the third method (step Sr_2c).

[0488] The first method, the second method, and the third method are different methods for generating the predicted image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.

[0489] [Intra Prediction Unit]

[0490] The intra prediction unit 216 performs intra prediction by referring to the blocks within the current picture stored in the block memory 210 based on the intra prediction mode read from the encoded bitstream, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 generates an intra prediction signal by referring to the samples (such as luminance values, chrominance differences) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0491] In addition, when an intra prediction mode of a reference luminance block is selected in the intra prediction of the color difference block, the intra prediction unit 216 may also predict the color difference component of the current block based on the luminance component of the current block.

[0492] In addition, when the information decoded from the coded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0493] [Inter prediction unit]

[0494] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using motion information (e.g., motion vectors) decoded from the coded bitstream (e.g., prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction signal for the current block or sub-block, and outputting the inter prediction signal to the prediction control unit 220.

[0495] When the information decoded from the coded bitstream indicates the adoption of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained through motion estimation but also the motion information of adjacent blocks.

[0496] In addition, when the information decoded from the coded bitstream indicates the adoption of the FRUC mode, the inter prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) decoded from the coded stream, thereby deriving motion information. And the inter prediction unit 218 uses the derived motion information for motion compensation (prediction).

[0497] In addition, when the BIO mode is adopted, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when the information decoded from the coded bitstream indicates the adoption of the affine motion compensation prediction mode, the inter prediction unit 218 derives a motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.

[0498] [MV derivation > Normal inter-frame mode]

[0499] When the information read from the coded bitstream indicates the application of the normal inter-frame mode, the inter prediction unit 218 derives an MV based on the information read from the coded bitstream, and uses the MV for motion compensation (prediction).

[0500] Figure 45 is a flowchart showing an example of inter prediction based on the normal inter-frame mode in the decoding apparatus 200.

[0501] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Ss_1). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0502] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Ss_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined order of priority (step Ss_2). In addition, this order of priority is predetermined for each of the N prediction MV candidates.

[0503] Next, the inter-frame prediction unit 218 decodes prediction motion vector selection information from the input stream (i.e., the encoded bitstream), and uses the decoded prediction motion vector selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Ss_3).

[0504] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction motion vector (step Ss_4).

[0505] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).

[0506] [Prediction control unit]

[0507] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device side.

[0508] [Installation example of decoding device]

[0509] Figure 46 is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 41 As shown, the multiple components of the decoding device 200 are through Figure 46It is installed with the processor b1 and the memory b2 shown.

[0510] The processor b1 is a circuit for information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes an encoded moving image (i.e., an encoded bitstream). The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor b1 can also play the role of Figure 41 multiple components of the decoding device 200 shown in etc., except for the components for storing information.

[0511] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode the encoded bitstream. The memory b2 can be an electronic circuit or can be connected to the processor b1. In addition, the memory b2 can also be included in the processor b1. In addition, the memory b2 can also be an aggregate of multiple electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disk, etc., and can also be represented as a storage or a recording medium, etc. In addition, the memory b2 can be either a non-volatile memory or a volatile memory.

[0512] For example, the memory b2 can store a moving image or an encoded bitstream. In addition, a program for the processor b1 to decode the encoded bitstream can also be stored in the memory b2.

[0513] In addition, for example, the memory b2 can play the role of Figure 41 components for storing information among multiple components of the decoding device 200 shown in etc. Specifically, the memory b2 can play the role of Figure 41 the block memory 210 and the frame memory 214 shown. More specifically, the memory b2 can store the reconstructed blocks and the reconstructed pictures, etc.

[0514] In addition, in the decoding device 200, all of the multiple components shown in etc. may not be installed, Figure 41 nor may all of the above-mentioned multiple processes be performed. Figure 41 A part of the multiple components shown in etc. can be included in other devices, or a part of the above-mentioned multiple processes can be performed by other devices.

[0515] [Definitions of each term]

[0516] As an example, each term can also be defined as follows.

[0517] The picture is an arrangement of multiple luminance samples in a monochrome format, or an arrangement of multiple luminance samples and two corresponding arrangements of multiple chrominance samples in a color format of 4:2:0, 4:2:2, or 4:4:4. The picture can be a frame or a field.

[0518] A frame is a composition of a top field that generates multiple sample lines 0, 2, 4,... and a bottom field that generates multiple sample lines 1, 3, 5,....

[0519] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same access unit.

[0520] A tile is a rectangular area of multiple coding tree blocks within a specific tile column and a specific tile row in a picture. A tile can still apply a loop filter across the edges of the tile, but it can also be a rectangular area of a frame intended to be decoded and encoded independently.

[0521] A block is an MxN (N rows and M columns) arrangement of multiple samples, or an MxN arrangement of multiple transform coefficients. A block can also be a square or rectangular area of multiple pixels composed of multiple matrices of one luminance and two chrominances.

[0522] A CTU (Coding Tree Unit) can be a coding tree block of multiple luminance samples of a picture with a 3-sample arrangement, or two corresponding coding tree blocks of multiple chrominance samples. Alternatively, a CTU can also be a coding tree block of any multiple samples in a monochrome picture and a picture encoded using syntax constructs used in the encoding of three separate color planes and multiple samples.

[0523] A superblock constitutes one or two mode information blocks, or it can also be a 64×64 pixel square block that can be recursively divided into four 32×32 blocks and thus can be divided.

[0524] [First form]

[0525] Hereinafter, an encoding device 100, a decoding device 200, an encoding method, and a decoding method according to the first form of the present invention will be described.

[0526] Figure 47 It is a flowchart showing an example of the inter-frame prediction process in the first form. Hereinafter, an example of the inter-frame prediction process in the decoding device 200 will be described.

[0527] The inter-frame prediction unit 218 in the decoding device 200 derives a reference motion vector for predicting a processing target block, derives a first motion vector different from the reference motion vector, derives a differential motion vector based on the difference between the reference motion vector and the first motion vector, determines whether the differential motion vector is greater than a threshold value, changes the first motion vector when it is determined that the differential motion vector is greater than the threshold value, does not change the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and decodes the processing target block using the changed first motion vector or the unchanged first motion vector. For example, the reference motion vector corresponds to a first pixel set within the processing target block, and the first motion vector corresponds to a second pixel set different from the first pixel set within the processing target block. Hereinafter, the inter-frame prediction process will be described more specifically with reference to the drawings. In addition, the reference motion vector and the first motion vector described below are examples and are not limited thereto.

[0528] First, in step S1001, the inter-frame prediction unit 218 derives a reference motion vector for the first pixel set within the processing target block. Hereinafter, with reference to Figure 48 The process of step S1001 will be described more specifically.

[0529] Figure 48 is a diagram showing an example of the processing target block. As Figure 48 shown, the processing target block may also be composed of a plurality of sub-blocks (sub-block 0 to sub-block 5). In addition, each sub-block may respectively have different first motion vectors (MV0 to MV5). An example of the first pixel set may also be sub-block 0. In this case, the reference motion vector is one of the plurality of first motion vectors within the processing target block. In this example, the reference motion vector is MV0.

[0530] Another example of the first pixel set may also be the entire processing target block. In this example, the reference motion vector is the average of the first motion vectors of all sub-blocks within the processing target block, that is, the average from MV0 to MV5.

[0531] Next, in step S1002, the inter-frame prediction unit 218 derives the first motion vector for the second pixel set within the processing target block. The second pixel set is different from the first pixel set. An example of the second pixel set may also be sub-block 2. In this example, the first motion vector of the second pixel set is MV2.

[0532] In addition, another example of the second pixel set may also be sub-block 1. In this example, the first motion vector is MV1.

[0533] Next, in step S1003, the inter-frame prediction unit 218 derives a differential vector based on the difference between the reference motion vector of the first pixel set and the first motion vector of the second pixel set. An example of the value of the reference motion vector may be (-3, 4). An example of the first motion vector may be (16, 5). Therefore, the differential motion vector between the first motion vector and the reference motion vector becomes (-19, -1).

[0534] Next, in step S1004, the inter-frame prediction unit 218 determines whether the differential motion vector derived in step S1003 is greater than a threshold. The threshold is the first value (hereinafter referred to as the first threshold) and the second value (hereinafter referred to as the second threshold) of one set. An example of the threshold may also be (10, 20). Hereinafter, an example of this determination process will be described.

[0535] When the absolute value of the horizontal component of the differential motion vector derived in step S1003 is greater than the first threshold, or the absolute value of the vertical component of the differential motion vector is greater than the second threshold, the inter-frame prediction unit 218 determines that the differential motion vector is greater than the threshold (Yes in step S1004). In other words, if either the absolute value of the horizontal component or the absolute value of the vertical component of the differential motion vector is greater than the threshold, it is determined that the differential motion vector is greater than the threshold. For example, when comparing the absolute value of the horizontal component of the differential motion vector (-19, -1) exemplified in step S1003 with the first threshold, since |-19| > 10, it is determined that the differential motion vector is greater than the threshold. In this case, the inter-frame prediction unit 218 changes the first motion vector. Details of the change will be described in step S1006.

[0536] On the other hand, when the absolute value of the horizontal component of the differential motion vector is not greater than the first threshold and the absolute value of the vertical component of the differential motion vector is not greater than the second threshold, the inter-frame prediction unit 218 determines that the differential motion vector is not greater than the threshold (No in step S1004). In other words, when both the absolute value of the horizontal component and the absolute value of the vertical component of the differential motion vector are below the threshold, it is determined that the differential motion vector is below the threshold. When the inter-frame prediction unit 218 determines that the differential motion vector is not greater than the threshold (No in step S1004), it does not change the first motion vector (step S1005).

[0537] Next, in step S1006, when it is determined that the differential motion vector is greater than the threshold (Yes in step S1004), the inter-frame prediction unit 218 changes the first motion vector using the value obtained by clipping the differential motion vector. At this time, the differential motion vector (-19, -1) is clipped to (-10, -1). As described above, since the absolute value |-19| of the horizontal component of the differential motion vector is greater than 10, which is the first threshold of the thresholds (10, 20), the absolute value of the horizontal component of the differential motion vector is clipped to be the same as the first threshold. Therefore, the clipped differential motion vector becomes (-10, -1).

[0538] Next, the inter-frame prediction unit 218 changes the first motion vector using the clipped differential motion vector and the reference motion vector. More specifically, the first motion vector can also be changed by adding the differential motion vector and the reference motion vector. For example, the changed first motion vector is (-3, 4) - (-10, -1) = (7, 5).

[0539] In step S1007, the inter-frame prediction unit 218 decodes the second pixel set using the changed first motion vector or the unchanged first motion vector. For example, sub-block 2 is decoded using the changed first motion vector (7, 5).

[0540] In addition, Figure 47 the processing 1000 shown can also be the processing of an encoding device.

[0541] The processing 1000 can be applied to all sub-blocks within the processing target block. When the processing 1000 is applied to all sub-blocks within the processing target block, the prediction mode of the processing target block can also be the affine mode. In addition, when the processing 1000 is applied to all sub-blocks within the processing target block, the prediction mode of the processing target block can also be the ATMVP (Alternative Temporal Motion Vector Prediction) mode.

[0542] In addition, the ATMVP mode is an example of a sub-block mode classified as a merge mode. For example, in the encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, the temporal MV reference block corresponding to the current block is identified, and for each sub-block within the current block, the MV used for encoding in the region corresponding to the sub-block within the temporal MV reference block is identified.

[0543] In addition, the first motion vector of other sub-blocks of the processing target block can also be updated using the changed first motion vector in a specific sub-block. Hereinafter, a processing example will be described.

[0544] For example, MV0 of sub-block 0 is determined as the reference motion vector, MV2 of sub-block 2 is determined as the first motion vector, and the first motion vector after the change is determined using process 1000. The first motion vector after the change of sub-block 2 is MV2'=(V 2x ',V 2y Using the reference motion vector MV0 and the changed first motion vector MV2', the changed first motion vector MV of the sub-block i (i=1, 3, 4, 5) other than the sub-block 2 in the processing target block is calculated by the following formula: i '=(V ix ',V iy ').

[0545] V ix '=(V 2x '-V 0x )*POS ix / W-(V 2y '-V 0y )*POS iy / W+V 0x

[0546] V iy '=(V 2y '-V 0y )*POS ix / W+(V 2x '-V 0x )*POS iy / W+V 0y

[0547] Then, in the MV i 'If it is not greater than the threshold, it is used as MV i 'Update MV i , use MV i Decode sub-block i. POS ix and POS iy are the horizontal and vertical positions of sub-block i.

[0548] exist Figure 48 In the example,

[0549] POS 1x =W / 2, POS 1y =0

[0550] POS 3x =0, POS 3y =H

[0551] POS 4x =W / 2, POS 4y =H

[0552] POS5x = W, POS 5y = H

[0553] W and H are the horizontal and vertical positions of the sub-block 2 (e.g., the X coordinate and Y coordinate of the lower left corner).

[0554] The threshold value can also correspond to the size of the block to be processed. For example, the larger the size of the block to be processed, the larger the threshold value.

[0555] In addition, the threshold value can also correspond to the number of reference pictures. For example, the more the number of reference pictures, the smaller the threshold value.

[0556] The threshold value can also be encoded into a header area such as the SPS header, PPS header, or slice header.

[0557] The threshold value can also be predetermined without being decoded for the stream.

[0558] The threshold value can also be determined to limit the following situation: in the case of performing motion compensation processing (prediction processing) on the block to be processed in a specified prediction mode, the worst-case memory access amount becomes less than or equal to the memory access amount in the case of performing bidirectional motion compensation processing (prediction processing) on the block to be processed at every 8×8 pixels in a prediction mode other than the specified prediction mode. The specified prediction mode is, for example, the affine mode. Hereinafter, an example of the limitation will be described.

[0559] "Mem_base" represents the worst-case memory access amount per unit of one 8×8 size block in the case of performing bidirectional predictive motion compensation processing (luminance value) on the block to be processed at every 8×8 size in a prediction mode other than the specified prediction mode.

[0560] The worst-case memory access amount in the case of performing motion compensation processing (luminance value) on the block to be processed (size: M×N) in a specified prediction mode is represented by "Mem_CU", and the threshold value "Mem_th" for this is represented by the following using "Mem_base".

[0561] Mem_th = M×N / (8×8)*Mem_base

[0562] Mem_CU is limited to be not greater than Mem_th. Hereinafter, an example of the calculation process of the threshold value will be described.

[0563] Assuming that the size of each pixel of the luminance signal is 1 byte and the number of taps of the filter for motion compensation processing is 8 taps, Mem_base = (8 + 7)*(8 + 7)*2 = 450 (in byte unit).

[0564] When it is assumed that the prediction mode of the processing target block is the affine mode and the size of the processing target block is 64×64 pixels, Mem_th = 64 / 8 * 64 / 8 * 450 = 28800 (in bytes).

[0565] Assume that “H” and “V” respectively represent the threshold of the first component of the threshold (the first threshold) and the threshold of the second component (the second threshold).

[0566] The worst condition in the case of performing prediction processing in an ordinary inter-frame mode other than the affine mode is as follows: The processing target block is divided into 8×8 pixel blocks, and all 8×8 pixel blocks are motion compensated by bidirectional prediction. In a way that does not exceed the memory access amount of this worst condition, the range of the memory that can be referred to in the case of performing prediction processing in the affine mode is determined. Assume that the range of the memory that can be referred to in the affine mode is calculated by the following formula [1]. Calculate H and V so that the range of memory access in the affine mode is below Mem_th (for example, 28800). Additionally, here, an example of the case of bidirectional prediction in this affine mode is described.

[0567] Mem_CU = (64 + 7 + 2*H)(64 + 7 + 2*V)*2 ≤ 28800 [1]

[0568] Assume H = V, solve formula [1], then H = V = 24.

[0569] In addition, the above calculation process of the threshold is also applied to processing target blocks with different sizes and different numbers of reference pictures. Figure 49 It is a diagram showing an example of the calculated threshold. Additionally, in Figure 49 , an example of H = V is shown, but it can also be H > V or H < V.

[0570] As Figure 49 shown, the threshold is different according to the case of unidirectional prediction (number of reference frames = 1) and bidirectional prediction (number of reference frames = 2) of the processing target block. In the example of the above formula [1], the size of the processing target block predicted in the affine mode is 64×64. In the case of bidirectional prediction, that is, when referring to 2 reference pictures, both the first threshold H and the second threshold V are 24. For example, in the case of a 64×64 size processing target block predicted in the affine mode by unidirectional prediction, that is, when only referring to 1 reference picture, both the first threshold H and the second threshold V are 49. In Figure 49 , the thresholds for all block sizes that can be predicted in the affine mode are calculated using the above formula [1] and tabulated.

[0571] [Technical Advantages of the First Form]

[0572] In the first aspect of the present invention, the derivation process of the first motion vector is introduced in the inter-frame prediction process. As described above, selecting an appropriate first motion vector such that the deviation of a plurality of first motion vectors within the processing target block converges within a specified range can reduce the memory bandwidth of the inter-frame prediction process.

[0573] [Supplement]

[0574] The encoding device 100 and the decoding device 200 in the present embodiment can be used as an image encoding device and an image decoding device, respectively, or can be used as a moving image encoding device and a moving image decoding device.

[0575] Alternatively, the encoding device 100 and the decoding device 200 can also be used as an entropy encoding device and an entropy decoding device, respectively. That is, the encoding device 100 and the decoding device 200 can also correspond only to the entropy encoding unit 110 and the entropy decoding unit 202, respectively. Moreover, other components can be included in other devices.

[0576] In addition, at least a part of the present embodiment can be used as an encoding method, a decoding method, an entropy encoding method, an entropy decoding method, or other methods.

[0577] In addition, in the present embodiment, each component is constituted by dedicated hardware, but can also be implemented by executing a software program suitable for each component. Each component can be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[0578] Specifically, each of the encoding device 100 and the decoding device 200 can include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to the processor a1 or b1, and the storage device corresponds to the memory a2 or b2.

[0579] The processing circuit includes at least one of dedicated hardware and a program execution unit, and uses the storage device to execute processing. In addition, when the processing circuit includes a program execution unit, the storage device stores a software program executed by the program execution unit.

[0580] Here, the software for implementing the encoding device 100 or the decoding device 200 of the present embodiment is the following program.

[0581] For example, the program can also cause a computer to execute the following encoding method: The encoding method is an encoding method for encoding a moving image, deriving a reference motion vector used in the prediction of a processing target block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on the difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector when it is determined that the differential motion vector is greater than the threshold value, not changing the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and encoding the processing target block using the changed first motion vector or the unchanged first motion vector.

[0582] In addition, for example, the program can also cause a computer to execute the following decoding method: The decoding method is a decoding method for decoding a moving image, deriving a reference motion vector used in the prediction of a processing target block, deriving a first motion vector different from the reference motion vector, deriving a differential motion vector based on the difference between the reference motion vector and the first motion vector, determining whether the differential motion vector is greater than a threshold value, changing the first motion vector when it is determined that the differential motion vector is greater than the threshold value, not changing the first motion vector when it is determined that the differential motion vector is not greater than the threshold value, and decoding the processing target block using the changed first motion vector or the unchanged first motion vector.

[0583] In addition, as described above, each component can also be a circuit. These circuits can either form a single circuit as a whole or be separate different circuits. In addition, each component can be implemented by a general-purpose processor or a dedicated processor.

[0584] In addition, the processing performed by a specific component can also be executed by other components. In addition, the order of executing the processing can be changed, or multiple processes can be executed simultaneously. In addition, it can also be that the encoding / decoding device includes an encoding device 100 and a decoding device 200.

[0585] In addition, ordinal numbers such as first and second used in the description can also be appropriately replaced. In addition, for components and the like, ordinal numbers can be newly assigned or removed.

[0586] As described above, the configurations of the encoding device 100 and the decoding device 200 have been described based on the embodiments, but the configurations of the encoding device 100 and the decoding device 200 are not limited to these embodiments. As long as the gist of the present invention is not deviated from, configurations obtained by various modifications conceivable by those skilled in the art to the present embodiments and configurations constructed by combining constituent elements in different embodiments may also be included in the scope of the configurations of the encoding device 100 and the decoding device 200.

[0587] This method may also be implemented in combination with at least a part of other methods in the present invention. In addition, a part of the processing described in the flowchart of this method, a part of the structure of the device, a part of the grammar, etc. may be combined with other methods for implementation.

[0588] (Embodiment 2)

[0589] [Implementation and Application]

[0590] In each of the above embodiments, each functional block or operative block can generally be implemented by an MPU (microprocessing unit), a memory, etc. In addition, it may be that the processing of each functional block is implemented by a program execution unit such as a processor that reads and executes software (program) recorded in a recording medium such as a ROM. This software can be distributed. This software can also be recorded in various recording media such as a semiconductor memory. In addition, each functional block can also be implemented by hardware (a dedicated circuit).

[0591] The processing described in each embodiment can be implemented by centralized processing using a single device (system), or can also be implemented by distributed processing using multiple devices. In addition, the processor that executes the above program can be single or multiple. That is, centralized processing or distributed processing can be performed.

[0592] The configuration of the present invention is not limited to the above embodiments, and various changes can be made, and they are also included in the scope of the configuration of the present invention.

[0593] Furthermore, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) described in each of the above embodiments and various systems for implementing such application examples are described here. It may also be that such a system is characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding / decoding device having both. Regarding other structures of such a system, they can be appropriately changed according to circumstances.

[0594] [Usage Example]

[0595] Figure 50It is a diagram showing the overall structure of a suitable content supply system ex100 for implementing a content distribution service. The provision of communication services is divided into desired sizes, and in each unit, base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are provided respectively.

[0596] In this content supply system ex100, various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 are connected via the Internet ex101 through an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. Some of the above devices may be combined and connected in this content supply system ex100. In various implementations, the devices may also be directly or indirectly connected to each other via a telephone network or short-range wireless without passing through base stations ex106 to ex110. Moreover, a streaming media server ex103 may also be connected to various devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smart phone ex115 via the Internet ex101 or the like. In addition, the streaming media server ex103 may also be connected to terminals in a hotspot in an aircraft ex117 via a satellite ex116.

[0597] Alternatively, a wireless access point or a hotspot or the like may be used instead of base stations ex106 to ex110. In addition, the streaming media server ex103 may be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the aircraft ex117 without passing through the satellite ex116.

[0598] The camera ex113 is a device such as a digital camera that can perform still image photography and moving image photography. In addition, the smart phone ex115 is a smart phone, a mobile phone, or a PHS (Personal Handyphone System) or the like corresponding to the modes of mobile communication systems called 2G, 3G, 3.9G, 4G, and 5G in the future.

[0599] The home appliances ex114 are a refrigerator or devices included in a household fuel cell cogeneration system or the like.

[0600] In the content supply system ex100, a terminal having a photographing function is connected to a streaming media server ex103 via a base station ex106 or the like, whereby live distribution or the like can be performed. In live distribution, terminals (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, and a terminal in an airplane ex117) can perform the encoding process described in the above embodiments on still image or moving image content photographed by a user using the terminal, can also multiplex video data obtained by encoding and audio data obtained by encoding sound corresponding to the video, and can send the obtained data to the streaming media server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present invention.

[0601] On the other hand, the streaming media server ex103 performs stream distribution on content data sent to a requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smart phone ex115, or a terminal in an airplane ex117 that can decode the data after the above encoding process. Each device that receives the distributed data performs decoding processing on the received data and reproduces it. That is, each device can also function as an image decoding device according to one aspect of the present invention.

[0602] [Distributed processing]

[0603] In addition, the streaming media server ex103 can also be a plurality of servers or a plurality of computers, and perform distributed processing or recording and distribution of data. For example, the streaming media server ex103 can be implemented by a CDN (Content Delivery Network), and content distribution is achieved through a network connecting many edge servers dispersed in the world to each other. In the CDN, a physically closer edge server is dynamically allocated according to the client. And by caching and distributing content to the edge server, latency can be reduced. In addition, in the case of several types of errors occurring or when the communication state changes due to an increase in traffic or the like, the processing can be distributed among multiple edge servers, or the distribution entity can be switched to another edge server, or a part of the network that has failed can be bypassed and distribution can continue, so high-speed and stable distribution can be achieved.

[0604] In addition, not limited to decentralized processing of distributing itself, the encoding process of the captured data can be performed by each terminal, on the server side, or can be shared among them. As an example, generally two processing loops are performed in the encoding process. In the first loop, the complexity or encoding amount of the image in units of frames or scenes is detected. In addition, in the second loop, a process of improving the encoding efficiency while maintaining the image quality is performed. For example, by performing the first encoding process by the terminal and the second encoding process by the server that receives the content, it is possible to improve the quality and efficiency of the content while reducing the processing load in each terminal. In this case, if there is a request to receive and decode almost in real time, the data completed by the first encoding performed by the terminal can also be received and reproduced by other terminals, so more flexible real-time distribution can also be performed.

[0605] As other examples, the camera ex113 etc. extracts feature amounts from the image, compresses the data regarding the feature amounts as metadata, and sends it to the server. The server, for example, judges the importance of the target based on the feature amounts and switches the quantization accuracy etc., and performs compression corresponding to the meaning of the image (or the importance of the content). The feature amount data is particularly effective for improving the accuracy and efficiency of motion vector prediction during re-compression in the server. In addition, simple encoding such as VLC (Variable Length Coding) can be performed by the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) can be performed by the server.

[0606] As other examples, in a stadium, a shopping mall, a factory, etc., there are cases where there are multiple video data obtained by multiple terminals capturing substantially the same scene. In this case, multiple terminals that have performed the shooting and, if necessary, other terminals and servers that have not performed the shooting are used, and decentralized processing is performed by respectively allocating the encoding process in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture, etc. Thereby, it is possible to reduce the delay and better achieve real-time performance.

[0607] Since the multiple video data are of substantially the same scene, the server can also manage and / or instruct to refer to the video data captured by each terminal with each other. In addition, it can also be that the server receives the encoded data from each terminal and changes the reference relationship among the multiple data, or corrects or replaces the picture itself and re-encodes it. Thereby, it is possible to generate a stream with improved quality and efficiency of each data.

[0608] Moreover, the server can also perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server can change the MPEG-like encoding method to the VP class (e.g., VP9), or can change H.264 to H.265.

[0609] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, the following descriptions use terms such as "server" or "terminal" as the processing entity. However, part or all of the processing performed by the server can also be performed by the terminal, and part or all of the processing performed by the terminal can also be performed by the server. In addition, the same applies to the decoding process regarding these.

[0610] [3D, Multi-angle]

[0611] The cases of merging and using images or videos of different scenes captured by multiple terminals such as cameras ex113 and / or smartphones ex115 that are roughly synchronized with each other, or the same scene captured from different angles are increasing. The videos captured by each terminal are merged based on the relative position relationship between the terminals obtained separately, or the regions where the feature points included in the videos are consistent, etc.

[0612] The server not only encodes two-dimensional moving images, but can also encode still images automatically or at a user-specified time based on scene analysis of the moving images and send them to the receiving terminal. When the server can obtain the relative position relationship between the shooting terminals, it can not only generate the three-dimensional shape of the scene based on the videos of the same scene captured from different angles in addition to two-dimensional moving images. The server can also encode the three-dimensional data generated by point clouds, etc. separately, or select or reconstruct from the videos captured by multiple terminals based on the results of identifying or tracking people or objects using the three-dimensional data to generate and send the video to the receiving terminal.

[0613] In this way, the user can not only arbitrarily select each image corresponding to each shooting terminal to view the scene, but also view the content of the video cut from the three-dimensional data reconstructed using multiple images or videos at the selected viewing point. Furthermore, together with the video, sound can also be collected from multiple different angles, and the server multiplexes the sound from a specific angle or space with the corresponding video and sends the multiplexed video and sound.

[0614] In addition, in recent years, content that establishes a correspondence between the real world and the virtual world such as Virtual Reality (VR) and Augmented Reality (AR) has been spreading. In the case of VR images, the server separately creates viewpoint images for the right eye and the left eye, and can perform encoding that allows reference between the viewpoint videos through Multi-View Coding (MVC), etc., or can encode them as different streams without referring to each other. When decoding different streams, they can be reproduced synchronously according to the user's viewpoint to reproduce a virtual three-dimensional space.

[0615] In the case of an AR image, it is also possible that the server overlaps the virtual object information in the virtual space with the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device acquires or holds the virtual object information and the three-dimensional data, generates a two-dimensional image according to the movement of the user's viewpoint, and creates the overlapping data by smoothly connecting them. Alternatively, it is also possible that the decoding device sends the movement of the user's viewpoint to the server in addition to the delegation of the virtual object information. It is also possible that the server creates the overlapping data according to the three-dimensional data held in the server, matches the received movement of the viewpoint, encodes the overlapping data, and distributes it to the decoding device. In addition, the overlapping data has an α value representing the transmittance in addition to RGB, and the server sets the α value of the part other than the target created according to the three-dimensional data to 0, etc., and encodes it in a state where it is transmitted in this part. Alternatively, the server can also set the RGB value of a specified value as the background like chroma key, and generate data with the part other than the target set as the background color.

[0616] Similarly, the decoding process of the distributed data can be performed by each terminal as a client, on the server side, or can be shared between them. As an example, it is also possible that a certain terminal first sends a reception request to the server, and another terminal receives the content corresponding to the request and performs the decoding process, and sends the decoded signal to the device with a display. By dispersing the processing and selecting appropriate content regardless of the performance of the communicable terminal itself, it is possible to reproduce data with better image quality. In addition, as another example, it is also possible that a TV or the like receives large-size image data, and a personal terminal of the viewer decodes and displays a part of the area such as tiles after the picture is divided. Thus, while making the whole image shared, it is possible to confirm one's own responsible area or the area that one wants to confirm in more detail at hand.

[0617] In a situation where multiple short-range, medium-range, or long-range wireless communications indoors and outdoors can be used, it may be possible to receive content seamlessly using a distribution system standard such as MPEG-DASH. The user can also freely select the user's terminal, decoding devices or display devices such as displays arranged indoors and outdoors, and switch in real time. In addition, it is possible to use one's own position information, etc., to switch the decoding terminal and the display terminal and perform decoding. Thus, it is also possible to map and display information on a part of the wall surface or the ground of the building next to the device that can be displayed during the user's movement to the destination. In addition, it is also possible to switch the bit rate of the received data based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed by the receiving terminal in a short time, or the encoded data being replicated in the edge server of the content distribution service.

[0618] [Scalable Coding]

[0619] Regarding the content switching, use Figure 51 the scalable stream shown, which is compression-encoded using the moving image encoding method represented in the above-described respective embodiments, for explanation. For the server, there may be multiple streams with the same content but different qualities as separate streams, or it may be a structure that switches content by utilizing the characteristics of the temporally / spatially scalable stream achieved by hierarchical encoding as shown in the figure. That is, the decoding side can freely switch between decoding low-resolution content and high-resolution content by determining which layer to decode based on internal factors such as performance and external factors such as the state of the communication band. For example, when a user wants to view the subsequent video that was viewed on a smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream to different layers, thus reducing the burden on the server side.

[0620] Furthermore, in addition to the structure where pictures are encoded for each layer as described above and the scalability of the enhancement layer above the base layer is realized, the enhancement layer may include meta-information based on statistical information of the image, etc. It may also be that the decoding side generates high-quality content by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can improve the signal-to-noise ratio while maintaining and / or expanding the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least squares operation used in the super-resolution process, etc.

[0621] Alternatively, a structure may be provided that divides a picture into tiles, etc. according to the meaning of an object, etc. within the image. The decoding side decodes only a part of the region by selecting the tiles to be decoded. Moreover, by saving the attributes of the object (person, car, ball, etc.) and the position within the image (coordinate position within the same image, etc.) as meta-information, the decoding side can determine the position of the desired object based on the meta-information and decide on the tiles including the object. For example, as Figure 52 shown, a data storage structure different from the pixel data, such as the SEI (supplemental enhancement information) message in HEVC, may be used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0622] The meta-information may also be stored in units composed of multiple pictures, such as a stream, sequence, or random access unit. The decoding side can obtain the moment when a specific person appears within the video, and by matching with the picture unit information and time information, can determine the picture in which the object exists and can decide on the position of the object within the picture.

[0623] [Optimization of Web Page]

[0624] Figure 53 It is a diagram showing an example of a display screen of a Web page in a computer ex111 or the like. Figure 54 It is a diagram showing an example of a display screen of a Web page in a smart phone ex115 or the like. As Figure 53 and Figure 54 shown, there is a case where a Web page includes a plurality of linked images as links to image content, and the visible manner thereof varies depending on the viewing device. When a plurality of linked images can be seen on the screen, before the user explicitly selects a linked image, or before the linked image approaches near the center of the screen or the whole of the linked image enters the screen, the display device (decoding device) can display a still image or an I picture that each content has as a linked image, can also display an image such as a gif animation using a plurality of still images or I pictures, or can also receive only the base layer and decode and display the image.

[0625] When a linked image is selected by the user, the display device gives the highest priority to the base layer and decodes it. In addition, if there is information indicating that the content is scalable in the HTML constituting the Web page, the display device can also decode up to the enhancement layer. Moreover, in order to ensure real-time performance or when the communication band is very tight before selection, the display device can reduce the delay between the decoding time and the display time of the start picture (the delay from the start of decoding of the content to the start of display) by decoding and displaying only the forward-referenced pictures (I pictures, P pictures, B pictures that perform only forward reference). Furthermore, the display device can also forcibly ignore the reference relationship of the pictures, set all B pictures and P pictures as forward reference and roughly decode them, and perform normal decoding as the pictures received over time increase.

[0626] [Automatic Driving]

[0627] In addition, when receiving or transmitting still image or video data such as two-dimensional or three-dimensional map information for the automatic driving or driving assistance of a vehicle, the receiving terminal can also receive information such as weather or construction information as meta information in addition to the image data belonging to one or more layers, and decode them in correspondence. In addition, the meta information can belong to a layer or can be multiplexed only with the image data.

[0628] In this case, since vehicles, drones, airplanes, etc. that include the receiving terminal are moving, the receiving terminal can perform seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting the location information of the receiving terminal. In addition, the receiving terminal can dynamically switch the degree of receiving meta information and / or the degree of updating map information according to the user's selection, the user's situation, and / or the status of the communication band.

[0629] In the content supply system ex100, the client can receive, decode, and reproduce the encoded information sent by the user in real time.

[0630] [Distribution of Personal Content]

[0631] In addition, in the content supply system ex100, not only high-quality, long-duration content provided by video distribution operators, but also unicast or multicast distribution of low-quality, short-duration content provided by individuals can be performed. It is conceivable that such personal content will increase in the future. In order to make personal content better, the server can also perform encoding processing after editing processing. This can be achieved, for example, with the following structure.

[0632] During or after shooting in real time or cumulatively, the server performs recognition processing such as shooting error, scene search, meaning analysis, and target detection based on the original image data or the encoded data. And based on the recognition result, the server manually or automatically performs editing such as correcting focus deviation or camera shake, deleting scenes with low importance such as scenes with lower brightness or out-of-focus than other pictures, emphasizing the edges of the target, or changing the color tone. Based on the editing result, the server encodes the edited data. In addition, it is known that the viewing rate will decrease if the shooting time is too long. The server can also automatically limit not only scenes with low importance as described above but also scenes with little movement based on the image processing result according to the shooting time to make the content within a specific time range. Or, the server can also generate a summary based on the result of the meaning analysis of the scene and encode it.

[0633] In the original state, personal content may be invaded by content that infringes copyright, the moral rights of the author, or the right of portrait, etc. There may also be inconvenient situations for individuals, such as the sharing scope exceeding the desired range. Therefore, for example, the server can also encode by forcibly changing the faces of people in the peripheral part of the screen or in the home, etc. into out-of-focus images. Moreover, the server can also identify whether a face of a person different from the pre-registered person is captured in the image to be encoded, and in the case of capture, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, from the perspective of copyright, etc., the user can specify the person or background area for which the image is to be processed. The server can also perform processing such as replacing the specified area with another image or blurring the focus. If it is a person, the person can be tracked in the moving image and the image of the face part of the person can be replaced.

[0634] The real-time requirement for viewing and listening to personal content with a small data volume is relatively strong. Therefore, although it also depends on the bandwidth, the decoding device first receives and decodes and reproduces the base layer with the highest priority. The decoding device can also receive the enhancement layer during this period. In the case where the reproduction is looped and reproduced more than twice, the enhancement layer is also included to reproduce a high-quality image. In this way, if it is a scalable-encoded stream, an experience can be provided where the moving image is rough at the stage of not being selected or just starting to watch, but the stream gradually becomes smooth and the image quality improves. In addition to scalable encoding, the same experience can also be provided when the first rough stream and the second stream encoded with reference to the first moving image form one stream.

[0635] [Other implementation application examples]

[0636] In addition, these encoding or decoding processes are usually processed in the LSIex500 provided in each terminal. The LSI (large-scale integration circuitry) ex500 (refer to Figure 50 ) can be either a single chip or a structure composed of multiple chips. In addition, software for encoding or decoding moving images can also be loaded into a certain recording medium (CD-ROM, floppy disk, hard disk, etc.) that can be read by a computer ex111, etc., and the encoding process and decoding process can be performed using this software. Furthermore, in the case where the smart phone ex115 is equipped with a camera, the moving image data obtained by this camera can also be sent. The moving image data at this time is the data after being encoded by the LSIex500 provided in the smart phone ex115.

[0637] In addition, the LSIex500 may also be configured to download and activate application software. In this case, the terminal first determines whether the terminal corresponds to the encoding method of the content or has the ability to execute a specific service. If the terminal does not correspond to the encoding method of the content or does not have the ability to execute a specific service, the terminal downloads a codec or application software and then acquires and reproduces the content.

[0638] In addition, not limited to the content supply system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can also be incorporated in a digital broadcasting system. Since the radio wave for broadcasting carries multiplexed data in which video and audio are multiplexed by using a satellite or the like and is transmitted and received, there is a difference suitable for multicast compared to the structure of the content supply system ex100 which is prone to unicast, but the encoding process and the decoding process can be applied in the same way.

[0639] [Hardware Structure]

[0640] Figure 55 is a further detailed representation Figure 50 of the smart phone ex115 shown. In addition, Figure 56 is a diagram showing a structural example of the smart phone ex115. The smart phone ex115 has an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying the video captured by the camera unit ex465 and decoding the data such as the video received by the antenna ex450. The smart phone ex115 also includes an operation unit ex466 such as a touch panel, a sound output unit ex457 such as a speaker for outputting sound or audio, a sound input unit ex456 such as a microphone for inputting sound, a memory unit ex467 capable of storing the captured video or still images, the recorded sound, the received video or still images, the encoded or decoded data of e-mails, etc., or a slot unit ex464 as an interface unit with the SIM ex468, and the SIM ex468 is used to identify the user and perform authentication for accessing various data represented by the network. In addition, an external memory may be used instead of the memory unit ex467.

[0641] The main control unit ex460 that comprehensively controls the display unit ex458, the operation unit ex466, etc. is interconnected with the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / demultiplexing unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 in synchronization via the bus ex470.

[0642] If the power key is turned on by the user's operation, the power supply circuit unit ex461 starts the smart phone ex115 to an operable state and supplies power to each unit from the battery pack.

[0643] Based on the control of the main control unit ex460 having a CPU, a ROM, a RAM, etc., the smart phone ex115 performs processes such as calls and data communication. During a call, the audio signal collected by the audio input unit ex456 is converted into a digital audio signal by the audio signal processing unit ex454, spectrum spreading processing is performed by the modulation / demodulation unit ex452, and digital-to-analog conversion processing and frequency conversion processing are performed by the transmission / reception unit ex451, and the signal of the result is transmitted via the antenna ex450. In addition, the received data is amplified and frequency conversion processing and analog-to-digital conversion processing are performed, spectrum inverse spreading processing is performed by the modulation / demodulation unit ex452, and after being converted into an analog audio signal by the audio signal processing unit ex454, it is output from the audio output unit ex457. During data communication, text, still images, or video data are sent to the main control unit ex460 via the operation input control unit ex462 based on the operations of the operation unit ex466, etc. of the main body unit. The same transmission and reception processing is performed. In the data communication mode, when sending video, still images, or video and audio, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 during the process of the camera unit ex465 shooting video or still images, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded audio data in a prescribed manner, and modulation processing and conversion processing are performed by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and it is transmitted via the antenna ex450.

[0644] In the case of receiving an image attached to an email or a chat tool, or an image linked on a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bitstream of video data and a bitstream of audio data by demultiplexing the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method described in the above embodiments, and displays the image or still image included in the linked moving image file from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal and outputs the sound from the audio output unit ex457. Since real-time streaming media is becoming more and more popular, depending on the user's situation, there may also be cases where the reproduction of sound is not suitable in society. Therefore, it may also be a structure in which, as a first value, it is preferable not to reproduce the audio signal but only reproduce the video data, and the sound is reproduced synchronously only when the user performs an operation such as clicking on the video data.

[0645] In addition, here, the smart phone ex115 is taken as an example for explanation, but as a terminal, in addition to the transceiver terminal having both an encoder and a decoder, three other installation forms can be considered: a transmitting terminal having only an encoder and a receiving terminal having only a decoder. In the digital broadcast system, it is assumed that the multiplexed data in which the audio data is multiplexed in the video data is received and transmitted for explanation. However, in the multiplexed data, character data associated with the video, etc. can also be multiplexed in addition to the audio data. In addition, it is also possible to receive or transmit the video data itself instead of the multiplexed data.

[0646] In addition, it is assumed that the main control unit ex460 including the CPU controls the encoding or decoding process for explanation, but in many cases, various terminals have a GPU. Therefore, it can also be configured to use the performance of the GPU to process a larger area together through a memory shared by the CPU and the GPU, or a memory that manages addresses in a shared manner. As a result, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is more effective if the processes of motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transform / quantization are performed not by the CPU but by the GPU together in units of pictures, etc.

[0647] Industrial availability

[0648] The present invention can be used in, for example, a television receiver, a digital video recorder, a car navigation system, a mobile phone, a digital camera, a digital video camera, a video conferencing system, or an endoscope, etc.

[0649] Description of Reference Numerals

[0650] 100 Encoding device

[0651] 102 Splitting section

[0652] 104 Subtraction section

[0653] 106 Transformation section

[0654] 108 Quantization section

[0655] 110 Entropy encoding section

[0656] 112, 204 Inverse quantization section

[0657] 114, 206 Inverse transformation section

[0658] 116, 208 Addition section

[0659] 118, 210 Block memory

[0660] 120, 212 Loop filtering section

[0661] 122, 214 Frame memory

[0662] 124, 216 Intra prediction section

[0663] 126, 218 Inter prediction section

[0664] 128, 220 Prediction control section

[0665] 200 Decoding device

[0666] 202 Entropy decoding section

[0667] 1000 Processing

[0668] 1201 Boundary determination section

[0669] 1202, 1204, 1206 Switch

[0670] 1203 Filtering determination section

[0671] 1205 Filtering processing section

[0672] 1207 Filtering characteristic determination section

[0673] 1208 Processing determination section

[0674] a1, b1 Processor

[0675] a2 and b2 memories

Claims

1. An encoding device that encodes a moving image, wherein, The encoding device includes: a circuit; and a memory connected to the circuit, wherein the prediction mode of the current block to be encoded is the affine mode, and in operation, the circuit: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; when it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and when it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and encodes the current block using the second motion vector.

2. A decoding device that decodes a moving image, wherein, The decoding device includes: a circuit; and a memory connected to the circuit, wherein the prediction mode of the current block to be decoded is the affine mode, and in operation, the circuit: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; when it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and when it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and decodes the current block using the second motion vector.

3. A computer-readable non-transitory storage medium storing a bitstream, wherein the prediction mode of the current block to be decoded is the affine mode, the bitstream includes an encoded signal and syntax information, and the decoding device performs, based on the encoded signal and the syntax information: derives a reference motion vector for predicting the current block; derives a first motion vector different from the reference motion vector; derives a motion vector difference based on the difference between the reference motion vector and the first motion vector; determines whether the motion vector difference is greater than a threshold; when it is determined that the motion vector difference is greater than the threshold, sets a first value for a second motion vector, and when it is determined that the motion vector difference is not greater than the threshold, sets a second value different from the first value for the second motion vector, the second motion vector being different from the reference motion vector and the first motion vector; and decodes the current block using the second motion vector.