Encoding device, decoding device, and non-transitory computer readable medium

By deriveing multiple gradient and differential value parameters in the video encoding device and the decoding device, and generating high-precision predicted images, the problem of excessive computing volume and circuit scale in bidirectional optical flow processing is solved, and the encoding efficiency and processing speed are improved.

CN120378620APending Publication Date: 2025-07-25PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510726572.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-06-21
Filing Date
2020-06-15
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing video encoding technology has a need for improvement in encoding efficiency, picture quality, processing volume, circuit scale and processing speed, especially when using bidirectional optical flow processing, the computing volume and circuit scale may be too large.

Method used

By derive multiple parameters, such as absolute and differential values of horizontal and vertical gradients in the encoding and decoding device, a prediction image is generated, and high-precision prediction images are generated while reducing the amount of computing and circuit size.

Benefits of technology

It improves coding efficiency, simplifies processing flow, reduces processing volume and circuit scale, improves processing speed, and appropriately selects encoding elements such as filters, block sizes and motion vectors, improving picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378620A_ABST
    Figure CN120378620A_ABST
Patent Text Reader

Abstract

The invention relates to an encoding device, a decoding device, and a non-transitory computer readable medium. The encoding device includes a memory and a processor connected to the memory, derives, as a first parameter, a sum of a plurality of horizontal gradients and absolute values derived for each of a plurality of relative pixel positions, derives, as a second parameter, a sum of a plurality of vertical gradients and absolute values, and derives, as a third parameter, a sum of a plurality of horizontal corresponding pixel difference values. The method includes deriving a sum of a plurality of vertically corresponding pixel difference values as a fourth parameter, deriving a sum of a plurality of vertically corresponding horizontal gradient sums as a fifth parameter, and generating a prediction image to encode the current block using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter, the horizontal gradient value representing a difference value between a right sample value and a left sample value, and the horizontal gradient value representing a difference value between the right sample value and the left sample value. The right sample is adjacent to the right side of a target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in the first prediction block and the second prediction block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of the patent application for an invention titled "Coding Device, Decoding Device, Coding Method, and Decoding Method", with an application date of June 15, 2020, an application number of 202080043963.3. Technical Field

[0002] The present invention relates to video coding, for example, systems, components, and methods in the coding and decoding of moving images. Background Art

[0003] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to process the continuously increasing amount of digital video data in various applications, there is always a need to provide improvements and optimizations to video coding technology.

[0004] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above video coding technology.

[0005] Prior Art Documents

[0006] Non-Patent Documents

[0007] Non-Patent Document 1: H.265(ISO / IEC 23008-2HEVC) / HEVC(High Efficiency Video Coding) Summary of the Invention

[0008] Problems to be Solved by the Invention

[0009] Regarding the coding methods as described above, it is desired to propose new methods for improving coding efficiency, improving image quality, reducing processing volume, reducing circuit scale, or appropriately selecting elements or actions such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks.

[0010] The present invention provides a structure or method that can contribute to one or more of, for example, improving coding efficiency, improving image quality, reducing processing volume, reducing circuit scale, improving processing speed, and appropriately selecting elements or actions. In addition, the present invention may include a structure or method that can contribute to benefits other than the above.

[0011] Means for Solving the Problems

[0012] For example, an encoding device according to an aspect of the present invention includes:

[0013] including:

[0014] a memory; and

[0015] a processor, connected to the memory, generating a prediction image based on the derived first parameter, second parameter, third parameter, fourth parameter, and fifth parameter in BDOF (Bidirectional Optical Flow) processing to encode the current block.

[0016] Based on

[0017] [Formula 1]

[0018] Σ [i,j]∈Ω abs(I x 1 +I x 0 )

[0019] the first parameter is derived.

[0020] Based on

[0021] [Formula 2]

[0022] ∑ [i,j]∈Ω abs (I y 1 +I y 0 )

[0023] the second parameter is derived.

[0024] Based on

[0025] [Formula 3]

[0026] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 ))

[0027] the third parameter is derived.

[0028] Based on

[0029] [Formula 4]

[0030] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I1))

[0031] Derive the fourth parameter, and

[0032] Based on

[0033] [Equation 5]

[0034] Σ [i,j]∈Ω (sign(I y 1 +I y 0 )×(I x 1 +I x 0 ))

[0035] Derive the fifth parameter,

[0036] where Ω represents a set of multiple relative pixel positions,

[0037] [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω,

[0038] I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position is determined based on the relative pixel position [i, j], and the first gradient image and the second gradient image correspond to the current block,

[0039] I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image,

[0040] I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block,

[0041] I 1 represents the pixel value at the first pixel position in the second interpolation image corresponding to the current block,

[0042] The abs function outputs the absolute value of the independent variable,

[0043] The sign function outputs the sign of the independent variable, and the sign of this independent variable is -1, 0 or 1,

[0044] I x 0Indicates the difference value between the right sample value and the left sample value. The right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in the first prediction block. Moreover

[0045] I x 1 Indicates the difference value between the right sample value and the left sample value. The right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in the second prediction block different from the first prediction block.

[0046] For example, a decoding device according to an aspect of the present invention includes:

[0047] A memory; and

[0048] A processor, connected to the memory, generates a prediction image based on the derived first parameter, second parameter, third parameter, fourth parameter, and fifth parameter in BDOF (Bidirectional Optical Flow) processing to decode the current block.

[0049] Based on

[0050] [Equation 1]

[0051] Σ [i,j]∈Ω abs(I x 1 +I x 0 )

[0052] Derive the first parameter.

[0053] Based on

[0054] [Equation 2]

[0055] ∑ [i,j]∈Ω abs(I y 1 +I y 0 )

[0056] Derive the second parameter.

[0057] Based on

[0058] [Equation 3]

[0059] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 ))

[0060] Derive the third parameter.

[0061] Based on

[0062] [Equation 4]

[0063] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 ))×(I 0 -I 1 ))

[0064] Derive the fourth parameter, and

[0065] Based on

[0066] [Equation 5]

[0067] ∑ [i,j]∈Ω (sign(I y 1 +I y 0 ))×(I x 1 +I x 0 ))

[0068] Derive the fifth parameter,

[0069] where Ω represents a set of multiple relative pixel positions,

[0070] [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω,

[0071] I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position is determined based on the relative pixel position [i, j], and the first gradient image and the second gradient image correspond to the current block,

[0072] I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image,

[0073] I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block,

[0074] I 1 represents the pixel value at the first pixel position in the second interpolation image corresponding to the current block

[0075] The abs function outputs the absolute value of the argument

[0076] The sign function outputs the sign of the argument, and the sign of the argument is -1, 0, or 1

[0077] I x 0 represents the difference value between the right sample value and the left sample value. The right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, the target sample is included in the first prediction block, and

[0078] I x 1 represents the difference value between the right sample value and the left sample value. The right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in a second prediction block different from the first prediction block

[0079] A non - transitory computer - readable medium according to one aspect of the present invention stores a bitstream. The bitstream includes prediction parameters. The plurality of prediction parameters represent prediction processes among prediction process candidates and are read by a decoding device to decode a current block provided in the bitstream. The plurality of prediction process candidates include BDOF, i.e., bidirectional optical flow processing. In the BDOF processing, a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter are used to generate a prediction image of the current block

[0080] Based on

[0081] [Equation 1]

[0082] Σ [i,j]∈Ω abs(I x 1 +I x 0 )

[0083] derive the first parameter

[0084] Based on

[0085] [Equation 2]

[0086] Σ [i,j]∈Ω abs(I y 1 +I y 0 )

[0087] derive the second parameter

[0088] Based on

[0089] [Equation 3]

[0090] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 ))

[0091] Derive the third parameter,

[0092] Based on

[0093] [Equation 4]

[0094] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I 1 ))

[0095] Derive the fourth parameter, and

[0096] Based on

[0097] [Equation 5]

[0098] ∑ [i,j]∈Ω (sign(I y 1 +I y 0 )×(I x 1 +I x 0 ))

[0099] Derive the fifth parameter,

[0100] where Ω represents a set of multiple relative pixel positions,

[0101] [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω,

[0102] I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position is determined based on the relative pixel position [i, j], and the first gradient image and the second gradient image correspond to the current block,

[0103] I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image,

[0104] I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block,

[0105] I 1 represents the pixel value at the first pixel position in the second interpolation image corresponding to the current block,

[0106] The abs function outputs the absolute value of the argument,

[0107] The sign function outputs the sign of the argument, and the sign of the argument is -1, 0, or 1,

[0108] I x 0 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, the target sample is included in the first prediction block, and

[0109] I x 1 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in the second prediction block different from the first prediction block.

[0110] The installation of several embodiments in the present invention can improve the coding efficiency, simplify the coding / decoding process, accelerate the coding / decoding processing speed, and also efficiently select appropriate components / actions used in coding and decoding, such as filters, block sizes, motion vectors, reference pictures, reference blocks, etc.

[0111] According to the description and the drawings, the further advantages and effects in one aspect of the present invention are clarified. These advantages and / or effects are obtained respectively through several embodiments and the features described in the description and the drawings, but it is not necessary to provide all of them in order to obtain one or more advantages and / or effects.

[0112] In addition, these general or specific aspects can also be implemented by a system, a method, an integrated circuit, a computer program, a recording medium, or any combination thereof.

[0113] Advantages of the Invention

[0114] The structure or method of one aspect of the present invention can contribute to, for example, improvement of coding efficiency, improvement of image quality, reduction of processing amount, reduction of circuit scale, improvement of processing speed, and appropriate selection of elements or operations, etc., for one or more of them. In addition, the structure or method of one aspect of the present invention can also contribute to benefits other than the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Figure 1 It is a block diagram showing the functional structure of the encoding device of the embodiment.

[0116] Figure 2 It is a flowchart showing an example of the overall encoding process performed by the encoding device.

[0117] Figure 3 It is a conceptual diagram showing an example of block division.

[0118] Figure 4A It is a conceptual diagram showing an example of the structure of a slice.

[0119] Figure 4B It is a conceptual diagram showing an example of the structure of a tile.

[0120] Figure 5A It is a table showing transform basis functions corresponding to various transform types.

[0121] Figure 5B It is a conceptual diagram showing an example of SVT (Spatially Varying Transform).

[0122] Figure 6A It is a conceptual diagram showing an example of the shape of the filter used in ALF (adaptive loop filter).

[0123] Figure 6B It is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0124] Figure 6C It is a conceptual diagram showing another example of the shape of the filter used in ALF.

[0125] Figure 7 It is a block diagram showing an example of the detailed structure of the loop filtering section that functions as a DBF (deblocking filter).

[0126] Figure 8 It is a conceptual diagram showing an example of deblocking filtering having filtering characteristics symmetric with respect to the block boundary.

[0127] Figure 9It is a conceptual diagram for explaining the block boundary for performing deblocking filtering processing.

[0128] Figure 10 It is a conceptual diagram showing an example of the Bs value.

[0129] Figure 11 It is a flowchart showing an example of the processing performed by the prediction processing unit of the encoding device.

[0130] Figure 12 It is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device.

[0131] Figure 13 It is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device.

[0132] Figure 14 It is a conceptual diagram showing an example of 67 intra prediction modes in the intra prediction of the embodiment.

[0133] Figure 15 It is a flowchart showing an example of the process of the basic processing of inter prediction.

[0134] Figure 16 It is a flowchart showing an example of motion vector derivation.

[0135] Figure 17 It is a flowchart showing another example of motion vector derivation.

[0136] Figure 18 It is a flowchart showing another example of motion vector derivation.

[0137] Figure 19 It is a flowchart showing an example of inter prediction based on the normal inter mode.

[0138] Figure 20 It is a flowchart showing an example of inter prediction based on the merge mode.

[0139] Figure 21 It is a conceptual diagram for explaining an example of the motion vector derivation process based on the merge mode.

[0140] Figure 22 It is a flowchart showing an example of the FRUC (frame rate up conversion) processing.

[0141] Figure 23 It is a conceptual diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along the motion trajectory.

[0142] Figure 24It is a conceptual diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture.

[0143] Figure 25A It is a conceptual diagram showing an example of the derivation of a motion vector in sub-block units based on motion vectors of multiple adjacent blocks.

[0144] Figure 25B It is a conceptual diagram showing an example of the derivation of a motion vector in sub-block units in an affine mode with three control points.

[0145] Figure 26A It is a conceptual diagram showing an affine merge mode.

[0146] Figure 26B It is a conceptual diagram showing an affine merge mode with two control points.

[0147] Figure 26C It is a conceptual diagram showing an affine merge mode with three control points.

[0148] Figure 27 It is a flowchart showing an example of the processing of an affine merge mode.

[0149] Figure 28A It is a conceptual diagram showing an affine inter-frame mode with two control points.

[0150] Figure 28B It is a conceptual diagram showing an affine inter-frame mode with three control points.

[0151] Figure 29 It is a flowchart showing an example of the processing of an affine inter-frame mode.

[0152] Figure 30A It is a conceptual diagram showing an affine inter-frame mode in which the current block has three control points and an adjacent block has two control points.

[0153] Figure 30B It is a conceptual diagram showing an affine inter-frame mode in which the current block has two control points and an adjacent block has three control points.

[0154] Figure 31A It is a flowchart showing a merge mode including DMVR (decoder motion vector refinement).

[0155] Figure 31B It is a conceptual diagram showing an example of DMVR processing.

[0156] Figure 32 It is a flowchart showing an example of the generation of a predicted image.

[0157] Figure 33 It is a flowchart showing another example of the generation of a predicted image.

[0158] Figure 34 It is a flowchart showing another example of the generation of a predicted image.

[0159] Figure 35 It is a flowchart for explaining an example of a predicted image correction process based on OBMC (overlapped block motion compensation) processing.

[0160] Figure 36 It is a conceptual diagram for explaining an example of a predicted image correction process based on OBMC processing.

[0161] Figure 37 It is a conceptual diagram for explaining the generation of predicted images of two triangles.

[0162] Figure 38 It is a conceptual diagram for explaining a model assuming uniform linear motion.

[0163] Figure 39 It is a conceptual diagram for explaining an example of a predicted image generation method using a luminance correction process based on LIC (local illumination compensation) processing.

[0164] Figure 40 It is a block diagram showing an installation example of an encoding device.

[0165] Figure 41 It is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0166] Figure 42 It is a flowchart showing an example of the overall decoding process performed by a decoding device.

[0167] Figure 43 It is a flowchart showing an example of the process performed by the prediction processing unit of a decoding device.

[0168] Figure 44 It is a flowchart showing another example of the process performed by the prediction processing unit of a decoding device.

[0169] Figure 45 It is a flowchart showing an example of inter-frame prediction based on a normal inter-frame mode in a decoding device.

[0170] Figure 46 It is a block diagram showing an installation example of a decoding device.

[0171] Figure 47 It is a flowchart showing an example of inter-frame prediction according to BIO.

[0172] Figure 48 It is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.

[0173] Figure 49 It is a flowchart showing a first specific example of decoding processing based on BIO in the embodiment. Figure 50 It is a conceptual diagram showing an example of the calculation of horizontal gradient values in the embodiment.

[0174] Figure 51 It is a conceptual diagram showing an example of the calculation of vertical gradient values in the embodiment.

[0175] Figure 52 It is a flowchart showing a third specific example of decoding processing based on BIO in the embodiment.

[0176] Figure 53 It is a flowchart showing a fourth specific example of decoding processing based on BIO in the embodiment.

[0177] Figure 54 It is a flowchart showing the operation of the encoding device in the embodiment.

[0178] Figure 55 It is a flowchart showing the operation of the decoding device in the embodiment.

[0179] Figure 56 It is a block diagram showing the overall structure of a content supply system that implements a content distribution service.

[0180] Figure 57 It is a conceptual diagram showing an example of an encoding structure during scalable encoding.

[0181] Figure 58 It is a conceptual diagram showing an example of an encoding structure during scalable encoding.

[0182] Figure 59 It is a conceptual diagram showing an example of a display screen of a web page.

[0183] Figure 60 It is a conceptual diagram showing an example of a display screen of a web page.

[0184] Figure 61 It is a block diagram showing an example of a smart phone.

[0185] Figure 62 It is a block diagram showing an example of the structure of a smart phone. Detailed implementation mode

[0186] In recent years, research has been conducted on encoding moving images using bi-directional optical flow. Bi-directional optical flow is also referred to as BIO or BDOF.

[0187] For example, in bi-directional optical flow, a predicted image is generated based on the optical flow equation. More specifically, in bi-directional optical flow, a predicted image with predicted values adjusted in pixel units is generated using parameters derived from the pixel values of a reference image in block units and the gradient values of the reference image in block units. By using bi-directional optical flow, there is a high possibility of generating a predicted image with high accuracy.

[0188] For example, an encoding device encodes a difference image between the predicted image and the original image. Then, a decoding device decodes the difference image and adds the difference image and the predicted image to generate a reconstructed image. By using a predicted image with high accuracy, the encoding amount of the difference image can be reduced. That is, by using bi-directional optical flow, there is a high possibility of reducing the encoding amount of a moving image.

[0189] On the other hand, parameters for bi-directional optical flow are derived based on the pixel values and gradient values at each pixel position of a reference image. Therefore, in the derivation of parameters for bi-directional optical flow, due to the operations performed for each pixel position, the amount of computation may increase and the circuit scale may increase.

[0190] Therefore, for example, an encoding device according to one aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit derives a horizontal gradient and an absolute value for each of a plurality of relative pixel positions. The plurality of relative pixel positions are a plurality of pixel positions commonly and relatively determined for both a first range of a first reference block including a current block and a second range of a second reference block including the current block, and are a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The sum of the plurality of horizontal gradients and absolute values derived for the plurality of relative pixel positions is derived as a first parameter. For each of the plurality of relative pixel positions, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient and absolute value, is derived. The sum of the plurality of vertical gradient and absolute values derived for the plurality of relative pixel positions is derived as a second parameter. For each of the plurality of relative pixel positions, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value, is derived. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the horizontal gradient sum, that is, the horizontal corresponding pixel difference value, is derived. The sum of the plurality of horizontal corresponding pixel difference values derived for the plurality of relative pixel positions is derived as a third parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding pixel difference value, is derived. The sum of the plurality of vertical corresponding pixel difference values derived for the plurality of relative pixel positions is derived as a fourth parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the vertical gradient sum, the positive or negative sign of the horizontal gradient sum is reversed or maintained, and the horizontal gradient sum whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding horizontal gradient sum, is derived. The sum of the plurality of vertical corresponding horizontal gradient sums derived for the plurality of relative pixel positions is derived as a fifth parameter.Generate a prediction image used in the encoding of the current block using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter.

[0191] Thus, in the operations performed for each pixel position, it is possible to reduce the substantial multiplication operations with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a prediction image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in encoding. In addition, it is possible to appropriately generate a prediction image based on a plurality of parameters including a parameter related to a horizontal gradient value, a parameter related to a vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[0192] For example, the circuit derives the first parameter through the following formula (11.1), the second parameter through the following formula (11.2), the third parameter through the following formula (11.3), the fourth parameter through the following formula (11.4), and the fifth parameter through the following formula (11.5). Ω represents the set of the plurality of relative pixel positions, [i, j] represents each of the plurality of relative pixel positions, and for each of the plurality of relative pixel positions, I x 0 represents the horizontal gradient value of the relative pixel position in the first range, I x 1 represents the horizontal gradient value of the relative pixel position in the second range, I y 0 represents the vertical gradient value of the relative pixel position in the first range, I y 1 represents the vertical gradient value of the relative pixel position in the second range, I 0 represents the pixel value of the relative pixel position in the first range, I 1 represents the pixel value of the relative pixel position in the second range, abs(I x 1 +I x 0 ) represents I x 1 +I x 0 the absolute value of, sign(I x 1 +I x 0 ) represents I x 1 +I x 0 the positive or negative sign of, abs(I y 1 +I y0 ) represents I y 1 +I y 0 The absolute value of, sign(I y 1 +I y 0 ) represents I y 1 +I y 0 The positive or negative sign of.

[0193] Thus, it is possible to derive a plurality of parameters with low computational complexity using pixel values, horizontal gradient values, and vertical gradient values.

[0194] And, for example, the circuit derives a sixth parameter by dividing the third parameter by the first parameter, and derives a seventh parameter by subtracting the product of the fifth parameter and the sixth parameter from the fourth parameter and dividing by the second parameter, and generates the predicted image using the sixth parameter and the seventh parameter.

[0195] Thus, it is possible to appropriately aggregate a plurality of parameters into two parameters corresponding to the horizontal direction and the vertical direction. It is possible to appropriately incorporate a parameter related to the horizontal gradient value into the parameter corresponding to the horizontal direction. It is possible to appropriately incorporate a parameter related to the vertical gradient value, a parameter related to both the horizontal gradient value and the vertical gradient value, and a parameter corresponding to the horizontal direction into the parameter corresponding to the vertical direction. Then, it is possible to generate a predicted image using these two parameters.

[0196] In addition, for example, the circuit derives the sixth parameter by the following formula (10.8), and derives the seventh parameter by the following formula (10.9), sG x represents the first parameter, sG y represents the second parameter, sG x dI represents the third parameter, sG y dI represents the fourth parameter, sG x G y represents the fifth parameter, u represents the sixth parameter, and Bits is a function that returns the value obtained by rounding up the binary logarithm of the independent variable to an integer.

[0197] Thus, it is possible to derive two parameters corresponding to the horizontal direction and the vertical direction with low computational complexity.

[0198] In addition, for example, the circuit derives a predicted pixel value of the processing target pixel position included in the current block by using the first pixel value of the first pixel position corresponding to the processing target pixel position in the first reference block, the first horizontal gradient value of the first pixel position, the first vertical gradient value of the first pixel position, the second pixel value of the second pixel position corresponding to the processing target pixel position in the second reference block, the second horizontal gradient value of the second pixel position, the second vertical gradient value of the second pixel position, the sixth parameter, and the seventh parameter, and generates the predicted image.

[0199] Thereby, it is possible to generate a predicted image by using two parameters corresponding to the horizontal direction and the vertical direction, etc., and the two parameters corresponding to the horizontal direction and the vertical direction can be appropriately reflected in the predicted image.

[0200] In addition, for example, the circuit derives the predicted pixel value by dividing the sum of the first pixel value, the second pixel value, the first correction value, and the second correction value by 2. The first correction value corresponds to the product of the difference between the first horizontal gradient value and the second horizontal gradient value and the sixth parameter, and the second correction value corresponds to the product of the difference between the first vertical gradient value and the second vertical gradient value and the seventh parameter.

[0201] Thereby, by using two parameters corresponding to the horizontal direction and the vertical direction, etc., it is possible to appropriately generate a predicted image.

[0202] In addition, for example, the circuit derives the predicted pixel value by Equation (10.10) described later. I 0 represents the first pixel value, I 1 represents the second pixel value, u represents the sixth parameter, I x 0 represents the first horizontal gradient value, I x 1 represents the second horizontal gradient value, v represents the seventh parameter, I y 0 represents the first vertical gradient value, I y 1 represents the second vertical gradient value.

[0203] Thereby, according to the formula associated with two parameters corresponding to the horizontal direction and the vertical direction, etc., it is possible to appropriately generate a predicted image.

[0204] In addition, for example, a decoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit derives a horizontal gradient and an absolute value for each of a plurality of relative pixel positions. The plurality of relative pixel positions are a plurality of pixel positions that are commonly and relatively determined for both a first range including a first reference block containing the current block and a second range including a second reference block containing the current block, and are a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The sum of the plurality of horizontal gradients and absolute values derived for the plurality of relative pixel positions is derived as a first parameter. For each of the plurality of relative pixel positions, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient and absolute value, is derived. The sum of the plurality of vertical gradient and absolute values derived for the plurality of relative pixel positions is derived as a second parameter. For each of the plurality of relative pixel positions, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value, is derived. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the horizontal gradient sum, that is, the horizontal corresponding pixel difference value, is derived. The sum of the plurality of horizontal corresponding pixel difference values derived for the plurality of relative pixel positions is derived as a third parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding pixel difference value, is derived. The sum of the plurality of vertical corresponding pixel difference values derived for the plurality of relative pixel positions is derived as a fourth parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the vertical gradient sum, the positive or negative sign of the horizontal gradient sum is reversed or maintained, and the horizontal gradient sum whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding horizontal gradient sum, is derived. The sum of the plurality of vertical corresponding horizontal gradient sums derived for the plurality of relative pixel positions is derived as a fifth parameter.Generate a predicted image used in the decoding of the current block using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter.

[0205] Thus, in the operation for each pixel position, it is possible to reduce the substantial multiplication operation with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a predicted image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in decoding. In addition, it is possible to appropriately generate a predicted image based on a plurality of parameters including a parameter related to a horizontal gradient value, a parameter related to a vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[0206] In addition, for example, the circuit can derive the first parameter through Equation (11.1) described later, derive the second parameter through Equation (11.2) described later, derive the third parameter through Equation (11.3) described later, derive the fourth parameter through Equation (11.4) described later, and derive the fifth parameter through Equation (11.5) described later. Ω represents the set of the plurality of relative pixel positions, [i, j] represents each of the plurality of relative pixel positions, and for each of the plurality of relative pixel positions, I x 0 represents the horizontal gradient value of this relative pixel position in the first range, I x 1 represents the horizontal gradient value of this relative pixel position in the second range, I y 0 represents the vertical gradient value of this relative pixel position in the first range, I y 1 represents the vertical gradient value of this relative pixel position in the second range, I 0 represents the pixel value of this relative pixel position in the first range, I 1 represents the pixel value of this relative pixel position in the second range, abs(I x 1 +I x 0 ) represents I x 1 +I x 0 the absolute value of, sign(I x 1 +I x 0 ) represents I x 1 +I x 0 the positive or negative sign of, abs(I y 1 +I y0 ) represents I y 1 +I y 0 the absolute value of, sign(I y 1 +I y 0 ) represents I y 1 +I y 0 the positive or negative sign of.

[0207] Thus, it is possible to derive a plurality of parameters with low computational complexity using pixel values, horizontal gradient values, and vertical gradient values.

[0208] And, for example, the circuit derives a sixth parameter by dividing the third parameter by the first parameter, and derives a seventh parameter by subtracting the product of the fifth parameter and the sixth parameter from the fourth parameter and dividing by the second parameter, and generates the predicted image using the sixth parameter and the seventh parameter.

[0209] Thus, it is possible to appropriately aggregate a plurality of parameters into two parameters corresponding to the horizontal direction and the vertical direction. It is possible to appropriately incorporate parameters related to the horizontal gradient value into the parameter corresponding to the horizontal direction. It is possible to appropriately incorporate parameters related to the vertical gradient value, parameters related to both the horizontal gradient value and the vertical gradient value, and the parameter corresponding to the horizontal direction into the parameter corresponding to the vertical direction. Then, it is possible to generate a predicted image using these two parameters.

[0210] In addition, for example, the circuit derives the sixth parameter by the following formula (10.8), and derives the seventh parameter by the following formula (10.9), sG x represents the first parameter, sG y represents the second parameter, sG x dI represents the third parameter, sG y dI represents the fourth parameter, sG x G y represents the fifth parameter, u represents the sixth parameter, Bits is a function that returns the value obtained by rounding up the binary logarithm of the independent variable to an integer.

[0211] Thus, it is possible to derive two parameters corresponding to the horizontal direction and the vertical direction with low computational complexity.

[0212] In addition, for example, the circuit derives a predicted pixel value of the processing target pixel position included in the current block by using the first pixel value at the first pixel position corresponding to the processing target pixel position in the first reference block, the first horizontal gradient value at the first pixel position, the first vertical gradient value at the first pixel position, the second pixel value at the second pixel position corresponding to the processing target pixel position in the second reference block, the second horizontal gradient value at the second pixel position, the second vertical gradient value at the second pixel position, the sixth parameter, and the seventh parameter, and generates the predicted image.

[0213] Accordingly, it is possible to generate a predicted image by using two parameters corresponding to the horizontal direction and the vertical direction or the like, and the two parameters corresponding to the horizontal direction and the vertical direction can be appropriately reflected in the predicted image.

[0214] In addition, for example, the circuit derives the predicted pixel value by dividing the sum of the first pixel value, the second pixel value, the first correction value, and the second correction value by 2. The first correction value corresponds to the product of the difference between the first horizontal gradient value and the second horizontal gradient value and the sixth parameter, and the second correction value corresponds to the product of the difference between the first vertical gradient value and the second vertical gradient value and the seventh parameter.

[0215] Accordingly, it is possible to appropriately generate a predicted image by using two parameters corresponding to the horizontal direction and the vertical direction or the like.

[0216] In addition, for example, the circuit derives the predicted pixel value by Equation (10.10) described later. I 0 represents the first pixel value, I 1 represents the second pixel value, u represents the sixth parameter, I x 0 represents the first horizontal gradient value, I x 1 represents the second horizontal gradient value, v represents the seventh parameter, I y 0 represents the first vertical gradient value, I y 1 represents the second vertical gradient value.

[0217] Accordingly, it is possible to appropriately generate a predicted image according to an equation associated with two parameters corresponding to the horizontal direction and the vertical direction or the like.

[0218] In addition, for example, in an encoding method according to one aspect of the present invention, horizontal gradients and absolute values are respectively derived for a plurality of relative pixel positions. The plurality of relative pixel positions are a plurality of pixel positions that are commonly and relatively determined for both a first range of a first reference block including the current block and a second range of a second reference block including the current block, and are a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The sum of the plurality of horizontal gradients and absolute values respectively derived for the plurality of relative pixel positions is derived as a first parameter. For each of the plurality of relative pixel positions, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient and absolute value, is derived. The sum of the plurality of vertical gradients and absolute values respectively derived for the plurality of relative pixel positions is derived as a second parameter. For each of the plurality of relative pixel positions, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value, is derived. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the horizontal gradient sum, that is, the horizontal corresponding pixel difference value, is derived. The sum of the plurality of horizontal corresponding pixel difference values respectively derived for the plurality of relative pixel positions is derived as a third parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding pixel difference value, is derived. The sum of the plurality of vertical corresponding pixel difference values respectively derived for the plurality of relative pixel positions is derived as a fourth parameter. For each of the plurality of relative pixel positions, according to the positive or negative sign of the vertical gradient sum, the positive or negative sign of the horizontal gradient sum is reversed or maintained, and the horizontal gradient sum whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding horizontal gradient sum, is derived. The sum of the plurality of vertical corresponding horizontal gradient sums respectively derived for the plurality of relative pixel positions is derived as a fifth parameter. A predicted image used in the encoding of the current block is generated using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter.

[0219] Thus, in the operation for each pixel position, it is possible to reduce substantial multiplication operations with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a prediction image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in encoding. In addition, it is possible to appropriately generate a prediction image based on a plurality of parameters including a parameter related to a horizontal gradient value, a parameter related to a vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[0220] In addition, for example, in a decoding method according to one aspect of the present invention, horizontal gradients and absolute values are respectively derived for a plurality of relative pixel positions. The plurality of relative pixel positions are a plurality of pixel positions that are commonly and relatively determined for both a first range of a first reference block including the current block and a second range of a second reference block including the current block, and are a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The sum of the plurality of horizontal gradients and absolute values respectively derived for the plurality of relative pixel positions is derived as a first parameter. For each of the plurality of relative pixel positions, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient and absolute value, is derived. The sum of the plurality of vertical gradient and absolute values respectively derived for the plurality of relative pixel positions is derived as a second parameter. For each of the plurality of relative pixel positions, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value, is derived. For each of the plurality of relative pixel positions, based on the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the horizontal gradient sum, that is, the horizontal corresponding pixel difference value, is derived. The sum of the plurality of horizontal corresponding pixel difference values respectively derived for the plurality of relative pixel positions is derived as a third parameter. For each of the plurality of relative pixel positions, based on the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding pixel difference value, is derived. The sum of the plurality of vertical corresponding pixel difference values respectively derived for the plurality of relative pixel positions is derived as a fourth parameter. For each of the plurality of relative pixel positions, based on the positive or negative sign of the vertical gradient sum, the positive or negative sign of the horizontal gradient sum is reversed or maintained, and the horizontal gradient sum whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding horizontal gradient sum, is derived. The sum of the plurality of vertical corresponding horizontal gradient sums respectively derived for the plurality of relative pixel positions is derived as a fifth parameter. A predicted image used in decoding the current block is generated using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter.

[0221] Accordingly, in the operations performed for each pixel position, it is possible to reduce substantial multiplication operations with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a prediction image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in decoding. In addition, it is possible to appropriately generate a prediction image based on a plurality of parameters including a parameter related to a horizontal gradient value, a parameter related to a vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[0222] Alternatively, for example, an encoding device according to an aspect of the present invention includes a segmentation unit, an intra prediction unit, an inter prediction unit, a transformation unit, a quantization unit, and an entropy encoding unit.

[0223] The segmentation unit divides an encoding target picture constituting the moving image into a plurality of blocks. The intra prediction unit performs intra prediction for generating a prediction image of an encoding target block in the encoding target picture using a reference image in the encoding target picture. The inter prediction unit performs inter prediction for generating a prediction image of the encoding target block using a reference image in a reference picture different from the encoding target picture.

[0224] The transformation unit transforms a prediction error signal between the prediction image generated by the intra prediction unit or the inter prediction unit and the image of the encoding target block to generate a transformation coefficient signal of the encoding target block. The quantization unit quantizes the transformation coefficient signal. The entropy encoding unit encodes the quantized transformation coefficient signal.

[0225] In addition, for example, the inter-frame prediction unit derives a horizontal gradient and an absolute value for each of a plurality of relative pixel positions, the plurality of relative pixel positions being a plurality of pixel positions commonly and relatively determined for both a first range of a first reference block including the current block and a second range of a second reference block including the current block, and being a plurality of pixel positions in each of the first range and the second range, the horizontal gradient and the absolute value being the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, derives the sum of the plurality of horizontal gradients and absolute values derived for each of the plurality of relative pixel positions as a first parameter, for each of the plurality of relative pixel positions, derives the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, i.e., the vertical gradient and absolute value, and derives the sum of the plurality of vertical gradients and absolute values derived for each of the plurality of relative pixel positions as a second parameter, for each of the plurality of relative pixel positions, derives the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, i.e., the pixel difference value, for each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, i.e., the horizontal gradient sum, reverses or maintains the positive or negative sign of the pixel difference value derived for the relative pixel position, and derives the pixel difference value whose positive or negative sign has been reversed or maintained according to the positive or negative sign of the horizontal gradient sum, i.e., the horizontal corresponding pixel difference value, and derives the sum of the plurality of horizontal corresponding pixel difference values derived for each of the plurality of relative pixel positions as a third parameter, for each of the plurality of relative pixel positions, according to the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, i.e., the vertical gradient sum, reverses or maintains the positive or negative sign of the pixel difference value derived for the relative pixel position, and derives the pixel difference value whose positive or negative sign has been reversed or maintained according to the positive or negative sign of the vertical gradient sum, i.e., the vertical corresponding pixel difference value, and derives the sum of the plurality of vertical corresponding pixel difference values derived for each of the plurality of relative pixel positions as a fourth parameter, for each of the plurality of relative pixel positions, according to the positive or negative sign of the vertical gradient sum, reverses or maintains the positive or negative sign of the horizontal gradient sum, and derives the horizontal gradient sum whose positive or negative sign has been reversed or maintained according to the positive or negative sign of the vertical gradient sum, i.e., the vertical corresponding horizontal gradient sum, and derives the sum of the plurality of vertical corresponding horizontal gradient sums derived for each of the plurality of relative pixel positions as a fifth parameter, and generates a prediction image used in the encoding of the current block using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter.

[0226] Alternatively, for example, a decoding apparatus according to one embodiment of the present invention includes an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, and an addition unit (reconstruction unit).

[0227] The entropy decoding unit decodes the quantized transform coefficient signal of the decoding target block in the decoding target picture constituting the moving image. The inverse quantization unit inverse quantizes the quantized transform coefficient signal. The inverse transform unit inverse transforms the transform coefficient signal to obtain the prediction error signal of the decoding target block.

[0228] The intra prediction unit performs intra prediction for generating a prediction image of the decoding target block using a reference image in the decoding target picture. The inter prediction unit performs inter prediction for generating a prediction image of the decoding target block using a reference image in a reference picture different from the decoding target picture. The addition unit adds the prediction image generated by the intra prediction unit or the inter prediction unit and the prediction error signal to reconstruct the image of the decoding target block.

[0229] In addition, for example, the inter-frame prediction unit derives a horizontal gradient and an absolute value for each of a plurality of relative pixel positions, the plurality of relative pixel positions being a plurality of pixel positions that are commonly and relatively determined for both a first range of a first reference block including the current block and a second range of a second reference block including the current block, and being a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The sum of the plurality of horizontal gradients and absolute values derived for each of the plurality of relative pixel positions is derived as a first parameter. For each of the plurality of relative pixel positions, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient and absolute value, is derived. The sum of the plurality of vertical gradients and absolute values derived for each of the plurality of relative pixel positions is derived as a second parameter. For each of the plurality of relative pixel positions, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value, is derived. For each of the plurality of relative pixel positions, based on the positive or negative sign of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the horizontal gradient sum, that is, the horizontal corresponding pixel difference value, is derived. The sum of the plurality of horizontal corresponding pixel difference values derived for each of the plurality of relative pixel positions is derived as a third parameter. For each of the plurality of relative pixel positions, based on the positive or negative sign of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum, the positive or negative sign of the pixel difference value derived for the relative pixel position is reversed or maintained, and the pixel difference value whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding pixel difference value, is derived. The sum of the plurality of vertical corresponding pixel difference values derived for each of the plurality of relative pixel positions is derived as a fourth parameter. For each of the plurality of relative pixel positions, based on the positive or negative sign of the vertical gradient sum, the positive or negative sign of the horizontal gradient sum is reversed or maintained, and the horizontal gradient sum whose positive or negative sign is reversed or maintained according to the positive or negative sign of the vertical gradient sum, that is, the vertical corresponding horizontal gradient sum, is derived. The sum of the plurality of vertical corresponding horizontal gradient sums derived for each of the plurality of relative pixel positions is derived as a fifth parameter. The first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter are used to generate a prediction image used in decoding the current block.

[0230] Furthermore, these inclusive or specific forms can also be implemented by a system, apparatus, method, integrated circuit, computer program, or non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.

[0231] Hereinafter, embodiments will be specifically described with reference to the accompanying drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and are not intended to limit the claims.

[0232] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processing and / or structure described in each form of the present invention. The processing and / or structure can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processing and / or structure applied to the embodiments, for example, any of the following can be performed.

[0233] (1) Among a plurality of constituent elements of the encoding device or decoding device of the embodiment described in each form of the present invention, a certain one can be replaced with another constituent element described in a certain one of the forms of the present invention, or they can be combined;

[0234] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, and deletion of functions or processing performed on a part of the constituent elements of the encoding device or decoding device can also be made. For example, any function or processing can be replaced with another function or processing described in a certain one of the forms of the present invention, or they can be combined;

[0235] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, and deletion can also be made to a part of the plurality of processes included in the method. For example, any process in the method can be replaced with another process described in a certain one of the forms of the present invention, or they can be combined;

[0236] (4) A part of the constituent elements of the encoding device or decoding device constituting the embodiment can be combined with the constituent elements described in a certain one of the forms of the present invention, or can be combined with the constituent elements having a part of the functions described in a certain one of the forms of the present invention, or can be combined with the constituent elements that perform a part of the processing performed by the constituent elements described in the forms of the present invention;

[0237] (5) A component that is part of the function of the encoding device or decoding device of the embodiment, or a component that processes part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced by a component described in one of the various aspects of the present invention, a component that is part of the function described in one of the various aspects of the present invention, or a component that processes part of the processing described in one of the various aspects of the present invention;

[0238] (6) In the method implemented by the encoding device or decoding device of the embodiment, one of the multiple processes included in the method is replaced by a process described in one of the various aspects of the present invention or a similar process, or they are combined;

[0239] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various aspects of the present invention.

[0240] (8) The manner of implementing the processes and / or structures described in the various aspects of the present invention is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the motion image encoding or motion image decoding disclosed in the embodiment.

[0241] [Encoding Device]

[0242] First, the encoding device of the embodiment will be described. Figure 1 It is a block diagram showing the functional structure of the encoding device 100 of the embodiment. The encoding device 100 is a motion image encoding device that encodes motion images in units of blocks.

[0243] As Figure 1 shown, the encoding device 100 is a device that encodes images in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0244] The encoding device 100 is implemented by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the splitting unit 102, subtraction unit 104, transformation unit 106, quantization unit 108, entropy encoding unit 110, inverse quantization unit 112, inverse transformation unit 114, addition unit 116, loop filtering unit 120, intra prediction unit 124, inter prediction unit 126, and prediction control unit 128. In addition, the encoding device 100 may also be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, subtraction unit 104, transformation unit 106, quantization unit 108, entropy encoding unit 110, inverse quantization unit 112, inverse transformation unit 114, addition unit 116, loop filtering unit 120, intra prediction unit 124, inter prediction unit 126, and prediction control unit 128.

[0245] Hereinafter, after explaining the overall processing flow of the encoding device 100, each component included in the encoding device 100 will be described.

[0246] [Overall Flow of Encoding Processing]

[0247] Figure 2 It is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0248] First, the splitting unit 102 of the encoding device 100 splits each picture included in the input image, which is a moving image, into a plurality of blocks of a fixed size (for example, 128×128 pixels) (step Sa_1). Then, the splitting unit 102 selects a splitting pattern (also referred to as a block shape) for the block of the fixed size (step Sa_2). That is, the splitting unit 102 further splits the block with a fixed size into a plurality of blocks constituting the selected splitting pattern. Then, for each of the plurality of blocks, the encoding device 100 performs the processing of steps Sa_3 to Sa_9 on the block (i.e., the encoding target block).

[0249] That is, the prediction processing unit constituted by all or a part of the intra prediction unit 124, inter prediction unit 126, and prediction control unit 128 generates a prediction signal (also referred to as a prediction block) of the encoding target block (also referred to as the current block) (step Sa_3).

[0250] Next, the subtraction unit 104 generates the difference between the encoding target block and the prediction block as a prediction residual (also referred to as a difference block) (step Sa_4).

[0251] Next, the transformation unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing transformation and quantization on the difference block (step Sa_5). In addition, the block composed of a plurality of quantization coefficients is also referred to as a coefficient block.

[0252] Next, the entropy encoding unit 110 generates an encoded signal (step Sa_6) by encoding (specifically, entropy encoding) the coefficient block and the prediction parameters related to the generation of the prediction signal. Additionally, the encoded signal is also referred to as an encoded bitstream, a compressed bitstream, or a stream.

[0253] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore a plurality of prediction residuals (i.e., difference blocks) by performing inverse quantization and inverse transform on the coefficient block (step Sa_7).

[0254] Next, the addition unit 116 reconstructs the current block into a reconstructed image (also referred to as a reconstructed block or a decoded image block) by adding the restored difference block to the prediction block (step Sa_8). Thereby, a reconstructed image is generated.

[0255] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).

[0256] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding is not completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.

[0257] In addition, in the above example, the encoding device 100 selects one segmentation pattern for blocks of a fixed size and encodes each block according to the segmentation pattern. However, each block may also be encoded according to each of multiple segmentation patterns. In this case, the encoding device 100 may evaluate the cost for each of the multiple segmentation patterns, and for example, may select the encoded signal obtained by encoding according to the segmentation pattern with the minimum cost as the output encoded signal.

[0258] As shown in the figure, the processing of these steps Sa_1 to Sa_10 is sequentially performed by the encoding device 100. Alternatively, a plurality of parts of these processes may be performed in parallel, or the order of these processes may be swapped.

[0259] [Segmentation Unit]

[0260] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128). Other fixed block sizes may also be used. Such blocks of a fixed size are sometimes referred to as coding tree units (CTUs). Further, the splitting unit 102 splits each block of the fixed size into blocks of a variable size (e.g., 64×64 or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects a splitting pattern. Such blocks of a variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks within the picture may be used as the processing units for CUs, PUs, and TUs.

[0261] Figure 3 It is a conceptual diagram showing an example of block splitting in the embodiment. In Figure 3 the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.

[0262] Here, block 10 is a square block of 128×128 pixels (128×128 block). This 128×128 block 10 is first split into four square 64×64 blocks (quadtree block splitting).

[0263] The upper-left 64×64 block is further vertically split into two rectangular 32×64 blocks, and the left 32×64 block is further vertically split into two rectangular 16×64 blocks (binary tree block splitting). As a result, the upper-left 64×64 block is split into two 16×64 blocks 11, 12 and a 32×64 block 13.

[0264] The upper-right 64×64 block is horizontally split into two rectangular 64×32 blocks 14, 15 (binary tree block splitting).

[0265] The lower-left 64×64 block is split into four square 32×32 blocks (quadtree block splitting). The upper-left block and the lower-right block among the four 32×32 blocks are further split. The upper-left 32×32 block is vertically split into two rectangular 16×32 blocks, and the right 16×32 block is further horizontally split into two 16×16 blocks (binary tree block splitting). The lower-right 32×32 block is horizontally split into two 32×16 blocks (binary tree block splitting). As a result, the lower-left 64×64 block is split into a 16×32 block 16, two 16×16 blocks 17, 18, two 32×32 blocks 19, 20, and two 32×16 blocks 21, 22.

[0266] The 64×64 block 23 in the lower right is not divided.

[0267] As described above, in Figure 3 , the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes referred to as QTBT (quad - tree plus binary tree) partitioning.

[0268] In addition, in Figure 3 , one block is divided into 4 or 2 blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these. For example, one block can also be divided into 3 blocks (ternary tree partitioning). Partitioning including such ternary tree partitioning is sometimes referred to as MBT (multi type tree) partitioning.

[0269] [Structural Slices / Tiles of the Picture]

[0270] To decode a picture in parallel, the picture is sometimes constructed in units of slices or tiles. The picture constructed in units of slices or tiles can be formed by the partitioning unit 102.

[0271] A slice is the basic encoding unit that makes up a picture. A picture consists of, for example, one or more slices. In addition, a slice consists of one or more consecutive CTUs (Coding Tree Units).

[0272] Figure 4A is a conceptual diagram showing an example of the structure of a slice. For example, a picture includes 11×8 CTUs and is divided into 4 slices (slice 1 to slice 4). Slice 1 consists of 16 CTUs, slice 2 consists of 21 CTUs, slice 3 consists of 29 CTUs, and slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to any one of the slices. The shape of the slice becomes the shape of dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can be any position among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. In addition, the slice contains header information and encoded data. In the header information, the characteristics of the slice, such as the address of the first CTU of the slice and the slice type, can also be described.

[0273] A tile is the unit of a rectangular area that makes up a picture. Numbers called TileId can also be assigned to each tile in raster scan order.

[0274] Figure 4BIt is a conceptual diagram showing an example of the structure of a tile. For example, the picture includes 11×8 CTUs and is divided into tiles (tiles 1 to 4) in 4 rectangular regions. When using tiles, the processing order of CTUs is changed compared to the case of not using tiles. When not using tiles, multiple CTUs within the picture are processed in raster scan order. When using tiles, in each of the multiple tiles, at least 1 CTU is processed in raster scan order. For example, as Figure 4B shown, the processing order of the multiple CTUs included in tile 1 is from the left end of the first row of tile 1 to the right end of the first row of tile 1, and then from the left end of the second row of tile 1 to the right end of the second row of tile 1.

[0275] In addition, one tile sometimes contains more than one slice, and one slice sometimes contains more than one tile.

[0276] [Subtraction unit]

[0277] The subtraction unit 104 subtracts the predicted signal (the predicted sample input from the prediction control unit 128 shown below) from the original signal (original sample) in block units input from and segmented by the segmentation unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). And the subtraction unit 104 outputs the calculated prediction error (residual) to the transformation unit 106.

[0278] The original signal is the input signal of the encoding device 100 and is a signal representing the images of each picture constituting the moving image (for example, a luminance (luma) signal and two chrominance (chroma) signals). Hereinafter, there are cases where the signal representing the image is also called a sample.

[0279] [Transformation unit]

[0280] The transformation unit 106 transforms the prediction error in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a prescribed discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain. The prescribed DCT or DST can also be determined in advance.

[0281] In addition, the transformation unit 106 can also adaptively select a transformation type from among multiple transformation types and use a transformation basis function corresponding to the selected transformation type to transform the prediction error into transformation coefficients. Such a transformation is called EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform) in some cases.

[0282] Multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table representing transform basis functions corresponding to the transform type examples. In Figure 5A where N represents the number of input pixels. The selection of the transform type from among these multiple transform types can depend, for example, on the type of prediction (intra prediction and inter prediction) or on the intra prediction mode.

[0283] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., bit sequence level, picture level, slice level, tile level, or CTU level).

[0284] In addition, the transform unit 106 can also perform a re-transformation on the transform coefficients (transformation results). Such a re-transformation can be a case called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (e.g., 4×4 sub-block) included in the block of transform coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are typically signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0285] In the transform unit 106, a separable transform and a non-separable transform can also be applied. The separable transform is a method of performing multiple separations in each direction according to the number of input dimensions, and the non-separable transform is a method of treating two or more dimensions as one dimension and performing a transform together when the input is multi-dimensional.

[0286] For example, as an example of the non-separable transform, when the input is a 4×4 block, it can be regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.

[0287] In addition, in a further example of the Non-Separable transform, after regarding a 4×4 input block as a permutation having 16 elements, a transform (Hypercube Givens Transform) that performs Givens rotations on the permutation multiple times may be performed.

[0288] In the transform in the transform unit 106, the type of basis to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform). In SVT, as Figure 5B shown, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The type of transform basis can be set for each region. For example, DST7 and DCT8 are used. In this example, only one of the two regions within the CU is transformed, and the other is not transformed. However, both regions can also be transformed. In addition, the splitting method is not limited to bisection and can be more flexible. For example, it can be quartered or the information indicating the splitting is additionally encoded and signaled in the same way as CU splitting. In addition, SVT is sometimes also referred to as SBT (Sub-block Transform).

[0289] [Quantization Unit]

[0290] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a prescribed scan order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112. The prescribed scan order may be determined in advance.

[0291] The prescribed scan order is the order used for quantization / inverse quantization of the transform coefficients. For example, the prescribed scan order may be defined in ascending order of frequency (order from low frequency to high frequency) or descending order of frequency (order from high frequency to low frequency).

[0292] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0293] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing the values sampled at a specified interval by associating them with specified levels. In this technical field, other expressions such as rounding, truncating, and scaling can also be used for reference, and rounding, truncating, and scaling can also be adopted. The specified interval and levels can also be determined in advance.

[0294] As methods of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device side and a method of using a default quantization matrix (default matrix). On the encoding device side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix.

[0295] On the other hand, there is also a method of quantizing in such a way that the coefficients of the high-frequency components and the coefficients of the low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix (flat matrix) in which all the coefficients are the same value.

[0296] The quantization matrix can be specified by, for example, SPS (Sequence Parameter Set) or PPS (Picture Parameter Set). SPS contains the parameters used for a sequence, and PPS contains the parameters used for a picture. SPS and PPS are sometimes simply referred to as parameter sets.

[0297] [Entropy Encoding Unit]

[0298] The entropy encoding unit 110 generates an encoded signal (encoded bitstream) based on the quantized coefficients input from the quantization unit 108. Specifically, the entropy encoding unit 110, for example, binarizes the quantized coefficients, performs arithmetic coding on the binary signal, and outputs a compressed bitstream or sequence.

[0299] [Inverse Quantization Unit]

[0300] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a specified scan order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114. The specified scan order can also be determined in advance.

[0301] [Inverse Transform Unit]

[0302] The inverse transform unit 114 restores the prediction error (residual) by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0303] In addition, since information is usually lost due to quantization in the restored prediction error, it is inconsistent with the prediction error calculated by the subtraction unit 104. That is, the restored prediction error usually contains quantization error.

[0304] [Addition unit]

[0305] The addition unit 116 reconstructs the current block by adding the prediction error input from the inverse transform unit 114 and the prediction sample input from the prediction control unit 128. Then, the addition unit 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes referred to as a local decoded block.

[0306] [Block memory]

[0307] The block memory 118 is, for example, a storage unit for storing blocks within the coded object picture (referred to as the current picture) that are referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the addition unit 116.

[0308] [Frame memory]

[0309] The frame memory 122 is, for example, a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0310] [Loop filter unit]

[0311] The loop filter unit 120 applies loop filtering to the block reconstructed by the addition unit 116 and outputs the filtered reconstructed block to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, deblocking filtering (DF or DBF), sample adaptive offset (SAO), and adaptive loop filtering (ALF).

[0312] In ALF, a least squares error filter for removing coding distortion is adopted. For example, for each 2×2 sub-block within the current block, one filter selected from multiple filters based on the direction and activity of the locality-based gradient is adopted.

[0313] Specifically, first, sub-blocks (e.g., 2×2 sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is performed based on the direction and activity of the gradient. For example, using the direction value D of the gradient (e.g., 0 to 2 or 0 to 4) and the activity value A of the gradient (e.g., 0 to 4), the classification value C is calculated (e.g., C = 5D + A). And based on the classification value C, the sub-blocks are classified into multiple classes.

[0314] The direction value D of the gradient is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients in multiple directions and quantifying the added result.

[0315] Based on the result of such classification, a filter for the sub-block is determined from among multiple filters.

[0316] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 6A to 6C It is a diagram showing multiple examples of the shape of the filter used in the ALF. Figure 6A It represents a 5×5 diamond-shaped filter, Figure 6B It represents a 7×7 diamond-shaped filter, Figure 6C It represents a 9×9 diamond-shaped filter. The information indicating the shape of the filter is usually signaled at the picture level. In addition, the signaling of the information indicating the shape of the filter does not need to be limited to the picture level and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0317] The on / off of the ALF can also be determined, for example, at the picture level or the CU level. For example, regarding luminance, it can be determined at the CU level whether to adopt the ALF, and regarding chrominance difference, it can be determined at the picture level whether to adopt the ALF. The information indicating the on / off of the ALF is usually signaled at the picture level or the CU level. In addition, the signaling of the information indicating the on / off of the ALF does not need to be limited to the picture level or the CU level and can also be other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0318] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are usually signaled at the picture level. In addition, the signaling of the coefficient sets does not need to be limited to the picture level and can also be other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0319] [Loop Filtering Section > Deblocking Filter]

[0320] In the deblocking filter, the loop filtering section 120 reduces the distortion generated at the block boundary by performing filtering processing on the block boundary of the reconstructed image.

[0321] Figure 7 It is a block diagram showing an example of the detailed structure of the loop filter unit 120 that functions as a deblocking filter.

[0322] The loop filter unit 120 includes a boundary determination unit 1201, a filtering determination unit 1203, a filtering processing unit 1205, a processing determination unit 1208, a filtering characteristic determination unit 1207, and switches 1202, 1204, and 1206.

[0323] The boundary determination unit 1201 determines whether there are pixels (i.e., target pixels) for which deblocking filtering processing is to be performed near the block boundary. Then, the boundary determination unit 1201 outputs the determination result to the switch 1202 and the processing determination unit 1208.

[0324] When it is determined by the boundary determination unit 1201 that target pixels exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1204. On the contrary, when it is determined by the boundary determination unit 1201 that target pixels do not exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1206.

[0325] The filtering determination unit 1203 determines whether to perform deblocking filtering processing on the target pixels based on the pixel values of at least one neighboring pixel located around the target pixels. Then, the filtering determination unit 1203 outputs the determination result to the switch 1204 and the processing determination unit 1208.

[0326] When it is determined by the filtering determination unit 1203 that deblocking filtering processing is to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the filtering processing unit 1205. On the contrary, when it is determined by the filtering determination unit 1203 that deblocking filtering processing is not to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the switch 1206.

[0327] When the image before filtering processing is obtained via the switches 1202 and 1204, the filtering processing unit 1205 performs deblocking filtering processing with the filtering characteristics determined by the filtering characteristic determination unit 1207 on the target pixels. Then, the filtering processing unit 1205 outputs the pixels after the filtering processing to the switch 1206.

[0328] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs the pixels that have not been subjected to deblocking filtering processing and the pixels that have been subjected to deblocking filtering processing by the filtering processing unit 1205.

[0329] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering processing is to be performed on the target pixel, the processing determination unit 1208 outputs the pixel after the deblocking filtering processing from the switch 1206. In addition, in cases other than the above, the processing determination unit 1208 outputs the pixel that has not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206.

[0330] Figure 8 It is a conceptual diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to the block boundary.

[0331] In the deblocking filtering process, for example, using the pixel value and the quantization parameter, either one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Figure 8 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the respective pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing operations shown in the following equations, for example.

[0332] q'0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0333] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0334] q'2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0335] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the respective pixels p0 to p2 and pixels q0 to q2. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering process are filtering coefficients.

[0336] Furthermore, in the deblocking filtering process, clipping processing may also be performed in such a way that the pixel value after the operation is not set to exceed the threshold value. In this clipping processing, using the threshold value determined according to the quantization parameter, the pixel value after the operation based on the above equation is clipped to "operation target pixel value ± 2×threshold value". Thereby, excessive smoothing can be prevented.

[0337] Figure 9 It is a conceptual diagram for explaining the block boundary where the deblocking filtering process is performed. Figure 10 It is a conceptual diagram showing an example of the Bs value.

[0338] The block boundaries for performing deblocking filtering are, for example, Figure 9 the boundaries of the PU (Prediction Unit) or TU (Transform Unit) of the 8×8 pixel block shown. The deblocking filtering process can be performed in units of 4 rows or 4 columns. First, for Figure 9 the blocks P and Q shown, as Figure 10 such, the Bs (Boundary Strength) value is determined.

[0339] According to Figure 10 the Bs value, it is determined whether to perform deblocking filtering with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filtering for the chrominance signal is performed. When the Bs value is 1 or more and a specified condition is satisfied, deblocking filtering for the luminance signal is performed. The specified condition can also be determined in advance. In addition, the determination condition of the Bs value is not limited to Figure 10 the conditions shown, and can also be determined based on other parameters.

[0340] [Prediction processing unit (intra prediction unit / inter prediction unit / prediction control unit)]

[0341] Figure 11 is a flowchart showing an example of the processing performed by the prediction processing unit of the encoding device 100. In addition, the prediction processing unit is composed of all or part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0342] The prediction processing unit generates a prediction image of the current block (step Sb_1). This prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there is an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit uses the reconstructed image that has been obtained by generating a prediction block, a differential block, a coefficient block, restoring the differential block, and generating a decoded image block, to generate a prediction image of the current block.

[0343] The reconstructed image can be, for example, an image of a reference picture, or an image of an encoded block within the current picture that includes the current block. The encoded blocks within the current picture are, for example, adjacent blocks of the current block.

[0344] Figure 12 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0345] The prediction processing unit generates a prediction image by a first method (step Sc_1a), generates a prediction image by a second method (step Sc_1b), and generates a prediction image by a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and may be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such a prediction method, the above-mentioned reconstructed image may also be used.

[0346] Next, the prediction processing unit selects any one of the plurality of prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). The selection of this prediction image, that is, the selection of the method or mode for obtaining the final prediction image, may also calculate the cost for each generated prediction image and be based on this cost. In addition, the selection of this prediction image may be performed based on the parameters for encoding processing. The encoding device 100 may signal the information for determining the selected prediction image, method, or mode as an encoding signal (also referred to as an encoding bitstream). This information may be, for example, a flag or the like. Thus, the decoding device can generate a prediction image based on this information in the manner or mode selected in the encoding device 100. In addition, in Figure 12 the example shown, after the prediction processing unit generates a prediction image by each method, it selects any one of the prediction images. However, before generating these prediction images, the prediction processing unit may select a method or mode based on the parameters for the above-mentioned encoding processing, and may generate a prediction image according to this method or mode.

[0347] For example, the first method and the second method are intra-frame prediction and inter-frame prediction, respectively, and the prediction processing unit may select the final prediction image for the current block from the prediction images generated according to these prediction methods.

[0348] Figure 13 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0349] First, the prediction processing unit generates a prediction image by intra-frame prediction (step Sd_1a), and generates a prediction image by inter-frame prediction (step Sd_1b). In addition, the prediction image generated by intra-frame prediction is also referred to as an intra-frame prediction image, and the prediction image generated by inter-frame prediction is also referred to as an inter-frame prediction image.

[0350] Next, the prediction processing unit evaluates each of the intra-predicted image and the inter-predicted image (step Sd_2). Cost can also be used in this evaluation. That is, the prediction processing unit calculates the cost C for each of the intra-predicted image and the inter-predicted image. This cost C can be calculated by an equation of the R-D optimization model, such as C = D + λ × R. In this equation, D is the coding distortion of the predicted image, and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. In addition, R is the generated coding amount of the predicted image. Specifically, it is the coding amount required for coding motion information, etc. used to generate the predicted image. In addition, λ is, for example, the undetermined multiplier of Lagrange.

[0351] Then, the prediction processing unit selects, as the final predicted image of the current block, the predicted image that calculates the minimum cost C from the intra-predicted image and the inter-predicted image (step Sd_3). That is, the prediction method or mode used to generate the predicted image of the current block is selected.

[0352] [Intra-Prediction Unit]

[0353] The intra-prediction unit 124 performs intra-prediction (also called intra-frame prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction signal (intra-prediction signal). Specifically, the intra-prediction unit 124 generates an intra-prediction signal by performing intra-prediction with reference to the samples (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 128.

[0354] For example, the intra-prediction unit 124 performs intra-prediction using one of a plurality of prescribed intra-prediction modes. The plurality of intra-prediction modes usually include one or more non-directional prediction modes and a plurality of directional prediction modes. The prescribed plurality of modes can also be determined in advance.

[0355] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode specified by the H.265 / HEVC standard.

[0356] The plurality of directional prediction modes include, for example, 33-direction prediction modes specified by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes can also include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Figure 14 is a conceptual diagram showing all 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra-prediction. The solid arrows indicate 33 directions specified by the H.265 / HEVC standard, and the dashed arrows indicate the additional 32 directions (the 2 non-directional prediction modes are not Figure 14 shown in the figure).

[0357] In various processing examples, in the intra prediction of a chrominance block, a luminance block may also be referred to. That is, the chrominance component of the current block may also be predicted based on the luminance component of the current block. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block in this way (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0358] The intra prediction unit 124 may also correct the intra-predicted pixel value based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes called PDPC (position dependent intraprediction combination). Information indicating whether PDPC is used (for example, called the PDPC flag) is generally signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and may also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0359] [Inter prediction unit]

[0360] The inter prediction unit 126 performs inter prediction (also called inter-picture prediction) of the current block by referring to a reference picture different from the current picture stored in the frame memory 122, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or the current sub-block (for example, 4×4 block) within the current block. For example, the inter prediction unit 126 performs a motion search (motion estimation) within the reference picture for the current block or the current sub-block to find the reference block or sub-block that most matches the current block or the current sub-block. And, the inter prediction unit 126 obtains motion information (for example, a motion vector) that compensates for the motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction unit 126 performs motion compensation (or motion prediction) based on this motion information, thereby generating an inter prediction signal for the current block or sub-block. And, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0361] The motion information used in motion compensation is signaled as an inter prediction signal in various forms. For example, a motion vector may also be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) may also be signaled.

[0362] [Basic process of inter prediction]

[0363] Figure 15 It is a flowchart showing an example of the basic process of inter-frame prediction.

[0364] First, the inter-frame prediction unit 126 generates a prediction image (Steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (Step Se_4).

[0365] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by determining the motion vector (MV) of the current block (Steps Se_1 and Se_2) and performing motion compensation (Step Se_3). In addition, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by selecting a candidate motion vector (candidate MV) (Step Se_1) and deriving the MV (Step Se_2). The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. In addition, in the derivation of the MV, the inter-frame prediction unit 126 may further select at least one candidate MV from at least one candidate MV and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching for regions of the reference picture indicated by the candidate MV for each of the selected at least one candidate MVs. In addition, the action of searching for regions of the reference picture may also be referred to as motion search.

[0366] In addition, in the above example, Steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126. However, for example, the processing of Step Se_1 or Step Se_2 may also be performed by other components included in the encoding device 100.

[0367] [Flow of Derivation of Motion Vector]

[0368] Figure 16 It is a flowchart showing an example of the derivation of a motion vector.

[0369] The inter-frame prediction unit 126 derives the MV of the current block in a mode in which motion information (e.g., MV) is encoded. In this case, for example, the motion information is encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the encoded signal (also referred to as the encoded bitstream).

[0370] Alternatively, the inter-frame prediction unit 126 derives the MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the encoded signal.

[0371] Here, the modes for MV derivation may also include the following common inter-frame modes, merge modes, FRUC modes, affine modes, etc. Among these modes, the modes for encoding motion information include common inter-frame modes, merge modes, and affine modes (specifically, affine inter-frame modes and affine merge modes), etc. In addition, the motion information may include not only MVs but also the predicted motion vector selection information described later. In addition, the modes that do not encode motion information include FRUC modes, etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0372] Figure 17 It is a flowchart showing another example of motion vector derivation.

[0373] The inter-frame prediction unit 126 derives the MV of the current block in the mode of encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0374] Alternatively, the inter-frame prediction unit 126 derives the MV in the mode of not encoding the differential MV. In this case, the encoded differential MV is not included in the encoded signal.

[0375] Here, as described above, the modes for MV derivation include the following common inter-frame modes, merge modes, FRUC modes, affine modes, etc. Among these modes, the modes for encoding the differential MV include common inter-frame modes and affine modes (specifically, affine inter-frame modes), etc. In addition, the modes that do not encode the differential MV include FRUC modes, merge modes, and affine modes (specifically, affine merge modes), etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.

[0376] [Flow of Motion Vector Derivation]

[0377] Figure 18It is a flowchart showing another example of motion vector derivation. There are multiple modes for MV derivation, i.e., inter-frame prediction modes, which are roughly classified into a mode for encoding differential MVs and a mode for not encoding differential motion vectors. The modes for not encoding differential MVs include the merge mode, the FRUC mode, and the affine mode (specifically, the affine merge mode). The detailed content of these modes will be described later. Briefly, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from the encoded blocks in the vicinity, the FRUC mode is a mode in which the MV of the current block is derived by searching between the encoded regions, and the affine mode is a mode in which the motion vectors of the respective sub-blocks constituting the current block are assumed to be an affine transformation and are derived as the MV of the current block.

[0378] Specifically, as shown in the figure, when the inter-frame prediction mode information indicates 0 (0 in Sf_1), the inter-frame prediction unit 126 derives a motion vector based on the merge mode (Sf_2). Further, when the inter-frame prediction mode information indicates 1 (1 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the FRUC mode (Sf_3). Further, when the inter-frame prediction mode information indicates 2 (2 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the affine mode (specifically, the affine merge mode) (Sf_4). Further, when the inter-frame prediction mode information indicates 3 (3 in Sf_1), the inter-frame prediction unit 126 derives a motion vector according to the mode for encoding differential MVs (e.g., the normal inter-frame mode) (Sf_5).

[0379] [MV Derivation > Normal Inter-Frame Mode]

[0380] The normal inter-frame mode is an inter-frame prediction mode in which the MV of the current block is derived based on a block similar to the image of the current block in the region of the reference picture represented by the candidate MV. Further, in this normal inter-frame mode, the differential MV is encoded.

[0381] Figure 19 It is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.

[0382] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially around the current block (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0383] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order (step Sg_2). Additionally, this priority order may be predetermined for each of the N candidate MVs.

[0384] Next, the inter-frame prediction unit 126 selects one prediction motion vector candidate from the N prediction motion vector candidates as the prediction motion vector (also referred to as prediction MV) of the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes prediction motion vector selection information for identifying the selected prediction motion vector into the stream. Additionally, the stream is the above-mentioned encoded signal or encoded bitstream.

[0385] Next, the inter-frame prediction unit 126 derives the MV of the current block with reference to the encoded reference picture (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction motion vector as the differential MV into the stream. Additionally, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0386] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). Additionally, the predicted image is the above-mentioned inter-frame prediction signal.

[0387] Furthermore, information indicating the inter-frame prediction mode (in the above example, the ordinary inter-frame mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, prediction parameters.

[0388] Additionally, the candidate MV list may also be used in common with the list used in other modes. Furthermore, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to this candidate MV list is, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs, etc.

[0389] [MV Derivation > Merge Mode]

[0390] The merge mode is an inter-frame prediction mode in which the MV of the current block is derived by selecting a candidate MV from the candidate MV list.

[0391] Figure 20 It is a flowchart showing an example of inter-frame prediction based on the merge mode.

[0392] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0393] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream.

[0394] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).

[0395] In addition, information indicating the inter-frame prediction mode (in the above example, the merge mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.

[0396] Figure 21 It is a conceptual diagram for explaining an example of the motion vector derivation process of the current picture based on the merge mode.

[0397] First, a prediction MV list registering candidates for the prediction MV is generated. As candidates for the prediction MV, there are: a spatially adjacent prediction MV, which is the MV possessed by a plurality of encoded blocks located in the spatial vicinity of the target block; a temporally adjacent prediction MV, which is the MV possessed by a nearby block obtained by projecting the position of the target block in the encoded reference picture; a combined prediction MV, which is an MV generated by combining the MV values of the spatially adjacent prediction MV and the temporally adjacent prediction MV; and a zero prediction MV, which is an MV with a value of zero, etc.

[0398] Next, the MV for the target block is determined by selecting one prediction MV from the plurality of prediction MVs registered in the prediction MV list.

[0399] Moreover, in the variable-length coding unit, the signal indicating which prediction MV is selected, that is, merge_idx, is described in the stream and encoded.

[0400] In addition, Figure 21 The prediction MVs registered in the prediction MV list described in are an example, and the number may be different from that in the figure, or it may be a structure that does not include some types of the prediction MVs in the figure, or a structure with additional prediction MVs other than the types of prediction MVs in the figure.

[0401] It is also possible to use the MV of the object block exported in the merge mode, and determine the final MV by performing the following DMVR (decode motion vector refinement) process.

[0402] In addition, the candidate for the predicted MV is the above-mentioned candidate MV, and the predicted MV list is the above-mentioned candidate MV list. In addition, the candidate MV list may also be referred to as the candidate list. In addition, merge_idx is MV selection information.

[0403] [MV Export>FRUC Mode]

[0404] The motion information may also be derived on the decoder side instead of being signaled from the encoder side. In addition, as described above, the merge mode defined by the H.265 / HEVC standard may also be used. In addition, for example, the motion information may be derived by performing a motion search on the decoder side. In the embodiment, the motion search is performed on the decoder side without using the pixel values of the current block.

[0405] Here, the mode of performing motion estimation on the decoder side will be described. The mode of performing motion estimation on the decoder side may be a mode called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0406] In the form of a flowchart Figure 22This represents an example of FRUC processing. First, referring to the motion vectors of the encoded blocks adjacent to the current block in space or time, a plurality of candidates each having a predicted motion vector (MV) are generated (i.e., it is a candidate MV list, and it can also be shared with the merge list) (step Si_1). Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected based on the evaluation value. And, based on the motion vector of the selected candidate, a motion vector for the current block is derived (step Si_4). Specifically, for example, the motion vector of the selected candidate (the best candidate MV) is directly derived as the motion vector for the current block. In addition, for example, the motion vector for the current block can also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate. That is, the peripheral region of the best candidate MV can be searched by using pattern matching and the evaluation value in the reference picture. In the case where there is an MV with a better evaluation value, the best candidate MV is updated to the above MV, and it is used as the final MV of the current block. It is also possible to configure it so that the process of updating to an MV with a better evaluation value is not performed.

[0407] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block by using the derived MV and the encoded reference picture (step Si_5).

[0408] The exact same process can also be performed when processing is done in sub-block units.

[0409] The evaluation value can also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the motion vector is compared with the reconstructed image of a specified region (for example, as shown below, this region can be a region of another reference picture or a neighboring block of the current picture). The specified region can also be determined in advance.

[0410] Then, the difference in pixel values of the two reconstructed images can be calculated for the evaluation value of the motion vector. In addition, it can also be that other information is used in addition to the difference value to calculate the evaluation value.

[0411] Next, an example of pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (for example, the merge list) is selected as the starting point for the search based on pattern matching. For example, as the pattern matching, the first pattern matching or the second pattern matching can be used. In some cases, the first pattern matching and the second pattern matching are respectively called bilateral matching and template matching.

[0412] [MV Derivation>FRUC>Two-way Matching]

[0413] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified area for calculating the evaluation value for candidates as described above, an area in another reference picture along the motion trajectory of the current block is used. The specified area may also be determined in advance.

[0414] Figure 23 It is a conceptual diagram showing an example of the first pattern matching (two-way matching) between two blocks in two reference pictures along the motion trajectory. As Figure 23 shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, and good results can be obtained.

[0415] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional motion vectors are derived.

[0416] [MV Derivation>FRUC>Template Matching]

[0417] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, as the specified area for calculating the evaluation value for candidates as described above, a block adjacent to the current block in the current picture is used.

[0418] Figure 24This is a conceptual diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture. As Figure 24 shown, in the second pattern matching, the motion vector of the current block is derived by searching for the block in the reference picture (Ref0) that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded region of both or one of the left and upper adjacent blocks and the reconstructed image at the equivalent position in the coded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. Among multiple candidate MVs, the candidate MV with the best evaluation value can be selected as the best candidate MV.

[0419] Information indicating whether to adopt the FRUC mode (e.g., referred to as the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the pattern matching method that can be adopted (the first pattern matching or the second pattern matching) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0420] [MV Derivation > Affine Mode]

[0421] Next, the affine mode of deriving the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks will be described. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0422] Figure 25A This is a conceptual diagram showing an example of the derivation of the motion vector in sub-block units based on the motion vectors of multiple adjacent blocks. In Figure 25A , the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the motion vectors of adjacent sub-blocks. Then, according to the following equation (1A), the two motion vectors v0 and v1 can be projected, and the motion vectors (v x , v y ) of each sub-block within the current block can also be derived.

[0423]

Equation 1

[0424]

[0425] Here, x and y represent the horizontal and vertical positions of the sub-blocks respectively, and w represents a prescribed weight coefficient. The prescribed weight coefficient can also be determined in advance.

[0426] The information indicating such an affine mode (e.g., called an affine flag) can be signaled as a signal at the CU level. In addition, the signaling of the information indicating the affine mode does not need to be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0427] In addition, in such an affine mode, several modes with different methods for deriving the motion vectors of the upper-left and upper-right control points can also be included. For example, in the affine mode, there are 2 modes: the affine inter-frame (also called affine ordinary inter-frame) mode and the affine merge mode.

[0428] [MV Derivation > Affine Mode]

[0429] Figure 25B is a conceptual diagram for explaining an example of the derivation of the motion vector of a sub-block unit in an affine mode with 3 control points. In Figure 25B , the current block includes 16 4×4 sub-blocks. Here, the motion vector v0 of the upper-left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector v1 of the upper-right control point of the current block is derived based on the motion vectors of adjacent blocks, and the motion vector v2 of the lower-left control point of the current block is derived based on the motion vectors of adjacent blocks. Then, according to the following equation (1B), the 3 motion vectors v0, v1, and v2 can be projected, and the motion vectors (v x , v y ) of each sub-block within the current block can also be derived.

[0430]

Equation 2

[0431]

[0432] Here, x and y represent the horizontal and vertical positions of the sub-block center respectively, w represents the width of the current block, and h represents the height of the current block.

[0433] Affine modes with different numbers of control points (e.g., 2 and 3) can also be switched and signaled at the CU level. In addition, the information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0434] In addition, in such an affine mode with three control points, several modes with different methods of deriving motion vectors of the upper left, upper right and lower left control points may be included. For example, in the affine mode, there are two modes, namely, the affine inter-frame (also called affine normal inter-frame) mode and the affine merge mode.

[0435] [MV Export > Affine Merge Mode]

[0436] Figure 26A , Figure 26B and Figure 26C This is a conceptual diagram for explaining the affine merge mode.

[0437] In affine merge mode, such as Figure 26A As shown, for example, based on a plurality of motion vectors corresponding to blocks encoded in an affine mode among the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, the predicted motion vector of each of the control points of the current block is calculated. Specifically, the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) are checked in this order to determine the first valid block encoded in the affine mode. The predicted motion vectors of the control points of the current block are calculated based on the plurality of motion vectors corresponding to the determined blocks.

[0438] For example, Figure 26B As shown in FIG. 1 , when a block A adjacent to the left side of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 projected to the positions of the upper left corner and the upper right corner of the encoded block including the block A are derived. Then, based on the derived motion vectors v3 and v4, a predicted motion vector v0 of the control point at the upper left corner of the current block and a predicted motion vector v1 of the control point at the upper right corner are calculated.

[0439] For example, Figure 26C As shown in FIG. 1 , when a block A adjacent to the left side of a current block is encoded in an affine mode having three control points, motion vectors v3, v4, and v5 are derived that are projected to the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, a predicted motion vector v0 of the control point at the upper left corner, a predicted motion vector v1 of the control point at the upper right corner, and a predicted motion vector v2 of the control point at the lower left corner of the current block are calculated.

[0440] In addition, in the following Figure 29 This predicted motion vector derivation method can also be used in the derivation of the predicted motion vectors of each control point of the current block in step Sj_1.

[0441] Figure 27 This is a flowchart showing an example of the affine merge mode.

[0442] In the affine merge mode, as shown in the figure, first, the inter-frame prediction unit 126 derives the predicted MVs of the respective control points of the current block (step Sk_1). The control points are, as Figure 25A shown, the upper left and upper right points of the current block, or, as Figure 25B shown, the upper left, upper right, and lower left points of the current block.

[0443] That is to say, as Figure 26A shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), and determines the initial valid block encoded in the affine mode.

[0444] Then, when block A is determined and block A has two control points, as Figure 26B shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner of the current block and the motion vector v1 of the control point at the upper right corner of the current block based on the motion vectors v3 and v4 of the upper left and upper right corners of the encoded block containing block A. For example, by projecting the motion vectors v3 and v4 of the upper left and upper right corners of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0445] Alternatively, when block A is determined and block A has three control points, as Figure 26C shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the control point at the upper left corner of the current block, the motion vector v1 of the control point at the upper right corner of the current block, and the motion vector v2 of the control point at the lower left corner of the current block based on the motion vectors v3, v4, and v5 of the upper left, upper right, and lower left corners of the encoded block containing block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left, upper right, and lower left corners of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the predicted motion vector v0 of the control point at the upper left corner of the current block, the predicted motion vector v1 of the control point at the upper right corner of the current block, and the motion vector v2 of the control point at the lower left corner of the current block.

[0446] Next, the inter-frame prediction unit 126 performs motion compensation on each of a plurality of sub-blocks included in the current block. That is, the inter-frame prediction unit 126 calculates the motion vector of each of the plurality of sub-blocks as an affine MV using two prediction motion vectors v0 and v1 and the above formula (1A), or three prediction motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-block (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0447] [MV Derivation>Affine Inter-frame Mode]

[0448] Figure 28A is a conceptual diagram for explaining the affine inter-frame mode with two control points.

[0449] In this affine inter-frame mode, as Figure 28A shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block.

[0450] Figure 28B is a conceptual diagram for explaining the affine inter-frame mode with three control points.

[0451] In this affine inter-frame mode, as Figure 28B shown, the motion vector selected from the motion vectors of the encoded blocks A, B, and C adjacent to the current block is used as the predicted motion vector v0 of the control point at the upper left corner of the current block. Similarly, the motion vector selected from the motion vectors of the encoded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point at the upper right corner of the current block. In addition, the motion vector selected from the motion vectors of the encoded blocks F and G adjacent to the current block is used as the predicted motion vector v2 of the control point at the lower left corner of the current block.

[0452] Figure 29 is a flowchart showing an example of the affine inter-frame mode.

[0453] As shown in the figure, in the affine inter-frame mode, first, the inter-frame prediction unit 126 derives the predicted MVs (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). As Figure 25A or Figure 25B shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0454] That is, the inter-frame prediction unit 126 derives the predicted motion vectors (v0, v1) or (v0, v1, v2) of the control points of the current block by selecting Figure 28A or Figure 28B a motion vector of one of the encoded blocks near each control point of the current block shown in the figure. At this time, the inter-frame prediction unit 126 encodes the predicted motion vector selection information for identifying the two selected motion vectors into the stream.

[0455] For example, the inter-frame prediction unit 126 can determine which block's motion vector to select from the encoded blocks adjacent to the current block as the predicted motion vector of the control point by using cost evaluation or the like, and can describe a flag indicating which predicted motion vector is selected in the bitstream.

[0456] Next, the inter-frame prediction unit 126 performs a motion search (steps Sj_3 and Sj_4) while using the predicted motion vectors selected or derived in the update step Sj_1 (step Sj_2). That is, the inter-frame prediction unit 126 uses the motion vectors of the respective sub-blocks corresponding to the predicted motion vectors to be updated as the affine MVs, and calculates using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference pictures (step Sj_4). As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted motion vector that can obtain the minimum cost as the motion vector of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted motion vector as a differential MV into the stream.

[0457] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference pictures (step Sj_6).

[0458] [MV Derivation>Affine Inter-frame Mode]

[0459] In the case of signaling in an affine mode that switches different numbers of control points (for example, 2 and 3) at the CU level, sometimes the number of control points in the encoded block and the current block is different. Figure 30A And Figure 30B is a conceptual diagram for explaining a method for deriving the predicted vector of the control point in the case where the number of control points in the encoded block and the current block is different.

[0460] For example, as Figure 30AAs shown, when the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points, motion vectors v3 and v4 projected onto the positions of the upper left corner and upper right corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated. In addition, the predicted motion vector v2 of the control point at the lower left corner is calculated based on the derived motion vectors v0 and v1.

[0461] For example, as Figure 30B shown, when the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated.

[0462] In Figure 29 the derivation of each predicted motion vector of the control points of the current block in step Sj_1, this predicted motion vector derivation method can also be used.

[0463] [MV Derivation > DMVR]

[0464] Figure 31A is a flowchart showing the relationship between the merge mode and DMVR.

[0465] The inter-frame prediction unit 126 derives the motion vector of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform a motion vector search, i.e., a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the motion vector derived in step Sl_1 as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is determined in the merge mode.

[0466] On the other hand, when it is determined in step Sl_1 that a motion search is to be performed (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final motion vector for the current block by searching the peripheral area of the reference picture represented by the motion vector derived in step Sl_1 (step Sl_3). That is, in this case, the motion vector of the current block is determined by DMVR.

[0467] Figure 31B is a conceptual diagram for explaining an example of the DMVR process for determining the MV.

[0468] First, set the best MVP set for the current block (e.g., in the merge mode) as the candidate MV. Then, according to the candidate MV (L0), determine the reference pixels based on the encoded picture in the L0 direction, i.e., the first reference picture (L0). Similarly, according to the candidate MV (L1), determine the reference pixels based on the encoded picture in the L1 direction, i.e., the second reference picture (L1). Generate a template by taking the average of these reference pixels.

[0469] Next, using the above template, search the surrounding areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) respectively, and determine the MV with the minimum cost as the final MV. In addition, the cost value can also be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values, etc.

[0470] In addition, typically, in the encoding device and the decoding device described later, the structures and operations of the processes described here are basically common.

[0471] Even if it is not the processing example itself described here, as long as it is a process capable of searching the surrounding of the candidate MV and deriving the final MV, any process can be used.

[0472] [Motion Compensation > BIO / OBMC]

[0473] In motion compensation, there is a mode of generating a prediction image and correcting the prediction image. This mode is, for example, BIO and OBMC described later.

[0474] Figure 32 It is a flowchart showing an example of the generation of a prediction image.

[0475] The inter-frame prediction unit 126 generates a prediction image (step Sm_1), and corrects the prediction image by, for example, any of the above modes (step Sm_2).

[0476] Figure 33 It is a flowchart showing another example of the generation of a prediction image.

[0477] The inter-frame prediction unit 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a prediction image (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final prediction image by correcting the prediction image (step Sn_4). On the other hand, when it is determined not to perform the correction process (No in step Sn_3), the inter-frame prediction unit 126 outputs the prediction image without correcting it as the final prediction image (step Sn_5).

[0478] In addition, in motion compensation, there is a mode of correcting luminance when generating a predicted image. This mode is, for example, LIC described later.

[0479] Figure 34 It is a flowchart showing another example of generating a predicted image.

[0480] The inter-frame prediction unit 126 derives the motion vector of the current block (step So_1). Next, the inter-frame prediction unit 126 determines whether to perform luminance correction processing (step So_2). Here, when it is determined to perform luminance correction processing (Yes in step So_2), the inter-frame prediction unit 126 generates a predicted image while performing luminance correction (step So_3). That is, the predicted image is generated by LIC. On the other hand, when it is determined not to perform luminance correction processing (No in step So_2), the inter-frame prediction unit 126 generates a predicted image by normal motion compensation without performing luminance correction (step So_4).

[0481] [Motion Compensation > OBMC]

[0482] Not only the motion information of the current block obtained by motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame prediction signal. Specifically, an inter-frame prediction signal can also be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search (within the reference picture) and a prediction signal based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0483] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (for example, referred to as the OBMC block size) can also be signaled at the sequence level. And information indicating whether to apply the OBMC mode (for example, referred to as the OBMC flag) can also be signaled at the CU level. In addition, the level of signaling of this information does not need to be limited to the sequence level and the CU level, and can also be other levels (for example, picture level, slice level, tile level, CTU level, or sub-block level).

[0484] A more specific example of the OBMC mode will be described. Figure 35 and Figure 36 are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction processing based on OBMC processing.

[0485] First, as Figure 36 shown, using the motion vector (MV) assigned to the block to be processed (current), a predicted image (Pred) based on normal motion compensation is obtained. InFigure 36 In this case, the arrow "MV" points to the reference picture and indicates which block the current block of the current picture refers to for obtaining the predicted image.

[0486] Next, the motion vector (MV_L) already derived for the encoded left adjacent block is applied (reused) to the coding target block to obtain the predicted image (Pred_L). The motion vector (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0487] Similarly, the motion vector (MV_U) already derived for the encoded upper adjacent block is applied (reused) to the coding target block to obtain the predicted image (Pred_U). The motion vector (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted image that has undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block where the boundary with the adjacent block is blended (smoothed).

[0488] In addition, the above example is a two-path correction method using the left adjacent and upper adjacent blocks, but this correction method can also be a three-path or more path correction method that also uses the right adjacent and / or lower adjacent blocks.

[0489] Furthermore, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.

[0490] In addition, the prediction image correction process of OBMC is described here. The prediction image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can also be applied to each of the multiple reference pictures. In this case, through the OBMC image correction based on multiple reference pictures, after obtaining the corrected predicted images from each reference picture, the final predicted image is obtained by further overlapping the obtained multiple corrected predicted images.

[0491] In addition, in OBMC, the unit of the target block can be the prediction block unit or the sub-block unit obtained by further dividing the prediction block.

[0492] As a method for determining whether to apply OBMC processing, for example, there is a method of using a signal indicating whether to apply OBMC processing, i.e., obmc_flag. As a specific example, the encoding device can also determine whether the target block belongs to a region with complex motion. When the encoding device belongs to a region with complex motion, it sets the obmc_flag value to 1 and applies OBMC processing for encoding. When it does not belong to a region with complex motion, it sets the obmc_flag value to 0 and encodes the block without applying OBMC processing. On the other hand, in the decoding device, by decoding the obmc_flag described in the stream (e.g., compressed sequence), it switches whether to apply OBMC processing according to this value for decoding.

[0493] In the above example, the inter prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter prediction unit 126 can generate multiple prediction images with shapes different from the rectangle for the rectangular current block, and can generate the final rectangular prediction image by combining these multiple prediction images. Shapes different from the rectangle can also be triangles, for example.

[0494] Figure 37 It is a conceptual diagram for explaining the generation of two triangular prediction images.

[0495] The inter prediction unit 126 generates a triangular prediction image by performing motion compensation on the first triangular partition within the current block using the first MV of the first partition. Similarly, the inter prediction unit 126 generates a triangular prediction image by performing motion compensation on the second triangular partition in the current block using the second MV of the second partition. Then, the inter prediction unit 126 generates a rectangular prediction image identical to the current block by combining these prediction images.

[0496] In addition, in Figure 37 the example shown, the first partition and the second partition are triangles respectively, but they can also be trapezoids, or can be of different shapes respectively. Moreover, in Figure 37 the example shown, the current block is composed of two partitions, but it can also be composed of three or more partitions.

[0497] In addition, the first partition and the second partition can also be repeated. That is, the first partition and the second partition can also include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition can be used to generate the prediction image of the current block.

[0498] In addition, in this example, an example where prediction images are generated for both partitions through inter prediction is shown, but prediction images can also be generated for at least one partition through intra prediction.

[0499] [Motion Compensation > BIO]

[0500] Next, a method for deriving motion vectors will be described. First, a mode for deriving motion vectors based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi-directional optical flow) mode.

[0501] Figure 38 is a conceptual diagram for explaining the model assuming uniform linear motion. In Figure 38 , (v x , v y ) represents the velocity vector, τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.

[0502] At this time, it can also be that, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) is adopted.

[0503] [Equation 3]

[0504]

[0505] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation means that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. It can also be that, based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a merge list or the like are corrected in pixel units.

[0506] In addition, motion vectors can also be derived on the decoding device side by a method different from the derivation of motion vectors based on the model assuming uniform linear motion. For example, motion vectors can also be derived in sub-block units based on the motion vectors of multiple adjacent blocks.

[0507] [Motion Compensation > LIC]

[0508] Next, an example of a mode for generating a predicted image (prediction) by performing LIC (local illumination compensation) processing will be described.

[0509] Figure 39 It is a conceptual diagram for explaining an example of a predicted image generation method using a luminance correction process based on LIC processing.

[0510] First, an MV is derived from the encoded reference picture, and a reference image corresponding to the current block is obtained.

[0511] Next, information indicating how the luminance values change in the reference picture and the current picture is extracted for the current block. This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same positions within the reference picture specified by the derived MV. Then, using the information indicating how the luminance values change, a luminance correction parameter is calculated.

[0512] By applying the above luminance correction parameter to the reference image within the reference picture specified by the MV, a predicted image for the current block is generated.

[0513] In addition, Figure 39 the shape of the above peripheral reference region in Figure 39 is an example, and shapes other than this can also be used.

[0514] Furthermore, the processing of generating a predicted image based on one reference picture has been described here, but the same applies when generating a predicted image based on multiple reference pictures. A predicted image can also be generated after performing luminance correction processing on the reference images obtained from each reference picture in the same manner as above.

[0515] As a method for determining whether to adopt LIC processing, for example, there is a method of using lic_flag as a signal indicating whether to adopt LIC processing. As a specific example, in an encoding device, it is determined whether the current block belongs to a region where a luminance change has occurred. In the case where it belongs to a region where a luminance change has occurred, the value 1 is set as lic_flag, and encoding is performed using LIC processing. In the case where it does not belong to a region where a luminance change has occurred, the value 0 is set as lic_flag, and encoding is performed without using LIC processing. On the other hand, in a decoding device, it is also possible to decode the lic_flag described in the stream and switch whether to adopt LIC processing according to its value for decoding.

[0516] As another method for determining whether to use the LIC process, there is a method of determining whether the LIC process is used in the surrounding blocks. As a specific example, when the current block is in the merge mode, it is determined whether the surrounding coded blocks selected when deriving the MV in the merge mode process are coded using the LIC process, and based on the result, it is switched whether to use the LIC process for coding. In addition, in the case of this example, the same process is also applied to the decoding device side.

[0517] use Figure 39 The form of the LIC process (luminance correction process) has been described, and the details thereof will be described below.

[0518] First, the inter prediction unit 126 derives a motion vector for acquiring a reference image corresponding to a current block to be encoded from a reference picture that is an already encoded picture.

[0519] Next, the inter-frame prediction unit 126 uses the brightness pixel values of the encoded surrounding reference areas adjacent to the left and above and the brightness pixel values at the same position in the reference picture specified by the motion vector to extract information indicating how the brightness values in the reference picture and the encoding target picture change, and calculates the brightness correction parameter. For example, the brightness pixel value of a certain pixel in the surrounding reference area in the encoding target picture is set to p0, and the brightness pixel value of the pixel in the surrounding reference area in the reference picture at the same position as the pixel is set to p1. The inter-frame prediction unit 126 calculates the coefficients A and B for optimizing A×p1+B=p0 as brightness correction parameters for multiple pixels in the surrounding reference area.

[0520] Next, the inter-frame prediction unit 126 generates a predicted image for the encoding target block by performing a brightness correction process on the reference image in the reference picture specified by the motion vector using the brightness correction parameter. For example, the brightness pixel value in the reference image is set to p2, and the brightness pixel value of the predicted image after the brightness correction process is set to p3. The inter-frame prediction unit 126 generates a predicted image after the brightness correction process by calculating A×p2+B=p3 for each pixel in the reference image.

[0521] also, Figure 39 The shape of the peripheral reference area in is an example, and other shapes may also be used. Figure 39 A portion of the peripheral reference area shown. For example, an area including a predetermined number of pixels thinned out from the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference area. In addition, the peripheral reference area is not limited to an area adjacent to the encoding object block, and may also be an area not adjacent to the encoding object block. The predetermined number related to the pixel may also be predetermined.

[0522] In addition,Figure 39 In the example shown, the peripheral reference area in the reference picture is the area specified by the motion vector of the coded object picture from among the peripheral reference areas in the coded object picture, but it may also be an area specified by another motion vector. For example, this other motion vector may also be the motion vector of the peripheral reference area in the coded object picture.

[0523] In addition, although the operation in the encoding device 100 has been described here, typically, the operation in the decoding device 200 is the same.

[0524] Furthermore, the LIC process can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0525] Furthermore, the LIC process can also be applied in units of sub-blocks. For example, correction parameters can be derived using the peripheral reference area of the current sub-block and the peripheral reference area of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0526] [Prediction control unit]

[0527] The prediction control unit 128 selects one of the intra-prediction signal (the signal output from the intra-prediction unit 124) and the inter-prediction signal (the signal output from the inter-prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0528] As Figure 1 shown, in various examples of encoding devices, the prediction control unit 128 may also output prediction parameters input to the entropy encoding unit 110. The entropy encoding unit 110 can generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device. The decoding device can also receive and decode the encoded bitstream and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters may include the selection of the prediction signal (e.g., the motion vector, the prediction type, or the prediction mode used by the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0529] [Installation example of encoding device]

[0530] Figure 40 is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 1The multiple components of the encoding device 100 shown are implemented by Figure 40 the processor a1 and the memory a2 shown.

[0531] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. The processor a1 can also be a processor such as a CPU. In addition, the processor a1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 can also play the role of Figure 1 multiple components among the multiple components of the encoding device 100 shown.

[0532] The memory a2 is a dedicated or general-purpose memory that stores information for the processor a1 to encode moving images. The memory a2 can be either an electronic circuit or connected to the processor a1. In addition, the memory a2 can also be included in the processor a1. In addition, the memory a2 can also be an aggregate of multiple electronic circuits. In addition, the memory a2 can be a magnetic disk or an optical disk, etc., or can be represented as a storage or a recording medium, etc. In addition, the memory a2 can be either a non-volatile memory or a volatile memory.

[0533] For example, the memory a2 can store the encoded moving image or can store the bit string corresponding to the encoded moving image. In addition, a program for the processor a1 to encode moving images can also be stored in the memory a2.

[0534] In addition, for example, the memory a2 can also play the role of Figure 1 the component for storing information among the multiple components of the encoding device 100 shown. For example, the memory a2 can play the role of Figure 1 the block memory 118 and the frame memory 122 shown. More specifically, the memory a2 can store the reconstructed blocks and the reconstructed pictures, etc.

[0535] In addition, in the encoding device 100, all of the multiple components shown Figure 1 etc. may not be installed, and all of the above-mentioned multiple processes may not be performed. Figure 1 A part of the multiple components shown

[0536] [Decoding device]

[0537] Next, a decoding device that can decode the encoded signal (encoded bitstream) output from the above-mentioned encoding device 100, for example, will be described. Figure 41FIG. 0 is a block diagram showing the functional configuration of a decoding apparatus 200 according to an embodiment. The decoding apparatus 200 is a moving image decoding apparatus that decodes a moving image in units of blocks.

[0538] As Figure 41 shown in FIG. 5, the decoding apparatus 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0539] The decoding apparatus 200 is implemented, for example, by a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoding apparatus 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filtering unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0540] Hereinafter, after explaining the flow of the overall processing of the decoding apparatus 200, each component included in the decoding apparatus 200 will be described.

[0541] [Overall Flow of Decoding Process]

[0542] Figure 42 FIG. 18 is a flowchart showing an example of the overall decoding process performed by the decoding apparatus 200.

[0543] First, the entropy decoding unit 202 of the decoding apparatus 200 determines a division pattern (step Sp_1) of a block of a fixed size (e.g., 128×128 pixels). This division pattern is the division pattern selected by the encoding apparatus 100. Then, the decoding apparatus 200 performs the processing of steps Sp_2 to Sp_6 on each of the plurality of blocks constituting the division pattern.

[0544] That is, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the block to be decoded (also referred to as the current block) (step Sp_2).

[0545] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore a plurality of prediction residuals (i.e., differential blocks) by performing inverse quantization and inverse transformation on the plurality of quantization coefficients (step Sp_3).

[0546] Next, a prediction processing unit, which is composed of all or part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220, generates a prediction signal (also referred to as a prediction block) for the current block (step Sp_4).

[0547] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction block to the differential block (step Sp_5).

[0548] Moreover, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).

[0549] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeated.

[0550] As shown in the figure, the processing of steps Sp_1 to Sp_7 is sequentially performed by the decoding device 200. Alternatively, multiple processes of a part of these processes can be performed in parallel, or the order can be changed, etc.

[0551] [Entropy Decoding Unit]

[0552] The entropy decoding unit 202 performs entropy decoding on the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetic decodes the encoded bitstream into a binary signal. Next, the entropy decoding unit 202 de-binarizes the binary signal. Thus, the entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the encoded bitstream to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment (refer to Figure 1 ) and the prediction parameters included therein. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device side.

[0553] [Inverse Quantization Unit]

[0554] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the decoding target block (hereinafter referred to as the current block) that is the input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. And the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0555] [Inverse Transform Unit]

[0556] The inverse transform unit 206 restores the prediction error by inversely transforming the transform coefficients that are the input from the inverse quantization unit 204.

[0557] For example, when the information read from the encoded bitstream indicates the use of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 inversely transforms the transform coefficients of the current block based on the information indicating the transform type read.

[0558] In addition, for example, when the information read from the encoded bitstream indicates the use of NSST, the inverse transform unit 206 applies an inverse re - transformation to the transform coefficients.

[0559] [Addition unit]

[0560] The addition unit 208 reconstructs the current block by adding the prediction error that is the input from the inverse transform unit 206 and the prediction sample that is the input from the prediction control unit 220. Also, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0561] [Block memory]

[0562] The block memory 210 is a storage unit for storing blocks within the decoded picture (hereinafter referred to as the current picture) that are referred to in intra - prediction. Specifically, the block memory 210 stores the reconstructed blocks output from the addition unit 208.

[0563] [Loop filter unit]

[0564] The loop filter unit 212 applies loop filtering to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214, the display device, etc.

[0565] When the information indicating the on / off of ALF read from the encoded bitstream indicates that ALF is on, one filter is selected from among multiple filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed block.

[0566] [Frame memory]

[0567] The frame memory 214 is a storage unit for storing reference pictures used in inter - prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0568] [Prediction processing unit (intra - prediction unit / inter - prediction unit / prediction control unit)]

[0569] Figure 43This is a flowchart showing an example of the processing performed by the prediction processing unit of the decoding device 200. In addition, the prediction processing unit is composed of all or some of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0570] The prediction processing unit generates a prediction image of the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there are an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit uses the reconstructed image that has been obtained by generating a prediction block, a differential block, a coefficient block, restoring the differential block, and generating a decoded image block, to generate a prediction image of the current block.

[0571] The reconstructed image can be, for example, an image of a reference picture, or an image of a decoded block within the current picture that includes the current block. The decoded block within the current picture is, for example, an adjacent block of the current block.

[0572] Figure 44 This is a flowchart showing another example of the processing performed by the prediction processing unit of the decoding device 200.

[0573] The prediction processing unit determines the method or mode for generating the prediction image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.

[0574] When it is determined that the first method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the first method (step Sr_2a). In addition, when it is determined that the second method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the second method (step Sr_2b). In addition, when it is determined that the third method is the mode for generating the prediction image, the prediction processing unit generates the prediction image according to the third method (step Sr_2c).

[0575] The first method, the second method, and the third method are different methods for generating the prediction image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.

[0576] [Intra Prediction Unit]

[0577] The intra prediction unit 216 performs intra prediction by referring to the blocks within the current picture stored in the block memory 210 based on the intra prediction mode read from the encoded bitstream, thereby generating a prediction signal (intra prediction signal). Specifically, the intra prediction unit 216 generates an intra prediction signal by referring to the samples (such as luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0578] In addition, when the intra prediction mode of the reference luminance block is selected in the intra prediction of the color difference block, the intra prediction unit 216 may also predict the color difference component of the current block based on the luminance component of the current block.

[0579] In addition, when the information decoded from the coded bitstream indicates the adoption of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0580] [Inter prediction unit]

[0581] The inter prediction unit 218 predicts the current block with reference to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4×4 blocks) within the current block. For example, the inter prediction unit 218 performs motion compensation using the motion information (e.g., motion vector) decoded from the coded bitstream (e.g., the prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction signal for the current block or sub-block, and outputting the inter prediction signal to the prediction control unit 220.

[0582] When the information decoded from the coded bitstream indicates the adoption of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained through motion estimation but also the motion information of adjacent blocks.

[0583] In addition, when the information decoded from the coded bitstream indicates the adoption of the FRUC mode, the inter prediction unit 218 performs motion estimation according to the pattern matching method (bidirectional matching or template matching) decoded from the coded stream, thereby deriving the motion information. And the inter prediction unit 218 uses the derived motion information for motion compensation (prediction).

[0584] In addition, when the BIO mode is adopted, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when the information decoded from the coded bitstream indicates the adoption of the affine motion compensation prediction mode, the inter prediction unit 218 derives a motion vector in units of sub-blocks based on the motion vectors of multiple adjacent blocks.

[0585] [MV derivation > Normal inter-frame mode]

[0586] When the information read from the coded bitstream indicates the application of the normal inter-frame mode, the inter prediction unit 218 derives an MV based on the information read from the coded bitstream, and uses the MV for motion compensation (prediction).

[0587] Figure 45 is a flowchart showing an example of inter prediction based on the normal inter-frame mode in the decoding apparatus 200.

[0588] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. The inter-frame prediction unit 218 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Ss_1). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0589] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Ss_1 as prediction motion vector candidates (also referred to as prediction MV candidates) in a prescribed priority order (step Ss_2). In addition, this priority order may also be determined in advance for each of the N prediction MV candidates.

[0590] Next, the inter-frame prediction unit 218 decodes prediction motion vector selection information from the input stream (i.e., the encoded bitstream), and uses the decoded prediction motion vector selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction motion vector (also referred to as the prediction MV) of the current block (step Ss_3).

[0591] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction motion vector (step Ss_4).

[0592] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Ss_5).

[0593] [Prediction control unit]

[0594] The prediction control unit 220 selects one of the intra-frame prediction signal and the inter-frame prediction signal, and outputs the selected signal as the prediction signal to the adder 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device side.

[0595] [Installation example of decoding device]

[0596] Figure 46 is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 41 as shown, the multiple components of the decoding device 200 pass through Figure 46It is installed using the processor b1 and the memory b2 shown.

[0597] The processor b1 is a circuit that processes information and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes an encoded moving image (i.e., an encoded bitstream). The processor b1 can also be a processor such as a CPU. Additionally, the processor b1 can also be an aggregate of multiple electronic circuits. Further, for example, the processor b1 can also perform the functions of Figure 41 multiple components among the multiple components of the decoding device 200 shown, etc.

[0598] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode the encoded bitstream. The memory b2 can be an electronic circuit or can be connected to the processor b1. Additionally, the memory b2 can also be included in the processor b1. Further, the memory b2 can also be an aggregate of multiple electronic circuits. Additionally, the memory b2 can be a magnetic disk or an optical disk, etc., and can also be represented as a storage or a recording medium, etc. Further, the memory b2 can be either a non-volatile memory or a volatile memory.

[0599] For example, the memory b2 can store a moving image or an encoded bitstream. Additionally, a program for the processor b1 to decode the encoded bitstream can also be stored in the memory b2.

[0600] Further, for example, the memory b2 can perform the functions of Figure 41 the component for storing information among the multiple components of the decoding device 200 shown, etc. Specifically, the memory b2 can perform the functions of Figure 41 the block memory 210 and the frame memory 214 shown. More specifically, the memory b2 can store the reconstructed blocks and the reconstructed pictures, etc.

[0601] Additionally, in the decoding device 200, not all of the multiple components shown, etc., need to be installed, Figure 41 nor do all of the above-mentioned multiple processes need to be performed. Figure 41 A part of the multiple components shown, etc., can be included in other devices, or a part of the above-mentioned multiple processes can be performed by other devices.

[0602] [Definitions of Each Term]

[0603] As an example, each term can also be defined as follows.

[0604] A picture is an arrangement of multiple luma samples in monochrome format, or an arrangement of multiple luma samples and 2 corresponding arrangements of multiple color difference samples in color formats of 4:2:0, 4:2:2, and 4:4:4. A picture can be a frame or a field.

[0605] A frame is a composition of a top field that generates a plurality of sample rows 0, 2, 4, ... and a bottom field that generates a plurality of sample rows 1, 3, 5, ...

[0606] A slice is an integer number of coding tree units contained in 1 independent slice segment and all subsequent dependent slice segments before the next independent slice segment (if any) in the same access unit.

[0607] A tile is a rectangular region of multiple coding tree blocks within a specific tile column and a specific tile row in a picture. A tile can still apply a loop filter across the edge of the tile, but can also be a rectangular region of a frame that is intended to be independently decodable and coded.

[0608] A block is an MxN (N rows and M columns) arrangement of multiple samples, or an MxN arrangement of multiple transform coefficients. A block can also be a square or rectangular area of multiple pixels consisting of multiple matrices of 1 luminance and 2 chrominance.

[0609] A CTU (Coding Tree Unit) can be a coding tree block of multiple luma samples of a picture with 3 sample arrangements, or 2 corresponding coding tree blocks of multiple chrominance samples. Alternatively, a CTU can be a coding tree block of any number of samples in a monochrome picture and a picture encoded using the syntax structure used in the encoding of 3 separate color planes and multiple samples.

[0610] A super block is composed of one or two mode information blocks, or may be recursively divided into four 32×32 blocks, or further divided into a square block of 64×64 pixels.

[0611] [BDOF (BIO) Processing Overview]

[0612] use Figure 38 , Figure 47 and Figure 48 The outline of the BDOF (BIO) process will be described again.

[0613] Figure 47 This is a flowchart showing an example of inter-frame prediction according to BIO. Figure 48 1 is a diagram showing an example of the functional structure of the inter-frame prediction unit 126 that performs inter-frame prediction according to the BIO.

[0614] like Figure 48As shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a prediction image correction unit 126f. Additionally, the memory 126a may also be the frame memory 122.

[0615] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 126 uses these two motion vectors (M0, M1) to derive the prediction image of the current block (step Sy_1). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0616] Next, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . Additionally, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_2). Here, the interpolation image I 0 is the image included in the reference picture Ref0 derived for the current block, and the interpolation image I 1 is the image included in the reference picture Ref1 derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Or, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each also be an image larger than the current block. Moreover, the interpolation image I 0 and I 1 can also include the motion vectors (M0, M1) and the reference pictures (L0, L1), and the prediction image derived by applying the motion compensation filter.

[0617] Furthermore, the gradient image derivation unit 126c derives the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) of the current block from the interpolation image I 0 and the interpolation image I 1 (step Sy_3). Additionally, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy0 , Iy 1 ). The gradient image derivation unit 126c can derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image only needs to represent the spatial change amount of pixel values along the horizontal direction or the vertical direction.

[0618] Next, the optical flow derivation unit 126d uses the interpolated image (I 0 , I 1 ) and the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) to derive the optical flow (vx, vy) as the above-mentioned velocity vector (step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and can be called a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can also be a 4×4 pixel sub-CU. In addition, the derivation of the optical flow can also be performed in other units such as pixel units instead of sub-block units.

[0619] Next, the inter-frame prediction unit 126 corrects the prediction image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives the correction value of the pixel values included in the current block using the optical flow (vx, vy) (step Sy_5). Then, the prediction image correction unit 126f can correct the prediction image of the current block using the correction value (step Sy_6). In addition, the correction value can be derived in units of each pixel, or in units of multiple pixels or sub-blocks.

[0620] In addition, the processing flow of BIO is not limited to Figure 47 the disclosed processing. It is possible to execute only Figure 47 a part of the processing disclosed in

[0621] For example, BDOF (BIO) can be defined as a process of generating a corrected prediction image using the gradient of the reference block. Or, BDOF can be defined as a process of generating a prediction image using the gradients of two respective reference blocks. Or, BDOF can be defined as a process of generating a prediction image using the gradient in a mode different from the affine mode. Or, BDOF can be defined as an optical flow process in a mode different from the affine mode. Here, the gradient refers to the spatial gradient of pixel values.

[0622] [The first specific example of BIO]

[0623] Next, using Figure 49 , Figure 50 and Figure 51, a first specific example of BIO-based decoding processing is described. For example, the decoding device 200 calculates BIO parameters based on two reference blocks that have undergone motion compensation, and decodes the current block using the calculated BIO parameters. The BIO parameters are parameters corresponding to the above-mentioned optical flow, and are also referred to as local motion estimation values, modified motion vectors, modified MV values, modified weighted motion vectors, or BIO correction values.

[0624] Figure 49 is a flowchart showing a first specific example of BIO-based decoding processing. Figure 50 is a conceptual diagram showing an example of calculating the gradient value in the horizontal direction, i.e., the horizontal gradient value. Figure 51 is a conceptual diagram showing an example of calculating the gradient value in the vertical direction, i.e., the vertical gradient value.

[0625] First, for the current block, the decoding device 200 calculates the first sum (S1001) using the horizontal gradient value of the first reference block and the horizontal gradient value of the second reference block. The current block can be a sub-block of the current coding unit (current CU) as shown in Figure 50 and Figure 51 .

[0626] The first reference block is a block referred to in the decoding of the current block, and is a block determined in the reference picture L0 according to the first motion vector of the current block or the current CU. The second reference block is a block referred to in the decoding of the current block, and is a block determined in the reference picture L1 according to the second motion vector of the current block or the current CU.

[0627] Basically, the reference picture L0 and the reference picture L1 are two different reference pictures, and the first reference block and the second reference block are two different reference blocks. In addition, here, the first reference block and the second reference block that are respectively adjusted with sub-pixel precision by the motion compensation filter are used, and they have the same size as the current block.

[0628] When the current block is a sub-block of the current coding unit, the first reference block and the second reference block can respectively be sub-blocks of the reference blocks of the current coding unit.

[0629] That is, Figure 50 and Figure 51 The multiple pixel values in the reference picture L0 can also be the multiple pixel values of the block determined as the reference block of the current coding unit with respect to the reference picture L0. Similarly, Figure 50 and Figure 51 The multiple pixel values in the reference picture L1 can also be the multiple pixel values of the block determined as the reference block of the current coding unit with respect to the reference picture L1.

[0630] The decoding device 200 calculates the first sum using the horizontal gradient value of the above-mentioned first reference block and the horizontal gradient value of the above-mentioned second reference block. The decoding device 200 is not limited to the horizontal gradient value of the first reference block and the horizontal gradient value of the second reference block, and may also use the horizontal gradient values around the first reference block and the horizontal gradient values around the second reference block to calculate the first sum. The following formulas (3.1) and (3.2) show examples of the calculation processing method of the first sum.

[0631]

Equation 4

[0632] G x =I x 0 +I x 1 (3.1)

[0633]

Equation 5

[0634] sG x =∑ [i,j]∈Ω sign(G x [i,j])×G x [i,j] (3.2)

[0635] Here, × represents multiplication, and + represents addition. In addition, sign represents the positive or negative sign. Specifically, sign is represented by 1 for positive and -1 for negative. More specifically, for example, sign represents the following.

[0636]

Equation 6

[0637]

[0638] As a result, sign(G x [i,j])×G x [i,j] becomes directly G x [i,j] when G x [i,j] is positive, and becomes -G x [i,j] when G x [i,j] is negative, that is, it is equal to the absolute value (abs) of G x [i,j] derived.

[0639] In formula (3.1), I x 0 represents the horizontal gradient value in the first reference block of the reference picture L0, and I x 1 represents the horizontal gradient value in the second reference block of the reference picture L1.

[0640] An example of a horizontal gradient filter for obtaining a horizontal gradient value is a 3-tap filter having a filter coefficient set of [-1, 0, 1]. The horizontal gradient value in the first reference block is calculated by applying the horizontal gradient filter to a plurality of reference pixels in the first reference block. The horizontal gradient value in the second reference block is calculated by applying the horizontal gradient filter to a plurality of reference pixels in the second reference block.

[0641] In Figure 50 the example shown, the horizontal gradient value I of the pixel at [3, 2] in the first reference block x 0 is calculated as the matrix product [-1, 0, 1] T [2, 3, 5], and its value is 3. The horizontal gradient value I of the pixel at [3, 2] in the second reference block x 1 is calculated as the matrix product [-1, 0, 1] T [5, 3, 2], and its value is -3. In addition, [a, b, c] represents a 3-row and 1-column matrix.

[0642] In Equation (3.2), sG x represents the first sum, which is calculated as the sum of the absolute values of G x over the window represented by Ω. The size of Ω can also be the same as the current block. Additionally, the size of Ω can be larger than the size of the current block. In the latter case, the values of G x at adjacent positions of the current block are included in the calculation process of the first sum.

[0643] In addition, the decoding device 200 calculates a second sum (S1002) for the current block in the same manner as the first sum, which is the sum of the horizontal gradient values, using the vertical gradient values of the first reference block and the vertical gradient values of the second reference block. The following Equations (3.3) and (3.4) show examples of the calculation processing method of the second sum.

[0644]

Equation 7

[0645] G y = I y 0 + I y 1 (3.3)

[0646]

Equation 8

[0647] sGy = ∑ [i,j]∈Ω sign(G y [i, j]) × G y [i, j] (3.4)

[0648] In Equation (3.3), I y 0Represents the vertical gradient value in the first reference block referring to picture L0, I y 1 Represents the vertical gradient value in the second reference block referring to picture L1.

[0649] An example of a vertical gradient filter for obtaining the vertical gradient value is a 3-tap filter having a filter coefficient set of [-1, 0, 1]. The vertical gradient value in the first reference block is calculated by applying the vertical gradient filter to a plurality of reference pixels in the first reference block. The vertical gradient value in the second reference block is calculated by applying the vertical gradient filter to a plurality of reference pixels in the second reference block.

[0650] In Figure 51 the example shown, the vertical gradient value I of the pixel at [3, 2] in the first reference block y 0 is calculated as the matrix product [-1, 0, 1] T [2, 3, 5], whose value is 3. The vertical gradient value I of the pixel at [3, 2] in the second reference block y 1 is calculated as the matrix product [-1, 0, 1] T [5, 3, 2], whose value is -3.

[0651] In Equation (3.4), sG y represents the second sum, calculated as the absolute value sum of G y over the window represented by Ω. In the case where the size of Ω is larger than the size of the current block, the value of G y at adjacent positions of the current block is included in the calculation process of the second sum.

[0652] Next, the decoding device 200 determines whether the first sum is greater than the second sum (S1003). In the case where it is determined that the first sum is greater than the second sum (yes in S1003), the decoding device 200 determines the BIO parameter without using the vertical gradient value for the current block (S1004). Equations (3.5) to (3.9) represent examples of the arithmetic processing for determining the BIO parameter in this case. In these equations, the BIO parameter represented by u is calculated using the horizontal gradient value.

[0653]

Equation 9

[0654] sG x dI = ∑ [i,j]∈Ω -sign(G x [i, j]) × (I 0 i,j - I 1 i,j ) (3.5)

[0655]

Equation 10

[0656] denominator = sG x (3.6)

[0657]

Equation 11

[0658] numerator = sG x dI (3.7)

[0659]

Equation 12

[0660] shift = Bits(denominator) - 6 (3.8)

[0661]

Equation 13

[0662] u = Clip(((numerator >> (shift - 12)) × BIOShift[demoninator >> shift]) >> 27) (3.9)

[0663] Here, - represents subtraction, and >> represents a shift operation. For example, a >> b means shifting a to the right by b bits. Additionally, Bits, BIOShift, and Clip represent the following respectively. Also, hereinafter, ceil represents rounding up a decimal, and floor represents truncating a decimal.

[0664]

Equation 14

[0665]

[0666] In Equation (3.5), sG x dI is calculated as the product sum of the difference between I 0 i,j and I 1 i,j throughout window Ω and sign(G x [i, j]). Here, I 0 i,j represents the pixel value at position [i, j] within the first reference block of reference picture L0, and I 1 i,j represents the pixel value at position [i, j] within the second reference block of reference picture L1. I 0 i,j and I 1 i,j may sometimes be represented only as I 0 and I 1 . The BIO parameter u is calculated by Equations (3.6) to (3.9) using sG x dI and sG x .

[0667] When it is determined that the first sum is not greater than the second sum (No in S1003), the decoding device 200 determines the BIO parameter u without using the horizontal gradient value for the current block (S1005). Expressions (3.10) to (3.14) represent examples of arithmetic processing for determining the BIO parameter u in this case. Expressions (3.10) to (3.14) are basically the same as Expressions (3.5) to (3.9), but in Expressions (3.10) to (3.14), the BIO parameter u is calculated using the vertical gradient value.

[0668]

Equation 15

[0669] sG y dI = Σ [i,j]∈Ω -sign(G y [i, j]) × (I 0 i,j -I 1 i,j ) (3.10)

[0670]

Equation 16

[0671] denominator = sG y (3.11)

[0672]

Equation 17

[0673] numerator = sG y dI (3.12)

[0674]

Equation 18

[0675] shift = Bits(denominator) - 6 (3.13)

[0676]

Equation 19

[0677] u = Clip(((numerator >> (shift - 12)) × BIOShift[demoninator >> shift]) >> 27) (3.14)

[0678] In Equation (3.10), sG y dI is calculated as the sum of the products of the difference between I 0 i,j and I 1 i,j throughout the window Ω and sign(G y [i, j]). The BIO parameter u is calculated using sG y dI and sG y through Expressions (3.11) to (3.14).

[0679] Then, the decoding device 200 decodes the current block using the BIO parameter u (S1006). Specifically, the decoding device 200 generates a prediction sample using the BIO parameter u and decodes the current block using the prediction sample. Equations (3.15) and (3.16) represent multiple examples of the arithmetic processing for generating the prediction sample.

[0680]

Equation 20

[0681] predictionsample = (I 0 + I 1 + u × (I x 0 - I x 1 )) >> 1 (3.15)

[0682]

Equation 21

[0683] prediction sample = (I 0 + I 1 + u × (I y 0 - I y 1 )) >> 1 (3.16)

[0684] In the case where it is determined that the first sum is greater than the second sum (Yes in S1003), Equation (3.15) is used. In the case where it is determined that the first sum is not greater than the second sum (No in S1003), Equation (3.16) is used.

[0685] The decoding device 200 may also repeatedly perform the above processing on all sub-blocks of the current CU (S1001 to S1006).

[0686] By using BIO, the decoding device 200 can improve the accuracy of the prediction sample in the current block. In addition, when calculating the BIO parameter, the decoding device 200 uses only one of the horizontal gradient value and the vertical gradient value, so an increase in the amount of computation can be suppressed.

[0687] In addition, the above equations are just an example, and the equations for calculating the BIO parameter are not limited to the above equations. For example, the plus and minus signs included in the above equations can be appropriately changed, or equations equivalent to the above equations can be used. Specifically, as equations corresponding to the above Equations (3.1) and (3.2), the following Equation (4.1) can also be used.

[0688]

Equation 22

[0689] sG x = ∑ [i,j]∈Ω abs(Ix 1 +I x 0 ) (4.1)

[0690] In addition, for example, as an equation corresponding to the above equation (3.5), the following equation (4.2) can also be used.

[0691]

Mathematical formula 23

[0692] sG x dI = ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 ) × (I 0 -I 1 )) (4.2)

[0693] In addition, for example, as an equation corresponding to the above equation (3.15), the following equation (4.3) can also be used.

[0694]

Mathematical formula 24

[0695] prediction sample = (I 1 +I 0 -u × (I x 1 -I x 0 )) >> 1 (4.3)

[0696] In addition, equations (3.6) to (3.9) essentially represent division, so they can also be expressed as follows in equation (4.4).

[0697]

Mathematical formula 25

[0698]

[0699] These equations (4.1) to (4.4) are substantially the same as the above equations (3.1), (3.2), (3.5) to (3.9), and (3.15).

[0700] Similarly, specifically, as equations corresponding to the above equations (3.3) and (3.4), the following equation (4.5) can also be used.

[0701]

Mathematical formula 26

[0702] sG y = ∑ [i,j]∈Ω abs(I y 1 +I y 0) (4.5)

[0703] In addition, for example, as an equation corresponding to the above equation (3.10), the following equation (4.6) can also be used.

[0704]

Mathematical formula 27

[0705] sG y dI = ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I 1 )) (4.6)

[0706] In addition, for example, as an equation corresponding to the above equation (3.16), the following equation (4.7) can also be used.

[0707]

Mathematical formula 28

[0708] predictionsample = (I 1 +I 0 -u×(I y 1 -I y 0 )) >> 1 (4.7)

[0709] In addition, equations (3.11) to (3.14) essentially represent division, so they can also be expressed as the following equation (4.8).

[0710]

Mathematical formula 29

[0711]

[0712] These equations (4.5) to (4.8) are substantially the same as the above equations (3.3), (3.4), (3.10) to (3.14), and (3.16).

[0713] In addition, in the above process, based on the comparison of the first sum and the second sum, a horizontal gradient value or a vertical gradient value is used, but the process of the decoding process is not limited to the above process. It is also possible to determine whether to use a horizontal gradient value or a vertical gradient value according to other coding parameters, etc. Moreover, instead of comparing the first sum and the second sum, a horizontal gradient value can be used to derive the BIO parameter, or a vertical gradient value can be used to derive the BIO parameter. In addition, it is also possible to calculate only one of the first sum and the second sum.

[0714] Even without comparing the first sum and the second sum, the decoding device 200 can reduce substantial multiplications with large computational amounts in the operations performed at each pixel position through the above equations, and can derive multiple parameters for generating a predicted image with low computational amounts. Specifically, in equations (3.2), (3.4), (3.5), (3.10), (4.1), (4.2), (4.5), and (4.6), etc., the positive and negative signs are changed, but substantial multiplication operations are not used. Therefore, the number of substantial multiplication operations in the BIO process can be significantly reduced.

[0715] That is, the decoding device 200 can calculate sG with low computational amounts x , sG x dI, sG y and sG y dI. Therefore, the decoding device 200 can reduce the processing amount in decoding.

[0716] In the above description, decoding processing is shown, but the same processing as above can also be applied to encoding processing. That is, the decoding in the above description can also be replaced with encoding.

[0717] In addition, the arithmetic expressions described here are just examples. As long as they are arithmetic expressions representing the same processing, the arithmetic expressions can be partially changed, deleted, or added. For example, one or more of the expressions for the first sum, the second sum, and the BIO parameters can be replaced with expressions of other specific examples, or can be replaced with other expressions different from those of other specific examples.

[0718] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0719] [Second Specific Example of BIO]

[0720] Next, a second specific example of the decoding process based on BIO is described. For example, similar to the first specific example, the decoding device 200 calculates BIO parameters based on two motion-compensated reference blocks, and decodes the current block using the calculated BIO parameters.

[0721] In this specific example, prediction samples for decoding the current block are generated through the following equations (5.1) to (5.8).

[0722]

Equation 30

[0723] s1 = ∑ [i,j]∈Ω (I x 1 + I x0 )×(I x 1 +I x 0 ) (5.1)

[0724]

Equation 31

[0725] s2 = ∑ [i,j]∈Ω (I y 1 +I y 0 )×(I y 1 +I y 0 ) (5.2)

[0726]

Equation 32

[0727] s3 = ∑ [i,j]∈Ω (-(I x 1 +I x 0 )×(I 0 -I 1 )) (5.3)

[0728]

Equation 33

[0729] s4 = ∑ [i,j]∈Ω (-(I y 1 +I y 0 )×(I 0 -I 1 )) (5.4)

[0730]

Equation 34

[0731]

[0732]

Equation 35

[0733]

[0734]

Equation 36

[0735]

[0736]

Equation 37

[0737]

[0738] s1 in Equation (5.1) in this specific example and sG in Equation (3.2) in the first specific example xcorrespond. Additionally, s2 in formula (5.2) in this specific example corresponds to sG in formula (3.4) in the first specific example y correspond. Additionally, s3 in formula (5.3) in this specific example corresponds to sG in formula (3.5) in the first specific example x dI corresponds. Additionally, s4 in formula (5.4) in this specific example corresponds to sG in formula (3.10) in the first specific example y dI corresponds.

[0739] Moreover, v in formulas (5.5) to (5.7) in this specific example x and v y respectively correspond to the BIO parameters and correspond to u in formulas (3.9) and (3.14) in the first specific example.

[0740] In the first specific example, in the calculation of sG x 、sG y 、sG x dI, and sG y dI, the change of positive and negative signs is performed, and no substantial multiplication operation is carried out. On the other hand, in this specific example, substantial multiplication operations are carried out for the calculation of s1, s2, s3, and s4. As a result, the amount of computation increases, but prediction samples are generated with higher precision. On the contrary, in the first specific example, the amount of computation is reduced.

[0741] Additionally, the arithmetic expressions described here are just examples. As long as they represent the same processing, the arithmetic expressions can be partially changed, deleted, or added. For example, one or more of the expressions for the first sum, the second sum, and the BIO parameters can be replaced with the expressions of other specific examples, or can be replaced with other expressions different from those of other specific examples.

[0742] Additionally, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0743] [The Third Specific Example of BIO]

[0744] Next, use Figure 52 to illustrate the third specific example of the decoding process based on BIO. For example, similar to the first specific example, the decoding device 200 calculates at least one BIO parameter from two motion-compensated reference blocks, and decodes the current block using the at least one calculated BIO parameter.

[0745] Figure 52 is a flowchart showing the third specific example of the decoding process based on BIO.

[0746] First, the decoding device 200 calculates the first sum (S2001) for the current block using the horizontal gradient value of the first reference block and the horizontal gradient value of the second reference block. The first sum calculation process (S2001) in this specific example can be the same as the first sum calculation process (S1001) in the first specific example. Specifically, the decoding device 200 calculates sG through the following equations (6.1) and (6.2) which are the same as equations (3.1) and (3.2) in the first specific example. x As the first sum.

[0747]

Equation 38

[0748] G x = I x 0 + I x 1 (6.1)

[0749]

Equation 39

[0750] sG x = ∑ [i,j]∈Ω sign(G x [i,j]) × G x [i,j] (6.2)

[0751] In addition, the decoding device 200 calculates the second sum (S2002) for the current block using the vertical gradient value of the first reference block and the vertical gradient value of the second reference block. The second sum calculation process (S2002) in this specific example can be the same as the second sum calculation process (S1002) in the first specific example. Specifically, the decoding device 200 calculates sG through the following equations (6.3) and (6.4) which are the same as equations (3.3) and (3.4) in the first specific example. y As the second sum.

[0752]

Equation 40

[0753] G y = I y 0 + I y 1 (6.3)

[0754]

Equation 41

[0755] sG y = Σ [i,j]∈Ω sign(G y [i,j]) × G y [i,j] (6.4)

[0756] Next, the decoding device 200 determines whether at least one of both the first sum and the second sum being greater than the first value and both the first sum and the second sum being less than the second value holds (S2003). The first value may be greater than the second value, may be less than the second value, or may be the same as the second value. The first value and the second value may also be represented as the first threshold value and the second threshold value, respectively.

[0757] For example, in the case where the first value = 100 and the second value = 100, the determination results based on the first sum and the second sum are as follows.

[0758] First sum = 300, second sum = 50: determination result = false

[0759] First sum = 50, second sum = 50: determination result = true

[0760] First sum = 300, second sum = 300: determination result = true

[0761] First sum = 50, second sum = 300: determination result = false

[0762] When the determination result is true (yes in S2003), the decoding device 200 uses both the horizontal gradient value and the vertical gradient value to determine two BIO parameters for the current block (S2004). Expressions (6.5) to (6.8) represent examples of the arithmetic processing for determining two BIO parameters in this case. In these expressions, the horizontal gradient value and the vertical gradient value are used to calculate two BIO parameters represented by u and v.

[0763]

Equation 42

[0764] sG x dI = ∑ [i,j]∈Ω sign(G x [i, j]) × (I 0 i,j -I 1 i,j ) (6.5)

[0765]

Equation 43

[0766] sG y dI = ∑ [i,j]∈Ω sign(G y [i, j]) × (I 0 i,j -I 1 i,j ) (6.6)

[0767]

Equation 44

[0768] u = Clip(((sG xdI >> ((Bits(sGx) - 6 - 12)) × BIOShift[sG x >> Bits(sG x ) - 6]) >> 27) (6.7)

[0769]

Equation 45

[0770] v = Clip(((sG y dI >> ((Bits(sGy) - 6 - 12)) × BIOShift[sG y >> Bits(sG y ) - 6]) >> 27) (6.8)

[0771] Equations (6.5) and (6.6) are the same as Equations (3.5) and (3.10) in the first specific example respectively. The BIO parameter u is calculated by Equation (6.7) using sG x dI and sG x The BIO parameter v is calculated by Equation (6.8) using sG y dI and sG y The BIO parameters u and v can be calculated using the following Equations (6.9) and (6.10) respectively instead of Equations (6.7) and (6.8).

[0772]

Equation 46

[0773] u = sG x dI >> Bits(sG x ) (6.9)

[0774]

Equation 47

[0775] V = sG y dI >> Bits(sG y ) (6.10)

[0776] In the case where the determination result is false (No in S2003), the decoding device 200 determines one BIO parameter for the current block without using the horizontal gradient value or the vertical gradient value (S2005). In this case, an example of the arithmetic processing for determining one BIO parameter is as follows.

[0777] For example, the decoding device 200 determines whether the first sum is greater than the second sum. In the case where the first sum is greater than the second sum, the decoding device 200 calculates only "u" as the BIO parameter by Equations (6.5) and (6.7) (or (6.9)). That is, in this case, the decoding device 200 determines the BIO parameter for the current block without using the vertical gradient value.

[0778] On the other hand, when the first sum is less than or equal to the second sum, the decoding device 200 calculates only "v" as the BIO parameter by equations (6.6) and (6.8) (or (6.10)). That is, in this case, the decoding device 200 determines the BIO parameter for the current block without using the horizontal gradient value.

[0779] Alternatively, for example, the decoding device 200 may also determine whether the first sum is greater than the first value and the second sum is less than the second value.

[0780] Then, when the first sum is greater than the first value and the second sum is less than the second value, the decoding device 200 may determine only "u" as the BIO parameter for the current block without using the vertical gradient value. When the first sum is less than or equal to the first value or the second sum is greater than or equal to the second value, the decoding device 200 may determine only "v" as the BIO parameter for the current block without using the horizontal gradient value.

[0781] Finally, the decoding device 200 decodes the current block using at least one BIO parameter (S2006). Specifically, the decoding device 200 uses at least one of the two BIO parameters u and v to generate a prediction sample, and uses the prediction sample to decode the current block. Equations (6.11) to (6.13) represent multiple examples of the arithmetic processing for generating the prediction sample.

[0782]

Equation 48

[0783] prediction sample = (I 0 +I 1 +u×(I x 0 -I x 1 )+v×(I y 0 -I y 1 )) >> 1 (6.11)

[0784]

Equation 49

[0785] prediction sample = (I 0 +I 1 +u×(I x 0 -I x 1 )) >> 1 (6.12)

[0786]

Equation 50

[0787] prediction sample = (I 0 +I 1 +v×(Iy 0 -I y 1 ))>1 (6.13)

[0788] When the result of the determination process (S2003) based on the first sum, the second sum, the first value, and the second value is true, Equation (6.11) is used. When the result of the determination process (S2003) is false, Equation (6.12) is used when only "u" is calculated. When the result of the determination process (S2003) is false, Equation (6.13) is used when only "v" is calculated.

[0789] The decoding device 200 can also repeatedly perform the above processes (S2001 to S2006) for all sub-blocks of the current CU.

[0790] In addition, a determination process different from the above determination process (S2003) can also be used. For example, the decoding device 200 can also determine whether both the first sum and the second sum are greater than the first value.

[0791] Moreover, when both the first sum and the second sum are greater than the first value, the decoding device 200 can use both the horizontal gradient value and the vertical gradient value to determine two BIO parameters for the current block. When at least one of the first sum and the second sum is not greater than the first value, the decoding device 200 can determine one BIO parameter for the current block without using the horizontal gradient value or the vertical gradient value.

[0792] Alternatively, for example, the decoding device 200 can also determine whether the first sum is greater than the first value and whether the second sum is greater than the second value.

[0793] When the first sum is greater than the first value and the second sum is greater than the second value, the decoding device 200 can use both the horizontal gradient value and the vertical gradient value to determine two BIO parameters for the current block. When the first sum is not greater than the first value or the second sum is not greater than the second value, the decoding device 200 can determine one BIO parameter for the current block without using the horizontal gradient value or the vertical gradient value.

[0794] Alternatively, for example, the decoding device 200 can also determine whether at least one of the two conditions, that is, the condition that the first sum is greater than the first value and the second sum is greater than the second value, and the condition that the first sum is less than the third value and the second sum is less than the fourth value, holds.

[0795] Then, when at least one of the two conditions is satisfied, the decoding device 200 can use both the horizontal gradient value and the vertical gradient value to determine two BIO parameters for the current block. When neither of the two conditions is satisfied, the decoding device 200 can determine one BIO parameter for the current block without using the horizontal gradient value or the vertical gradient value.

[0796] The decoding device 200 can improve the accuracy of the predicted samples in the current block by using BIO. In addition, the decoding device 200 sometimes uses only one of the horizontal gradient value and the vertical gradient value based on conditions. Therefore, the decoding device 200 can suppress an increase in the amount of calculation.

[0797] In addition, the above formula is an example, and the formula for calculating the BIO parameter is not limited to the above formula. For example, the plus and minus signs included in the above formula can be appropriately changed, or a formula equivalent to the above formula can be used. Specifically, as formulas corresponding to the above formulas (6.1) and (6.2), the following formula (7.1) can also be used.

[0798]

Equation 51

[0799] sG x =Σ [i,j]∈Ω abs(I x 1 +I x 0 ) (7.1)

[0800] In addition, for example, as a formula corresponding to the above formula (6.5), the following formula (7.2) can also be used.

[0801]

Equation 52

[0802] sG x dI=Σ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 )) (7.2)

[0803] In addition, for example, as a formula corresponding to the above formula (6.12), the following formula (7.3) can also be used.

[0804]

Equation 53

[0805] prediction sample=(I 1 +I 0 -u×(I x 1 -I x0 )) >> 1 (7.3)

[0806] In addition, equations (6.7) and (6.9) essentially represent division, so they can also be expressed as in the following equation (7.4).

[0807]

Mathematical formula 54

[0808]

[0809] These equations (7.1) to (7.4) are substantially the same as the above equations (6.1), (6.2), (6.5), (6.7), (6.9), and (6.12).

[0810] Similarly, specifically, as equations corresponding to the above equations (6.3) and (6.4), the following equation (7.5) can also be used.

[0811]

Mathematical formula 55

[0812] sG y = ∑ [i,j]∈Ω abs(I y 1 + I y 0 )) (7.5)

[0813] In addition, for example, as an equation corresponding to the above equation (6.6), the following equation (7.6) can also be used.

[0814]

Mathematical formula 56

[0815] sG y dI = ∑ [i,j]∈Ω (-sign(I y 1 + I y 0 ) × (I 0 - I1)) (7.6)

[0816] In addition, for example, as an equation corresponding to the above equation (6.13), the following equation (7.7) can also be used.

[0817]

Mathematical formula 57

[0818] prediction sample = (I 1 + I 0 - v × (I y 1 - I y 0 )) >> 1 (7.7)

[0819] In addition, Equations (6.8) and (6.10) essentially represent division, so they can also be expressed as Equation (7.8) below.

[0820]

Mathematical Formula 58

[0821]

[0822] Equations (7.5) to (7.8) are substantially the same as the above-mentioned Equations (6.3), (6.4), (6.6), (6.8), (6.10), and (6.13). In addition, for example, as an equation corresponding to the above-mentioned Equation (6.11), Equation (7.9) below can also be used. Equation (7.9) below is substantially the same as the above-mentioned Equation (6.11).

[0823]

Mathematical Formula 59

[0824] prediction sample=(I 0 +I 1 -u×(I x 1 -I x 0 )-v×(I y 1 -I y 0 ))>>1 (7.9)

[0825] In addition, in the above process, based on the first sum and the second sum, at least one of the horizontal gradient value and the vertical gradient value is used, but the process of the decoding process is not limited to the above process. It is possible to determine whether to use the horizontal gradient value, the vertical gradient value, or both according to other coding parameters, etc.

[0826] Moreover, regardless of the first sum or the second sum, at least one BIO parameter can be derived using the horizontal gradient value, can be derived using the vertical gradient value, or can be derived using both the horizontal gradient value and the vertical gradient value.

[0827] Regardless of the first sum or the second sum, according to the above equations, the decoding device 200 can reduce the substantial multiplications with large computational amounts in the operations performed for each pixel position, and can derive a plurality of parameters for generating a prediction image with a low computational amount. Specifically, in Equations (6.2), (6.4), (6.5), (6.6), (7.1), (7.2), (7.5), and (7.6), etc., the positive and negative signs are changed, but no substantial multiplication operations are used. Therefore, the number of substantial multiplication operations in the BIO process can be significantly reduced.

[0828] That is, the decoding device 200 can calculate sG x 、sGx dI, sG y and sG y dI. Therefore, the decoding device 200 can reduce the processing amount in decoding. In particular, regardless of the first sum or the second sum, the decoding device 200 can derive at least one BIO parameter using both the horizontal gradient value and the vertical gradient value. Therefore, the decoding device 200 can appropriately generate a prediction image using both the horizontal gradient value and the vertical gradient value while reducing the processing amount in decoding.

[0829] In the above, the decoding process is shown, but the same process as above can also be applied to the encoding process. That is, the decoding in the above description can be replaced with encoding.

[0830] In addition, the formula for calculating the BIO parameter can be replaced with other formulas as long as it is a formula for calculating the BIO parameter using the horizontal gradient value or the vertical gradient value.

[0831] For example, in the third specific example, the formula for calculating the BIO parameter u only needs to be a formula for calculating the BIO parameter u based on the horizontal gradient value instead of the vertical gradient value, and is not limited to formulas (6.7), (6.9), and (7.4), etc. In addition, the formula for calculating the BIO parameter v only needs to be a formula for calculating the BIO parameter v based on the vertical gradient value instead of the horizontal gradient value, and is not limited to formulas (6.8), (6.10), and (7.8), etc.

[0832] In addition, in the first specific example and the third specific example, the first sum corresponds to the horizontal gradient value, and the second sum corresponds to the vertical gradient value. However, this order can be exchanged. That is, the first sum can correspond to the vertical gradient value, and the second sum can correspond to the horizontal gradient value.

[0833] In addition, the arithmetic expressions described here are just examples, and as long as they are arithmetic expressions representing the same process, the arithmetic expressions can also be partially changed, deleted, or added. For example, one or more of the formulas for the first sum, the second sum, and the BIO parameter can be replaced with the formulas of other specific examples, or can be replaced with other formulas different from the formulas of other specific examples.

[0834] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0835] [Fourth Specific Example of BIO]

[0836] The calculation process of BIO parameters and the generation process of the prediction image in the first specific example, the second specific example, and the third specific example are one example, and other calculation processes and generation processes can also be applied. For example, the processes shown in the flowchart of Figure 53 can be applied.

[0837] Figure 53 is a flowchart showing a fourth specific example of the decoding process based on BIO. In the above-mentioned multiple specific examples, as Figure 49 and Figure 52 show, the decoding device 200 switches the method of deriving the prediction sample based on the size of the first sum and the size of the second sum. On the other hand, in this specific example, the decoding device 200 always calculates the optical flow components in the vertical and horizontal directions respectively to derive the prediction sample. Thus, the prediction accuracy may be further improved.

[0838] Specifically, in the example of Figure 53 , similar to the example of Figure 49 and the example of Figure 52 , the decoding device 200 calculates the first sum (S2001) for the current block using the horizontal gradient value of the first reference block and the horizontal gradient value of the second reference block. In addition, the decoding device 200 calculates the second sum (S2002) for the current block using the vertical gradient value of the first reference block and the vertical gradient value of the second reference block.

[0839] Moreover, in the example of Figure 53 , regardless of the size of the first sum and the size of the second sum, the decoding device 200 uses both the horizontal gradient value and the vertical gradient value to determine the BIO parameter (S2004) for the current block. The operation of using both the horizontal gradient value and the vertical gradient value to determine the BIO parameter for the current block can be the same as that in the third specific example. For example, the decoding device 200 can use the formulas (6.1) to (6.10) and the like described in the third specific example as the arithmetic expressions for determining the BIO parameter.

[0840] Moreover, the decoding device 200 decodes the current block using the BIO parameter (S2006). For example, the decoding device 200 uses two BIO parameters u and v to generate the prediction sample. At this time, the decoding device 200 can also derive the prediction sample through formulas such as (6.11). The decoding device 200 can also use the formulas described in the first specific example or other formulas. Then, the decoding device 200 decodes the current block using the prediction sample.

[0841] In addition, sign(x) appearing in formulas (6.2), (6.4), (6.5), and (6.6) etc. can be defined by the aforementioned binary formula (a), or can be defined by the following formula (b).

[0842]

Equation 60

[0843]

[0844] The sign function of formula (a) returns a value indicating whether the argument provided to the sign function is positive or negative. The sign function of formula (b) returns a value indicating whether the argument provided to the sign function is positive, negative, or zero.

[0845] In the original formula for deriving the optical flow, sign(G x [i, j]) and sign(G y [i, j]) in formulas (6.2), (6.4), (6.5), and (6.6) are G x [i, j] and G y [i, j] respectively. Therefore, for example, when G x [i, j] = 0, calculate G x [i, j] × (I 0 i,j -I 1 i,j ) = 0 as an intermediate value. When G y [i, j] = 0, calculate G y [i, j] × (I 0 i,j -I 1 i,j ) = 0 as an intermediate value.

[0846] However, in the simplified formula for deriving the optical flow, when sign(0) = 1, an appropriate intermediate value cannot be obtained. For example, in formula (6.5), when G x [i, j] = 0, calculate sign(G x [i, j]) × (I 0 i,j -I 1 i,j ) = (I 0 i,j -I 1 i,j ) as an intermediate value. Additionally, in formula (6.6), when G y [i, j] = 0, calculate sign(G y [i, j]) × (I 0 i,j -I 1 i,j ) = (I 0 i,j -I 1 i,j ) as an intermediate value. That is, the intermediate value is not 0, but a value different from 0 remains.

[0847] In the definition of sign in the above formula (b), when G x [i, j] = 0, calculate sign(G x [i, j]) × (I 0 i,j -I 1 i,j ) = 0 as an intermediate value, and when G y [i, j] = 0, calculate sign(G y [i, j]) × (I 0 i,j -I 1 i,j ) = 0 as an intermediate value. Therefore, in these cases, calculate the same intermediate value as the original optical flow derivation formula.

[0848] Therefore, sign defined by the above formula (b) has a value more similar to the formula of the original optical flow than sign defined by the above formula (a). Therefore, the prediction accuracy may become higher.

[0849] In addition, the above changes can also be combined throughout this disclosure. For example, the definition of sign(x) in the first specific example can be replaced with the above formula (b), and the definition of sign(x) in the third specific example can be replaced with the above formula (b).

[0850] In addition, in the above, decoding processing is shown, but the same processing as above can also be applied to encoding processing. That is, decoding in the above description can also be replaced with encoding.

[0851] In addition, for example, in the above description, as an example of a filter for obtaining the horizontal gradient value and the vertical gradient, a filter having a filter coefficient set of [-1, 0, 1] is used, but a filter having a filter coefficient set of [1, 0, -1] can also be used, etc.

[0852] In addition, formula (b) is an example of a sign function that returns a value indicating whether the independent variable is positive, negative, or zero. A sign function that returns a value indicating whether the independent variable is positive, negative, or zero can be represented by other formulas that obtain three values according to whether the independent variable is positive, negative, or zero.

[0853] [The fifth specific example of BIO]

[0854] Next, the fifth specific example of the decoding process based on BIO is described. In the fifth specific example, in the same way as the fourth specific example, the optical flow components in the vertical direction and the horizontal direction are always obtained to derive the prediction sample. The following operation formulas are the operation formulas in the fifth specific example.

[0855]

Equation 61

[0856] G x =I x 0 +I x 1 (8.1)

[0857]

Equation 62

[0858] sG x =Σ [i,j]∈Ω sign(G x [i,j])×G x [i,j] (8.2)

[0859]

Equation 63

[0860] G y =I y 0 +I y 1 (8.3)

[0861]

Equation 64

[0862] sG y =∑ [i,j]∈Ω sign(G y [i,j])×G y [i,j] (8.4)

[0863]

Equation 65

[0864] sG x dI=∑ [i,j]∈Ω sign(G x [i,j])×(I 0 i,j -I 1 i,j ) (8.5)

[0865]

Equation 66

[0866] sG y dI=∑ [i,j]∈Ω sign(G y [i,j])×(I 0 i,j -I 1 i,j ) (8.6)

[0867]

Equation 67

[0868] sG x G y =∑ [i,j]∈Ω sign(Gy [i,j]) × G x [i,j] (8.7)

[0869]

Mathematical Formula 68

[0870] u = sG x dI >> Bits(sG x ) (8.8)

[0871]

Mathematical Formula 69

[0872] v = (sG y dI - u × sG x G y ) >> Bits(sG y ) (8.9)

[0873]

Mathematical Formula 70

[0874] prediction sample = (I 0 + I 1 + u × (I x 0 - I x 1 ) + v × (I y 0 - I y 1 )) >> 1 (8.10)

[0875] The above formulas (8.1) to (8.6), (8.8), and (8.10) are the same as formulas (6.1) to (6.6), (6.9), and (6.11) in the third specific example. In this specific example, formula (8.7) is added, and formula (6.10) is replaced with formula (8.9). That is, by means of arithmetic processing depending on the BIO parameter "u" in the x direction and the parameter "sG x G y " related to the gradients in the x direction and y direction, the BIO parameter "v" in the y direction is derived. Thereby, it is possible to derive BIO parameters with higher precision, and the possibility of improving the coding efficiency becomes higher. Note that sG x G y can also be expressed as the third sum.

[0876] In addition, formula (b) described in the fourth specific example can be used as sign(x). As defined by formula (b), by sign(x) corresponding to three values, it is possible to further improve the coding efficiency. In addition, as sign(x), formula (a) described in the first specific example can also be used. Therefore, compared with the case where sign(x) corresponds to three values, the formula is simplified, so it is possible to improve the coding efficiency while reducing the processing load.

[0877] In addition, the arithmetic expressions described here are just examples. As long as they represent the same processing, the arithmetic expressions can be partially modified, deleted, or added. For example, one or more of the expressions in the multiple expressions of the first sum, the second sum, the third sum, and the BIO parameter can be replaced with expressions of other specific examples, or can be replaced with other expressions different from those of other specific examples.

[0878] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0879] In addition, this arithmetic expression is not only applicable to the Figure 53 flowchart in the fourth specific example, but can also be applicable to the Figure 49 flowchart in the first specific example, the Figure 52 flowchart in the third specific example, or other flowcharts.

[0880] In addition, in the above, the decoding process is shown, but the same processing as above can also be applied to the encoding process. That is, the decoding in the above description can also be replaced with encoding.

[0881] [Sixth Specific Example of BIO]

[0882] Next, the sixth specific example of the decoding process based on BIO is described. In the sixth specific example, similar to the fourth specific example, etc., the optical flow components in the vertical direction and the horizontal direction are always obtained to derive the predicted samples. The following arithmetic expressions are the arithmetic expressions in the sixth specific example.

[0883]

Equation 71

[0884] G x =I x 0 +I x 1 (9.1)

[0885]

Equation 72

[0886] sG x =∑ [i,j]∈Ω G x [i,j]×G x [i,j] (9.2)

[0887]

Equation 73

[0888] G y =I y 0 +I y1 (9.3)

[0889]

Equation 74

[0890] sG y =Σ [i,j]∈Ω G y [i,j]×G y [i,j] (9.4)

[0891]

Equation 75

[0892] sG x dI=∑ [i,j]∈Ω G x [i,j]×(I 0 i,j -I 1 i,j ) (9.5)

[0893]

Equation 76

[0894] sG y dI=∑ [i,j]∈Ω G y [i,j]×(I 0 i,j -I 1 i,j ) (9.6)

[0895]

Equation 77

[0896] u=sG x dI>>Bits(sG x ) (9.7)

[0897]

Equation 78

[0898] v=sG y dI>>Bits(sG u ) (9.8)

[0899]

Equation 79

[0900] prediction sample=(I 0 +I 1 +u×(I x 0 -I x 1 )+v×(I y 0 -I y 1 ))>>1 (9.9)

[0901] The above formulas (9.1), (9.3), and (9.7) - (9.9) are the same as formulas (6.1), (6.3), and (6.9) - (6.11) in the third specific example. In this specific example, formulas (6.2) and (6.4) - (6.6) are respectively replaced by formulas (9.2) and (9.4) - (9.6).

[0902] For example, in formulas (6.2) and (6.4) - (6.6), instead of G x [i, j] and G y [i, j], sign(G x [i, j]) and sign(G y [i, j]) are used to eliminate the substantial multiplication operation. On the other hand, in formulas (9.2) and (9.4) - (9.6), instead of sign(G x [i, j]) and sign(G y [i, j]), the values of G x [i, j] and G y [i, j] are directly used. That is, the substantial multiplication operation is used.

[0903] As a result, the amount of arithmetic processing increases. However, it is possible to derive BIO parameters with higher precision, and the possibility of improving the coding efficiency becomes higher.

[0904] In addition, the arithmetic expressions described here are just examples. As long as they represent the same processing, the arithmetic expressions can be partially modified, deleted, or added. For example, one or more of the expressions for the first sum, the second sum, the third sum, and the BIO parameters can be replaced with the expressions of other specific examples, or can be replaced with other expressions different from those of other specific examples.

[0905] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0906] In addition, the arithmetic expressions corresponding to this specific example are not only applied to the flowchart in the fourth specific example Figure 53 , but can also be applied to the flowchart in the first specific example Figure 49 , the flowchart in the third specific example Figure 52 , or other flowcharts.

[0907] In addition, in the above, the decoding process is shown, but the same process as above can also be applied to the encoding process. That is, the decoding in the above description can be replaced with encoding.

[0908] [The Seventh Specific Example of BIO]

[0909] Next, a seventh specific example of the decoding process based on BIO will be described. In the seventh specific example, in the same manner as in the fourth specific example and others, the optical flow components in the vertical and horizontal directions are always obtained to derive the predicted samples. The following arithmetic expressions are those in the seventh specific example.

[0910]

Equation 80

[0911] G x = I x 0 + I x 1 (10.1)

[0912]

Equation 81

[0913] sG x = ∑ [i,j]∈Ω abs(G x [i,j]) (10.2)

[0914]

Equation 82

[0915] G y = I y 0 + I y 1 (10.3)

[0916]

Equation 83

[0917] sG y = ∑ [i,j]∈Ω abs(G y [i,j]) (10.4)

[0918]

Equation 84

[0919] sG x dI = ∑ [i,j]∈Ω sign(G x [i,j])×(I 0 i,j - I 1 j,j ) (10.5)

[0920]

Equation 85

[0921] sG y dI = ∑ [i,j]∈Ω sign(G y [i,j])×(I 0 i,j - I 1 i,j ) (10.6)

[0922]

Equation 86

[0923] sG x G y = ∑ [i,j]∈Ω sign(G y [i, j]) × G x [i, j] (10.7)

[0924]

Equation 87

[0925] u = sG x dI > Bits(sG x ) (10.8)

[0926]

Equation 88

[0927] v = (sG y dI - u × sG x G y ) >> Bits(sG y ) (10.9)

[0928]

Equation 89

[0929] prediction sample = (I 0 + I 1 + u × (I x 0 - I x 1 ) + v × (I y 0 - I y 1 )) >> 1 (10.10)

[0930] The above equations (10.1), (10.3), and (10.5) - (10.10) are the same as equations (8.1), (8.3), and (8.5) - (8.10) in the fifth specific example. In this specific example, equations (8.2) and (8.4) are replaced by equations (10.2) and (10.4) respectively. Specifically, instead of using the sign function to change the positive and negative signs, the abs function is used.

[0931] The results obtained by using the abs function as in equations (10.2) and (10.4) are the same as the results obtained by using the sign function to change the positive and negative signs as in equations (8.2) and (8.4). That is, this specific example is substantially the same as the fifth specific example.

[0932] In the fifth specific example, the product of the positive and negative signs of G x is used, and the positive and negative signs of G x is used, and the positive and negative signs of G y is used, and the positive and negative signs of Gy The product. In this specific example, the part of the product is replaced with the absolute value. Thus, low processing can be achieved.

[0933] In addition, the positive and negative signs included in the above formula can be appropriately changed, and an equivalent formula to the above formula can also be used. Specifically, the following formulas (11.1) to (11.5) corresponding to the above formulas (10.1) to (10.7) can be used.

[0934]

Equation 90

[0935] sG x = ∑ [i,j]∈Ω abs(I x 1 + I x 0 )(11.1)

[0936]

Equation 91

[0937] sG y = ∑ [i,j]∈Ω abs(I y 1 + I y 0 )(11.2)

[0938]

Equation 92

[0939] sG x dI = ∑ [i,j]∈Ω (-sign(I x 1 + I x 0 ) × (I 0 - I 1 ))(11.3)

[0940]

Equation 93

[0941] sG y dI = ∑ [i,j]∈Ω (-sign(I y 1 + I y 0 ) × (I 0 - I 1 ))(11.4)

[0942]

Equation 94

[0943] sG x G y = ∑ [i,j]∈Ω (sign(I y 1 + Iy 0 )×(I x 1 +I x 0 )) (11.5)

[0944] In addition, since the expressions (10.8) and (10.9) essentially represent division, they can also be expressed as in the following expressions (11.6) and (11.7).

[0945]

Mathematical formula 95

[0946]

[0947]

Mathematical formula 96

[0948]

[0949] In addition, the arithmetic expressions described here are just examples. As long as they represent the same processing, the arithmetic expressions can be partially changed, deleted, or added. For example, one or more of the expressions among the expressions of the first sum, the second sum, the third sum, and the BIO parameters can be replaced with expressions of other specific examples, or can also be replaced with other expressions different from those of other specific examples.

[0950] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] was used, but a filter with a filter coefficient set of [1, 0, -1] can also be used, etc.

[0951] In addition, the arithmetic expressions corresponding to this specific example are not only applied to the Figure 53 flowchart in the fourth specific example, but can also be applied to the Figure 49 flowchart in the first specific example, the Figure 52 flowchart in the third specific example, or other flowcharts.

[0952] In addition, in the above, the decoding process was shown, but the same process as above can also be applied to the encoding process. That is, the decoding in the above description can also be replaced with encoding.

[0953] [Eighth specific example of BIO]

[0954] Next, the eighth specific example of the decoding process based on BIO will be described. In the eighth specific example, similar to the fourth specific example, etc., the optical flow components in the vertical direction and the horizontal direction are always obtained to derive the predicted samples. The following arithmetic expressions are the arithmetic expressions in the eighth specific example.

[0955]

Mathematical formula 97

[0956] Gx = I x 0 + I x 1 (12.1)

[0957]

Equation 98

[0958] sG x = ∑ [i,j]∈Ω G x [i,j] (12.2)

[0959]

Equation 99

[0960] G y = I y 0 + I y 1 (12.3)

[0961]

Equation 100

[0962] sG y = ∑ [i,j]∈Ω G y [i,j] (12.4)

[0963]

Equation 101

[0964] sG x dI = ∑ [i,j]∈Ω (I 0 i,j - I 1 i,j ) (12.5)

[0965]

Equation 102

[0966] sG y dI = ∑ [i,j]∈Ω (I 0 i,j - I 1 i,j ) (12.6)

[0967]

Equation 103

[0968] sG x G y = ∑ [i,j]∈Ω G x [i,j] (12.7)

[0969]

Equation 104

[0970] u = sG x dI >> EBits(sG x ) (12.8)

[0971]

Equation 105

[0972] v = (sG y dI - u × sG x G y ) >> Bits(sG y ) (12.9)

[0973]

Equation 106

[0974] predictionsample = (I 0 + I 1 + u × (I x 0 - I x 1 ) + v × (I y 0 - I y 1 )) >> 1 (12.10)

[0975] The above equations (12.1), (12.3), and (12.8) - (12.10) are the same as equations (8.1), (8.3), and (8.8) - (8.10) in the fifth specific example.

[0976] On the other hand, regarding equations (8.2) and (8.4) - (8.7), assuming that the values representing the positive and negative signs of G x and G y are always 1, equations (8.2) and (8.4) - (8.7) are respectively replaced by equations (12.2) and (12.4) - (12.7). This is based on the assumption that in the small region Ω, both the absolute value and the positive and negative signs of the gradient of the pixel values are constant.

[0977] Since it is assumed that the positive and negative signs are constant, the processing of calculating the positive and negative signs of G x and G y in pixel units can be reduced. In addition, the arithmetic expressions of sG x are equal to the arithmetic expressions of sG x G y . Thus, the processing can be further reduced.

[0978] In addition, the arithmetic expressions described here are just examples. As long as they represent the same processing, the arithmetic expressions can be partially changed, deleted, or added. For example, one or more of the expressions in the first sum, the second sum, the third sum, and the multiple expressions of the BIO parameter can be replaced by the expressions of other specific examples, or can be replaced by other expressions different from those of other specific examples.

[0979] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter having a filter coefficient set of [-1, 0, 1] was used, but a filter having a filter coefficient set of [1, 0, -1] etc. may also be used.

[0980] In addition, this arithmetic expression is applicable not only to the Figure 53 flowchart in the 4th specific example, but also to the Figure 49 flowchart in the 1st specific example, the Figure 52 flowchart in the 3rd specific example, or other flowcharts.

[0981] In addition, in the above, decoding processing was shown, but the same processing as above can also be applied to encoding processing. That is, the decoding in the above description can also be replaced with encoding.

[0982] [9th Specific Example of BIO]

[0983] Next, the 9th specific example of the decoding processing based on BIO will be described. In the 9th specific example, similar to the 4th specific example etc., the optical flow components in the vertical direction and the horizontal direction are always obtained to derive the predicted samples. The following arithmetic expressions are the arithmetic expressions in the 9th specific example.

[0984] [Equation 107]

[0985] G x = I x 0 + I x 1 (13.1)

[0986] [Equation 108]

[0987] sG x = ∑ [i,j]∈Ω G x [i,j] (13.2)

[0988] [Equation 109]

[0989] G y = I y 0 + I y 1 (13.3)

[0990] [Equation 110]

[0991] sG y = ∑ [i,j]∈Ω G y [i,j] (13.4)

[0992] [Equation 111]

[0993] sG x dI = ∑ [i,j]∈Ω (I 0 i,j -I 1 i,j ) (13.5)

[0994]

Equation 112

[0995] sG y dI = ∑ [i,j]∈Ω (I 0 i,j -I 1 i,j ) (13.6)

[0996]

Equation 113

[0997] u = sG x dI >> Bits(sG x ) (13.7)

[0998]

Equation 114

[0999] v = (sG y dI - u × sG x ) >> Bits(sG y ) (13.8)

[1000]

Equation 115

[1001] prediction sample = (I 0 +I 1 +u × (I x 0 -I x 1 )+v × (I y 0 -I y 1 )) >> 1 (13.9)

[1002] The above formulas (13.1) to (13.9) are the same as formulas (12.1) to (12.6) and (12.8) to (12.10) in the 8th specific example. In the 8th specific example, based on the case where the operation formula of sG x and the operation formula of sG x G y are equal, in this specific example, the mutually related derivation processes in the 3rd and are deleted. Then, sG x is used in the derivation of v.

[1003] In addition, the arithmetic expressions described here are just examples. As long as the arithmetic expressions represent the same processing, they can be partially modified, deleted, or added. For example, one or more of the expressions in the multiple expressions of the first sum, the second sum, and the BIO parameter can be replaced with expressions of other specific examples, or can be replaced with other expressions different from those of other specific examples.

[1004] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter with a filter coefficient set of [-1, 0, 1] is used, but a filter with a filter coefficient set of [1, 0, -1] can also be used.

[1005] In addition, the arithmetic expression corresponding to this specific example is not only applied to the Figure 53 flowchart in the fourth specific example, but can also be applied to the Figure 49 flowchart in the first specific example, the Figure 52 flowchart in the third specific example, or other flowcharts.

[1006] In addition, in the above, the decoding process is shown, but the same process as above can also be applied to the encoding process. That is, the decoding in the above description can also be replaced with encoding.

[1007] [Tenth Specific Example of BIO]

[1008] This specific example represents a variant that can be applied to other specific examples. For example, a lookup table can be used for division in each specific example. For example, a lookup table (divSigTable) with 6 bits, that is, 64 entries, can be used. Thus, equations (8.8) and (8.9) in the fifth specific example can be replaced with the following equations (14.1) and (14.2) respectively. The same kind of equations in other specific examples can also be replaced in the same way.

[1009]

Equation 116

[1010] u = sG x dI × (divSigTalbe[Upper6digits(sG x )][32) >> [log(sG x )] (14.1)

[1011]

Equation 117

[1012]

[1013] Upper6digits in the above equations (14.1) and (14.2) represents the following.

[1014]

Equation 118

[1015]

[1016] In addition, a gradient image is obtained by obtaining gradient values for a plurality of pixels. The gradient values can be derived, for example, by applying a gradient filter to the plurality of pixels. By increasing the number of taps of the gradient filter, it is possible to improve the accuracy of the gradient values, improve the accuracy of the prediction image, and improve the coding efficiency.

[1017] On the other hand, since the processing is performed for each pixel position, the amount of calculation increases when there are many pixel positions to be processed. Therefore, a 2-tap filter can be used as the gradient filter. That is, the gradient value can be the difference value between two pixels above, below, to the left, or to the right of the pixel for which the gradient value is calculated. Alternatively, the gradient value can also be the difference value between the pixel for which the gradient value is calculated and a pixel adjacent to the pixel for which the gradient value is calculated (specifically, pixels such as above, below, left, or right). Therefore, the amount of processing may decrease compared to the case where the number of taps is large.

[1018] In addition, the plurality of pixels for which the gradient values are calculated can be a plurality of integer pixels or can include fractional pixels.

[1019] In addition, for example, in the above description, as an example of the filter for obtaining the horizontal gradient value and the vertical gradient, a filter having a filter coefficient set of [-1, 0, 1] was used, but a filter having a filter coefficient set of [1, 0, -1] etc. can also be used.

[1020] [Representative examples of structure and processing]

[1021] The following shows representative examples of the structure and processing of the above-described encoding device 100 and decoding device 200. This representative example mainly corresponds to the above-described 5th specific example, 7th specific example, etc.

[1022] Figure 54 is a flowchart showing the operations performed by the encoding device 100. For example, the encoding device 100 includes a circuit and a memory connected to the circuit. The circuit and memory included in the encoding device 100 can also correspond to Figure 40 the processor a1 and memory a2 shown. The circuit of the encoding device 100 performs the following during operation.

[1023] For example, the circuit of the encoding device 100 derives, for each relative pixel position, the absolute value of the sum of the horizontal gradient value at the relative pixel position in the first range and the horizontal gradient value at the relative pixel position in the second range, that is, the horizontal gradient sum absolute value (S3101).

[1024] Here, the first range includes the first reference block of the current block, and the second range includes the second reference block of the current block. Each relative pixel position is a pixel position that is commonly and relatively determined with respect to both the first range and the second range, and is a pixel position in each of the first range and the second range.

[1025] In addition, a pixel position that is commonly and relatively determined with respect to both the first range and the second range means a pixel position that is relatively and identically determined with respect to both the first range and the second range. For example, when calculating one horizontal gradient and absolute value, the horizontal gradient values of pixel positions that are relatively the same within the first range and the second range are used. Specifically, for example, the horizontal gradient value of the pixel position at the top-leftmost position within the first range and the horizontal gradient value of the pixel position at the top-leftmost position within the second range are used to derive one horizontal gradient and absolute value.

[1026] Then, the circuit of the encoding device 100 derives the sum of the multiple horizontal gradients and absolute values respectively derived for multiple relative pixel positions as the first parameter (S3102).

[1027] In addition, the circuit of the encoding device 100 derives, for each relative pixel position, a vertical gradient and absolute value, which is the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range (S3103).

[1028] Then, the circuit of the encoding device 100 derives the sum of the multiple vertical gradients and absolute values respectively derived for multiple relative pixel positions as the second parameter (S3104).

[1029] In addition, the circuit of the encoding device 100 derives, for each relative pixel position, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value (S3105). For example, at this time, the circuit of the encoding device 100 derives, for each relative pixel position, a signed pixel difference value by subtracting one of the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range from the other.

[1030] Then, the circuit of the encoding device 100, for each relative pixel position, reverses or maintains the sign of the pixel difference value derived for the relative pixel position according to the sign of the sum of the horizontal gradients, and derives a horizontally corresponding pixel difference value (S3106). Here, for each relative pixel position, the sum of the horizontal gradients is the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range. The horizontally corresponding pixel difference value is the pixel difference value whose sign is reversed or maintained according to the sign of the sum of the horizontal gradients.

[1031] Then, the circuit of the encoding device 100 derives the sum of a plurality of horizontally corresponding pixel difference values respectively derived for a plurality of relative pixel positions as the third parameter (S3107).

[1032] In addition, for each relative pixel position, the circuit of the encoding device 100 reverses or maintains the positive or negative sign of the pixel difference value derived for the relative pixel position by the positive or negative sign of the sum of the vertical gradients, and derives the vertically corresponding pixel difference value (S3108). Here, for each relative pixel position, the sum of the vertical gradients is the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range. The vertically corresponding pixel difference value is the pixel difference value whose positive or negative sign is reversed or maintained by the positive or negative sign of the sum of the vertical gradients.

[1033] Then, the circuit of the encoding device 100 derives the sum of a plurality of vertically corresponding pixel difference values respectively derived for a plurality of relative pixel positions as the fourth parameter (S3109).

[1034] In addition, for each relative pixel position, the circuit of the encoding device 100 reverses or maintains the positive or negative sign of the sum of the horizontal gradients by the positive or negative sign of the sum of the vertical gradients, and derives the vertically corresponding sum of the horizontal gradients (S3110). Here, the vertically corresponding sum of the horizontal gradients is the sum of the horizontal gradients whose positive or negative sign is reversed or maintained by the positive or negative sign of the sum of the vertical gradients.

[1035] Then, the circuit of the encoding device 100 derives the sum of a plurality of vertically corresponding sums of the horizontal gradients respectively derived for a plurality of relative pixel positions as the fifth parameter (S3111).

[1036] Then, the circuit of the encoding device 100 uses the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter to generate a prediction image used in the encoding of the current block (S3112).

[1037] Accordingly, in the operations performed for each pixel position, it is possible to reduce the substantial multiplication operations with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a prediction image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in encoding. In addition, it is possible to appropriately generate a prediction image based on a plurality of parameters including a parameter related to the horizontal gradient value, a parameter related to the vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[1038] In addition, for example, the circuit of the encoding device 100 can derive the first parameter through the above formula (11.1), and can derive the second parameter through the above formula (11.2). The circuit of the encoding device 100 can also derive the third parameter through the above formula (11.3), and derive the fourth parameter through the above formula (11.4). In addition, the circuit of the encoding device 100 can also derive the fifth parameter through the above formula (11.5).

[1039] Here, Ω represents a set of multiple relative pixel positions, and [i, j] represents each relative pixel position. In addition, for each relative pixel position, I x 0 represents the horizontal gradient value of this relative pixel position in the first range, and I x 1 represents the horizontal gradient value of this relative pixel position in the second range. In addition, for each relative pixel position, I y 0 represents the vertical gradient value of this relative pixel position in the first range, and I y 1 represents the vertical gradient value of this relative pixel position in the second range.

[1040] In addition, for each relative pixel position, I 0 represents the pixel value of this relative pixel position in the first range, and I 1 represents the pixel value of this relative pixel position in the second range. And, abs(I x 1 +I x 0 ) represents the absolute value of I x 1 +I x 0 , sign(I x 1 +I x 0 ) represents the positive or negative sign of I x 1 +I x 0 , abs(I y 1 +I y 0 ) represents the absolute value of I y 1 +I y 0 , sign(I y 1 +I y 0 ) represents the positive or negative sign of I y 1 +Iy 0 plus or minus sign

[1041] Accordingly, it is possible to derive a plurality of parameters with a low amount of computation using pixel values, horizontal gradient values, and vertical gradient values.

[1042] In addition, for example, the circuit of the encoding device 100 can also derive the sixth parameter by dividing the third parameter by the first parameter. Further, the circuit of the encoding device 100 can derive the seventh parameter by subtracting the product of the fifth parameter and the sixth parameter from the fourth parameter and then dividing by the second parameter. Then, the circuit of the encoding device 100 can generate a prediction image using the sixth parameter and the seventh parameter.

[1043] Accordingly, it is possible to appropriately aggregate a plurality of parameters into two parameters corresponding to the horizontal direction and the vertical direction. It is possible to appropriately incorporate a parameter related to the horizontal gradient value into the parameter corresponding to the horizontal direction. It is possible to appropriately incorporate a parameter related to the vertical gradient value, a parameter related to both the horizontal gradient value and the vertical gradient value, and a parameter corresponding to the horizontal direction into the parameter corresponding to the vertical direction. Then, it is possible to generate a prediction image using these two parameters.

[1044] In addition, for example, the circuit of the encoding device 100 can derive the sixth parameter by the above formula (10.8). Further, the circuit of the encoding device 100 can derive the seventh parameter by the above formula (10.9).

[1045] Here, sG x represents the first parameter, sG y represents the second parameter, sG x dI represents the third parameter, sG y dI represents the fourth parameter, sG x G y represents the fifth parameter, and u represents the sixth parameter. Bits is a function that returns a value obtained by rounding up the binary logarithm of the independent variable to an integer.

[1046] Accordingly, it is possible to derive two parameters corresponding to the horizontal direction and the vertical direction with a low amount of computation.

[1047] In addition, for example, the circuit of the encoding device 100 can generate a prediction image by deriving a prediction pixel value using the first pixel value, the first horizontal gradient value, the first vertical gradient value, the second pixel value, the second horizontal gradient value, the second vertical gradient value, the sixth parameter, and the seventh parameter.

[1048] Here, the predicted pixel value is the predicted pixel value of the processing target pixel position included in the current block. The first pixel value is the pixel value of the first pixel position corresponding to the processing target pixel position in the first reference block. The first horizontal gradient value is the horizontal gradient value of the first pixel position. The first vertical gradient value is the vertical gradient value of the first pixel position. The second pixel value is the pixel value of the second pixel position corresponding to the processing target pixel position in the second reference block. The second horizontal gradient value is the horizontal gradient value of the second pixel position. The second vertical gradient value is the vertical gradient value of the second pixel position.

[1049] Thus, it is possible to generate a predicted image using two parameters corresponding to the horizontal direction and the vertical direction, etc., and the two parameters corresponding to the horizontal direction and the vertical direction can be appropriately reflected in the predicted image.

[1050] Also, for example, the circuit of the encoding device 100 can also derive the predicted pixel value by dividing the sum of the first pixel value, the second pixel value, the first correction value, and the second correction value by 2. Here, the first correction value corresponds to the product of the sixth parameter and the difference between the first horizontal gradient value and the second horizontal gradient value, and the second correction value corresponds to the product of the seventh parameter and the difference between the first vertical gradient value and the second vertical gradient value.

[1051] Thus, it is possible to appropriately generate a predicted image using two parameters corresponding to the horizontal direction and the vertical direction, etc.

[1052] In addition, for example, the circuit of the encoding device 100 can also derive the predicted pixel value through the above formula (10.10). Here, I 0 represents the first pixel value, I 1 represents the second pixel value, u represents the sixth parameter, I x 0 represents the first horizontal gradient value, I x 1 represents the second horizontal gradient value, v represents the seventh parameter, I y 0 represents the first vertical gradient value, I y 1 represents the second vertical gradient value. Thus, according to the formula associated with two parameters corresponding to the horizontal direction and the vertical direction, etc., it is possible to appropriately generate a predicted image.

[1053] In addition, for example, the circuit of the encoding device 100 can derive one or more parameters of the bidirectional optical flow using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter, and generate a predicted image using one or more parameters of the bidirectional optical flow and the bidirectional optical flow. Thus, the encoding device 100 can appropriately generate a predicted image.

[1054] One or more parameters of the bidirectional optical flow may also be at least one of the above-described sixth parameter and seventh parameter.

[1055] In addition, the inter-frame prediction unit 126 of the encoding device 100 may also perform the above-described operations as the circuit of the encoding device 100.

[1056] Figure 55 It is a flowchart showing the operations performed by the decoding device 200. For example, the decoding device 200 includes a circuit and a memory connected to the circuit. The circuit and memory included in the decoding device 200 may also correspond to Figure 46 the shown processor b1 and memory b2. The circuit of the decoding device 200 performs the following operations during operation.

[1057] For example, the circuit of the decoding device 200 derives, for each relative pixel position, the absolute value of the sum of the horizontal gradient value of the relative pixel position in the first range and the horizontal gradient value of the relative pixel position in the second range, that is, the horizontal gradient sum absolute value (S3201).

[1058] Here, the first range includes the first reference block of the current block, and the second range includes the second reference block of the current block. Each relative pixel position is a pixel position that is determined relatively with respect to both the first range and the second range, and is a pixel position in each of the first range and the second range.

[1059] Then, the circuit of the decoding device 200 derives the sum of the multiple horizontal gradient sum absolute values respectively derived for the multiple relative pixel positions as the first parameter (S3202).

[1060] In addition, the circuit of the decoding device 200 derives, for each relative pixel position, the absolute value of the sum of the vertical gradient value of the relative pixel position in the first range and the vertical gradient value of the relative pixel position in the second range, that is, the vertical gradient sum absolute value (S3203).

[1061] Then, the circuit of the decoding device 200 derives the sum of the multiple vertical gradient sum absolute values respectively derived for the multiple relative pixel positions as the second parameter (S3204).

[1062] In addition, the circuit of the decoding device 200 derives, for each relative pixel position, the difference between the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range, that is, the pixel difference value (S3205). For example, at this time, the circuit of the decoding device 200 derives, for each relative pixel position, a signed pixel difference value by subtracting one of the pixel value of the relative pixel position in the first range and the pixel value of the relative pixel position in the second range from the other.

[1063] Then, the circuit of the decoding device 200 derives a horizontally corresponding pixel difference value for each relative pixel position by reversing or maintaining the sign of the pixel difference value derived for that relative pixel position based on the sign of the sum of horizontal gradients (S3206). Here, for each relative pixel position, the sum of horizontal gradients is the sum of the horizontal gradient value of that relative pixel position in the first range and the horizontal gradient value of that relative pixel position in the second range. The horizontally corresponding pixel difference value is the pixel difference value whose sign is reversed or maintained based on the sign of the sum of horizontal gradients.

[1064] Then, the circuit of the decoding device 200 derives the sum of the multiple horizontally corresponding pixel difference values respectively derived for multiple relative pixel positions as the third parameter (S3207).

[1065] In addition, the circuit of the decoding device 200 derives a vertically corresponding pixel difference value for each relative pixel position by reversing or maintaining the sign of the pixel difference value derived for that relative pixel position based on the sign of the sum of vertical gradients (S3208). Here, for each relative pixel position, the sum of vertical gradients is the sum of the vertical gradient value of that relative pixel position in the first range and the vertical gradient value of that relative pixel position in the second range. The vertically corresponding pixel difference value is the pixel difference value whose sign is reversed or maintained based on the sign of the sum of vertical gradients.

[1066] Then, the circuit of the decoding device 200 derives the sum of the multiple vertically corresponding pixel difference values respectively derived for multiple relative pixel positions as the fourth parameter (S3209).

[1067] In addition, the circuit of the decoding device 200 derives a vertically corresponding sum of horizontal gradients for each relative pixel position by reversing or maintaining the sign of the sum of horizontal gradients based on the sign of the sum of vertical gradients (S3210). Here, the vertically corresponding sum of horizontal gradients is the sum of horizontal gradients whose sign is reversed or maintained by using the sign of the sum of vertical gradients.

[1068] Then, the circuit of the decoding device 200 derives the sum of the multiple vertically corresponding sums of horizontal gradients respectively derived for multiple relative pixel positions as the fifth parameter (S3211).

[1069] Then, the circuit of the decoding device 200 generates a prediction image used in the decoding of the current block by using the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter (S3212).

[1070] Thus, in the operations performed for each pixel position, it is possible to reduce the substantial multiplication operations with a large amount of calculation, and it is possible to derive a plurality of parameters for generating a prediction image with a low amount of calculation. Therefore, it is possible to reduce the processing amount in decoding. In addition, it is possible to appropriately generate a prediction image based on a plurality of parameters including a parameter related to a horizontal gradient value, a parameter related to a vertical gradient value, and a parameter related to both the horizontal gradient value and the vertical gradient value.

[1071] In addition, for example, the circuit of the decoding device 200 can derive the first parameter by the above formula (11.1), and can derive the second parameter by the above formula (11.2). In addition, the circuit of the decoding device 200 can derive the third parameter by the above formula (11.3), and derive the fourth parameter by the above formula (11.4). In addition, the circuit of the decoding device 200 can derive the fifth parameter by the above formula (11.5).

[1072] Here, Ω represents a set of a plurality of relative pixel positions, and [i, j] represents each relative pixel position. In addition, for each relative pixel position, I x 0 represents the horizontal gradient value of this relative pixel position in the first range, and I x 1 represents the horizontal gradient value of this relative pixel position in the second range. In addition, for each relative pixel position, I y 0 represents the vertical gradient value of this relative pixel position in the first range, and I y 1 represents the vertical gradient value of this relative pixel position in the second range.

[1073] In addition, for each relative pixel position, I 0 represents the pixel value of this relative pixel position in the first range, and I 1 represents the pixel value of this relative pixel position in the second range. And, abs(I x 1 +I x 0 ) represents the absolute value of I x 1 +I x 0 , and sign(I x 1 +I x 0 ) represents the positive or negative sign of I x 1 +I x 0 , and abs(I y 1 +I y0 ) represents I y 1 +I y 0 the absolute value of, sign(I y 1 +I y 0 ) represents I y 1 +I y 0 the positive or negative sign of.

[1074] Accordingly, it is possible to derive a plurality of parameters with a low computational load using pixel values, horizontal gradient values, and vertical gradient values.

[1075] In addition, for example, the circuit of the decoding device 200 can also derive the sixth parameter by dividing the third parameter by the first parameter. In addition, the circuit of the decoding device 200 can also derive the seventh parameter by subtracting the product of the fifth parameter and the sixth parameter from the fourth parameter and then dividing by the second parameter. Then, the circuit of the decoding device 200 can also generate a prediction image using the sixth parameter and the seventh parameter.

[1076] Accordingly, it is possible to appropriately aggregate a plurality of parameters into two parameters corresponding to the horizontal direction and the vertical direction. It is possible to appropriately incorporate parameters related to the horizontal gradient value into the parameter corresponding to the horizontal direction. It is possible to appropriately incorporate parameters related to the vertical gradient value, parameters related to both the horizontal gradient value and the vertical gradient value, and the parameter corresponding to the horizontal direction into the parameter corresponding to the vertical direction. Then, it is possible to generate a prediction image using these two parameters.

[1077] In addition, for example, the circuit of the decoding device 200 can derive the sixth parameter through the above formula (10.8). In addition, the circuit of the decoding device 200 can derive the seventh parameter through the above formula (10.9).

[1078] Here, sG x represents the first parameter, sG y represents the second parameter, sG x dI represents the third parameter, sG y dI represents the fourth parameter, sG x G y represents the fifth parameter, u represents the sixth parameter. Bits is a function that returns the value obtained by rounding up the binary logarithm of the independent variable to an integer.

[1079] Accordingly, it is possible to derive two parameters corresponding to the horizontal direction and the vertical direction with a low computational load.

[1080] In addition, for example, the circuit of the decoding device 200 can also generate a predicted image by deriving a predicted pixel value using the first pixel value, the first horizontal gradient value, the first vertical gradient value, the second pixel value, the second horizontal gradient value, the second vertical gradient value, the sixth parameter, and the seventh parameter.

[1081] Here, the predicted pixel value is the predicted pixel value of the processing target pixel position included in the current block. The first pixel value is the pixel value of the first pixel position corresponding to the processing target pixel position in the first reference block. The first horizontal gradient value is the horizontal gradient value of the first pixel position. The first vertical gradient value is the vertical gradient value of the first pixel position. The second pixel value is the pixel value of the second pixel position corresponding to the processing target pixel position in the second reference block. The second horizontal gradient value is the horizontal gradient value of the second pixel position. The second vertical gradient value is the vertical gradient value of the second pixel position.

[1082] Thus, it is possible to generate a predicted image using two parameters corresponding to the horizontal and vertical directions, etc., and the two parameters corresponding to the horizontal and vertical directions can be appropriately reflected in the predicted image.

[1083] Also, for example, the circuit of the decoding device 200 can also derive the predicted pixel value by dividing the sum of the first pixel value, the second pixel value, the first correction value, and the second correction value by 2. Here, the first correction value corresponds to the product of the sixth parameter and the difference between the first horizontal gradient value and the second horizontal gradient value, and the second correction value corresponds to the product of the seventh parameter and the difference between the first vertical gradient value and the second vertical gradient value.

[1084] Thus, by using two parameters corresponding to the horizontal and vertical directions, etc., it is possible to appropriately generate a predicted image.

[1085] In addition, for example, the circuit of the decoding device 200 can also derive the predicted pixel value by the above formula (10.10). Here, I 0 represents the first pixel value, I 1 represents the second pixel value, u represents the sixth parameter, I x 0 represents the first horizontal gradient value, I x 1 represents the second horizontal gradient value, v represents the seventh parameter, I y 0 represents the first vertical gradient value, I y 1 represents the second vertical gradient value. Thus, according to the formula associated with two parameters corresponding to the horizontal and vertical directions, etc., it is possible to appropriately generate a predicted image.

[1086] In addition, for example, the circuit of the decoding device 200 can also use the first parameter, the second parameter, the third parameter, the fourth parameter, and the fifth parameter to derive one or more parameters of the bidirectional optical flow, and use one or more parameters of the bidirectional optical flow and the bidirectional optical flow to generate a predicted image. Thus, the decoding device 200 can appropriately generate a predicted image.

[1087] One or more parameters of the bidirectional optical flow can also be at least one of the above-described sixth parameter and seventh parameter.

[1088] In addition, the inter-frame prediction unit 218 of the decoding device 200 can also perform the above-described operations as the circuit of the decoding device 200.

[1089] [Other Examples]

[1090] In each of the above examples, the encoding device 100 and the decoding device 200 can be used as an image encoding device and an image decoding device, respectively, or can be used as a moving image encoding device and a moving image decoding device.

[1091] Alternatively, the encoding device 100 and the decoding device 200 can also be used as prediction devices, respectively. That is, the encoding device 100 and the decoding device 200 can respectively correspond only to the inter-frame prediction unit 126 and the inter-frame prediction unit 218. Moreover, other components can be included in other devices.

[1092] In addition, at least a part of each of the above examples can be used as an encoding method, a decoding method, a prediction method, or other methods.

[1093] In addition, each component is constituted by dedicated hardware, but can also be implemented by executing a software program suitable for each component. Each component can be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded in a recording medium such as a hard disk or a semiconductor memory.

[1094] Specifically, each of the encoding device 100 and the decoding device 200 can include a processing circuit and a storage device that is electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to the processor a1 or b1, and the storage device corresponds to the memory a2 or b2.

[1095] The processing circuit includes at least one of dedicated hardware and a program execution unit, and uses the storage device to perform processing. In addition, when the processing circuit includes a program execution unit, the storage device stores a software program executed by the program execution unit.

[1096] Here, the software for implementing the above-described encoding device 100 or decoding device 200, etc. is a program as follows.

[1097] For example, it can also be that the program causes the computer to execute the following encoding method: respectively deriving a horizontal gradient and an absolute value for a plurality of relative pixel positions, where the plurality of relative pixel positions are a plurality of pixel positions commonly and relatively determined for both a first range of a first reference block including the current block and a second range of a second reference block including the current block, and are a plurality of pixel positions in each of the first range and the second range. The horizontal gradient and the absolute value are the absolute value of the sum of the horizontal gradient value of this relative pixel position in the first range and the horizontal gradient value of this relative pixel position in the second range. Deriving the sum of the plurality of horizontal gradients and absolute values respectively derived for the plurality of relative pixel positions as a first parameter. For each of the plurality of relative pixel positions, deriving the absolute value of the sum of the vertical gradient value of this relative pixel position in the first range and the vertical gradient value of this relative pixel position in the second range, that is, the vertical gradient and absolute value, and deriving the sum of the plurality of vertical gradient and absolute values respectively derived for the plurality of relative pixel positions as a second parameter. For each of the plurality of relative pixel positions, deriving the difference between the pixel value of this relative pixel position in the first range and the pixel value of this relative pixel position in th...

Claims

1. An encoding device, wherein, Comprising: A memory; And A processor, connected to the memory, in BDOF (Bidirectional Optical Flow), generating a predicted image based on the derived first parameter, second parameter, third parameter, fourth parameter, and fifth parameter to encode the current block. Based on [Equation 1] ∑ [i,j]∈Ω abs(I x 1 +I x 0 ) Derive the first parameter. Based on [Equation 2] ∑ [i,j]∈Ω abs(I y 1 +I y 0 ) Derive the second parameter. Based on [Equation 3] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 )) Derive the third parameter. Based on [Equation 4] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I 1 )) Derive the fourth parameter, and Based on [Equation 5] ∑ [i,j]∈Ω (sign(I y 1 +I y 0 )×(I x 1 +I x 0 )) Derive the fifth parameter. Where Ω represents a set of multiple relative pixel positions. [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω. I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position being determined based on the relative pixel position [i, j], the first gradient image and the second gradient image corresponding to the current block I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block, I 1 represents the pixel value at the first pixel position in the second interpolated image corresponding to the current block The abs function outputs the absolute value of the independent variable. The sign function outputs the sign of the independent variable, and the sign of the independent variable is -1, 0, or 1. I x 0 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, the target sample is included in the first prediction block, and I x 1 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in a second prediction block different from the first prediction block.

2. A decoding device, wherein, Comprising: A memory; And A processor, connected to the memory, in BDOF (Bidirectional Optical Flow), generating a predicted image based on the derived first parameter, second parameter, third parameter, fourth parameter, and fifth parameter to decode the current block. Based on [Equation 1] Σ [i,j]∈Ω abs(I x 1 +I x 0 ) Derive the first parameter. Based on [Equation 2] Σ [i,j]∈Ω abs(I y 1 +I y 0 ) Derive the second parameter. Based on [Equation 3] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 )) Derive the third parameter. Based on [Equation 4] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I 1 )) Derive the fourth parameter, and Based on [Equation 5] ∑ [i,j]∈Ω (sign(I y 1 +I y 0 )×(I x 1 +I x 0 )) Derive the fifth parameter. Where Ω represents a set of multiple relative pixel positions. [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω. I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position being determined based on the relative pixel position [i, j], the first gradient image and the second gradient image corresponding to the current block I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block I 1 represents the pixel value at the first pixel position in the second interpolation image corresponding to the current block The abs function outputs the absolute value of the independent variable. The sign function outputs the sign of the independent variable, and the sign of the independent variable is -1, 0, or 1. I x 0 represents a difference value between a right sample value and a left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, the target sample is included in the first prediction block, and I x 1 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in a second prediction block different from the first prediction block.

3. A non-transitory computer-readable medium storing a bitstream, the bitstream including prediction parameters, the multiple prediction parameters representing prediction processing candidates in the prediction processing, and being read by a decoding device to decode the current block provided in the bitstream, the multiple prediction processing candidates including BDOF (Bidirectional Optical Flow), and in the BDOF processing, using the first parameter, second parameter, third parameter, fourth parameter, and fifth parameter to generate a predicted image of the current block. Based on [Equation 1] Σ [i,j]∈Ω abs(I x 1 +I x 0 ) Derive the first parameter. Based on [Equation 2] Σ [i,j]∈Ω abs(I y 1 +I y 0 ) Derive the second parameter. Based on [Equation 3] ∑ [i,j]∈Ω (-sign(I x 1 +I x 0 )×(I 0 -I 1 )) Derive the third parameter. Based on [Equation 4] ∑ [i,j]∈Ω (-sign(I y 1 +I y 0 )×(I 0 -I 1 )) Derive the fourth parameter, and Based on [Equation 5] ∑ [i,j]∈Ω (sign(I y 1 +I y 0 )×(I x 1 +I x 0 )) Derive the fifth parameter. Among them, Ω represents a set of multiple relative pixel positions. [i, j] is defined by the horizontal position i and the vertical position j, representing the relative pixel position in the set Ω. I x 0 represents the horizontal gradient value at the first pixel position in the first gradient image, I x 1 represents the horizontal gradient value at the first pixel position in the second gradient image, the first pixel position is determined based on the relative pixel position [i, j], and the first gradient image and the second gradient image correspond to the current block I y 0 represents the vertical gradient value at the first pixel position in the first gradient image, I y 1 represents the vertical gradient value at the first pixel position in the second gradient image I 0 represents the pixel value at the first pixel position in the first interpolation image corresponding to the current block I 1 represents the pixel value at the first pixel position in the second interpolation image corresponding to the current block The abs function outputs the absolute value of the independent variable. The sign function outputs the sign of the independent variable, and the sign of the independent variable is -1, 0, or 1. I x 0 represents a differential value between a right sample value and a left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, the target sample is included in the first prediction block, and I x 1 represents the difference value between the right sample value and the left sample value, where the right sample is adjacent to the right side of the target sample, the left sample is adjacent to the left side of the target sample, and the target sample is included in a second prediction block different from the first prediction block.