Image Encoding Device, Image Decoding Device, and Bitstream Generation Device

By dividing image blocks into non-rectangular partitions and using specialized motion vectors and weighting techniques, the system addresses inefficiencies in video coding, enhancing encoding and decoding speed and efficiency for non-rectangular blocks.

JP7705989B2Active Publication Date: 2025-07-10PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Patent Information

Application Number
JP2024101219
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-07-16
Filing Date
2024-06-24
Publication Date
2025-07-10
Estimated Expiration
2038-08-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently handling the increasing amount of digital video data, particularly in optimizing inter prediction and intra prediction processes, especially when dealing with non-rectangular image blocks.

Method used

The implementation of a circuit and memory system that divides image blocks into non-rectangular partitions, such as triangles or polygons, and uses specific motion vectors and weighting techniques for improved encoding and decoding efficiency, including boundary smoothing operations.

Benefits of technology

Enhances encoding and decoding speed and efficiency by optimizing processing for non-rectangular image blocks, leading to improved image quality and coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705989000003
    Figure 0007705989000003
  • Figure 0007705989000004
    Figure 0007705989000004
  • Figure 0007705989000005
    Figure 0007705989000005
Patent Text Reader

Abstract

To improve the video coding technique.SOLUTION: An image encoder 100 includes circuitry and a memory. The circuitry, in operation, performs processing along a boundary between a first partition and a second partition. The processing includes: selecting a first motion vector from a first motion vector candidate set; first-predicting a plurality of first values of a set of pixels by using the first motion vector; selecting a second motion vector from the first motion vector candidate set; second-predicting a plurality of second values of the set of pixels by using the second motion vector; and weighting the plurality of first values and the plurality of second values. When the ratio between the width of the image block and the height of the image block is greater than 4, or the ratio between the height and the width is greater than 4, the circuitry disables processing.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to video coding, for example, constructing a current block based on a reference frame by performing inter prediction, or constructing a current block based on an encoded / decoded reference block within a current frame by performing intra prediction, and relates to systems, components, and methods in video encoding and decoding of moving images, etc.

Background Art

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this advancement, there has always been a need to provide improvements and optimizations to video coding technology to handle the continuously increasing amount of digital video data in various applications. The present disclosure further relates to advancements, improvements, and optimizations in video coding, particularly inter prediction or intra prediction that divides an image block into a plurality of partitions including at least a first partition having a non-rectangular shape (e.g., a triangle) and a second partition.

Summary of the Invention

[0003] According to one aspect, there is provided an image encoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit performs processing along a boundary between a first partition having a non-rectangular shape and a second partition in an image block, the processing including: selecting a first motion vector of the first partition from a first set of motion vector candidates; using the first motion vector to perform a first prediction of a plurality of first values of a pixel set of the first partition; selecting a second motion vector of the second partition from the first set of motion vector candidates; using the second motion vector to perform a second prediction of a plurality of second values of the pixel set of the first partition; and weighting the plurality of first values and the plurality of second values. When a ratio of a width to a height of the image block is greater than 4, or a ratio of the height to the width of the image block is greater than 4, the circuit invalidates the processing.

[0004] In some examples of embodiments of the present disclosure, appropriate components / operations used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc., may be efficiently selected to improve encoding efficiency, simplify encoding / decoding processing, accelerate encoding / decoding processing speed.

[0005] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. The benefits and / or advantages may be obtained individually by the various embodiments as well as the features of the specification and drawings, and it is not necessary to provide all of the various embodiments as well as the features of the specification and drawings in order to obtain one or more of such benefits and / or advantages.

[0006] Note that the comprehensive or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any combination thereof.

Brief Description of Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 9D

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21A

Figure 21B

Figure 21C

Figure 21D

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

DETAILED DESCRIPTION OF THE INVENTION

[0008] According to one aspect, there is provided an image encoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, encodes the first partition using the first motion vector, and encodes the second partition using the second motion vector.

[0009] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0010] According to another aspect, the predicting includes selecting the first motion vector from a first set of motion vector candidates and selecting the second motion vector from a second set of motion vector candidates. For example, the first set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the first partition, and the second set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the second partition. The plurality of partitions adjacent to the first partition and the plurality of partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The plurality of adjacent partitions may be one or both of a plurality of spatially adjacent partitions and a plurality of temporally adjacent partitions. The first set of motion vector candidates may be the same as or different from the second set of motion vector candidates.

[0011] According to another aspect, the predicting includes selecting a first motion vector candidate from a first set of motion vector candidates, deriving the first motion vector by adding a first differential motion vector to the first motion vector candidate, selecting a second motion vector candidate from a second set of motion vector candidates, and deriving the second motion vector by adding a second differential motion vector to the second motion vector candidate.

[0012] According to another aspect, in operation, a splitting unit that receives an original image and splits it into a plurality of blocks, in operation, an adding unit that receives the plurality of blocks from the splitting unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual, in operation, a conversion unit that performs a conversion on the plurality of residuals output from the adding unit and outputs a plurality of conversion coefficients, in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients, in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bitstream, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture, in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture, and a prediction control unit connected to a memory are provided. The prediction control unit, in operation, splits the plurality of blocks into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, encodes the first partition using the first motion vector, and encodes the second partition using the second motion vector.

[0013] According to another aspect, there is provided an image encoding method including splitting an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition, predicting a second motion vector for the second partition, encoding the first partition using the first motion vector, and encoding the second partition using the second motion vector.

[0014] According to one aspect, an image decoding apparatus is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, decodes the first partition using the first motion vector, and decodes the second partition using the second motion vector.

[0015] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0016] According to another aspect, in operation, an entropy decoding unit that receives and decodes a coded bitstream to obtain a plurality of quantized transform coefficients, an inverse quantization unit and an inverse transform unit that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals, an addition unit that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, an inter prediction unit that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit that, in operation, generates a prediction of a current block based on a decoded reference block in a current picture, and the prediction control unit connected to the memory are provided. In operation, the prediction control unit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, decodes the first partition using the first motion vector, and decodes the second partition using the second motion vector.

[0017] According to another aspect, an image decoding method is provided that includes dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition, predicting a second motion vector for the second partition, decoding the first partition using the first motion vector, and decoding the second partition using the second motion vector.

[0018] According to one aspect, an image encoding apparatus is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0019] According to a further aspect, the non-rectangular shape is a triangle. According to another aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles. According to yet another aspect, the second partition has a non-rectangular shape.

[0020] According to another aspect, at least one of the first prediction and the second prediction is an inter prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in an encoded reference picture. The inter prediction process may predict the plurality of first values of the plurality of pixels of the first partition including the pixel set, or may predict the plurality of second values only of the pixel set of the first partition.

[0021] According to another aspect, at least one of the first prediction and the second prediction is an intra prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in a current picture.

[0022] According to another aspect, the prediction method used for the first prediction is different from the prediction method used for the second prediction.

[0023] According to a further aspect, the number of the pixel sets in each row or each column for predicting the plurality of first values and the plurality of second values is an integer. For example, when the number of the pixel sets in each row or each column is 4, a plurality of weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the plurality of first values of the 4 pixels in the pixel set respectively, and a plurality of weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the plurality of second values of the 4 pixels in the pixel set respectively. As another example, when the number of the pixel sets in each row or each column is 2, a plurality of weights of 1 / 3 and 2 / 3 may be applied to the plurality of first values of the 2 pixels in the pixel set respectively, and a plurality of weights of 2 / 3 and 1 / 3 may be applied to the plurality of second values of the 2 pixels in the pixel set respectively.

[0024] According to another aspect, the plurality of weights may be integer values or fractional values.

[0025] According to another aspect, in operation, a dividing unit that receives an original image and divides it into a plurality of blocks; in operation, a adding unit that receives the plurality of blocks from the dividing unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; in operation, a conversion unit that performs a conversion on the plurality of residuals output from the adding unit and outputs a plurality of conversion coefficients; in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients; in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bitstream; in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture; in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture; and an image encoding apparatus including the prediction control unit connected to a memory are provided. The prediction control unit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block, in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0026] According to another aspect, there is provided an image encoding method that performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The method generally includes four steps of: first predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition; second predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition; weighting the plurality of first values and the plurality of second values; and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0027] According to a further aspect, there is provided an image decoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The boundary smoothing operation includes: first predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition; second predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition; weighting the plurality of first values and the plurality of second values; and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0028] According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles. According to another aspect, the second partition has a non-rectangular shape.

[0029] According to another aspect, at least one of the first prediction and the second prediction is an inter prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in an encoded reference picture. The inter prediction process may predict the plurality of first values of the plurality of pixels of the first partition including the pixel set, or may predict the plurality of second values of only the pixel set of the first partition.

[0030] According to another aspect, at least one of the first prediction and the second prediction is an intra prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in a current picture.

[0031] According to another aspect, in operation, an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, in operation, an inverse quantization unit that inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and an inverse transform unit that inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, in operation, an addition unit that adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and the plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in a decoded reference picture, in operation, an intra prediction unit that generates a prediction of the current block based on a decoded reference block in a current picture, and a prediction control unit connected to a memory. The prediction control unit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block, in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0032] According to another aspect, there is provided an image decoding method for performing a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The method generally includes four steps: predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition; predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values and the plurality of second values; and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.

[0033] According to one aspect, there is provided an image encoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit performs a partition syntax operation, which includes dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the division, encoding the first partition and the second partition, and writing one or more parameters including the partition parameter into a bit stream.

[0034] According to a further aspect, the partition parameter indicates that the first partition has a triangular shape.

[0035] According to another aspect, the partition parameter indicates that the second partition has a non-rectangular shape.

[0036] According to another aspect, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0037] According to another aspect, the partition parameter encodes, as a whole, the division direction applied to divide the image block into the plurality of partitions. For example, the division direction may include from the upper left corner to the lower right corner of the image block and from the upper right corner to the lower left corner of the image block. The partition parameter may encode, as a whole, at least the first motion vector of the first partition.

[0038] According to another aspect, one or more parameters other than the partition parameter encode the division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode, as a whole, at least the first motion vector of the first partition.

[0039] According to another aspect, the partition parameter may encode, as a whole, at least the first motion vector of the first partition. The partition parameter may encode, as a whole, the second motion vector of the second partition.

[0040] According to another aspect, one or more parameters other than the partition parameter may encode at least the first motion vector of the first partition.

[0041] According to another aspect, the one or more parameters are binarized according to a binarization method selected according to at least one value of the one or more parameters.

[0042] According to a further aspect, in operation, a splitting unit that receives an original image and splits it into a plurality of blocks, in operation, an addition unit that receives the plurality of blocks from the splitting unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual, in operation, a conversion unit that performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients, in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients, in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bit stream, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture, in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture, and an image encoding apparatus including the prediction control unit connected to a memory are provided. The prediction control unit, in operation, splits an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the splitting, and encodes the first partition and the second partition. The entropy encoding unit, in operation, writes one or more parameters including the partition parameter into the bit stream.

[0043] According to another aspect, an image encoding method including a partition syntax operation is provided. The method generally includes three steps: splitting an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the splitting, encoding the first partition and the second partition, and writing one or more parameters including the partition parameter into the bit stream.

[0044] According to another aspect, an image decoding apparatus is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit performs a partition syntax operation, and the partition syntax operation includes reading one or more parameters including a partition parameter indicating dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition from a bitstream, dividing the image block into the plurality of partitions based on the partition parameter, and decoding the first partition and the second partition.

[0045] According to a further aspect, the partition parameter indicates that the first partition has a triangular shape.

[0046] According to another aspect, the partition parameter indicates that the second partition has a non-rectangular shape.

[0047] According to another aspect, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0048] According to another aspect, the partition parameter encodes the division direction applied to divide the image block into the plurality of partitions as a unit. For example, the division direction includes from the upper left corner to the lower right corner of the image block and from the upper right corner to the lower left corner of the image block. The partition parameter may encode at least the first motion vector of the first partition as a unit.

[0049] According to another aspect, one or more parameters other than the partition parameter encode the division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode at least the first motion vector of the first partition as a unit.

[0050] According to another aspect, the partition parameter may encode at least the first motion vector of the first partition integrally. The partition parameter may encode the second motion vector of the second partition integrally.

[0051] According to another aspect, one or more parameters other than the partition parameter may encode at least the first motion vector of the first partition.

[0052] According to another aspect, the one or more parameters are binarized according to a binarization method selected according to at least one value of the one or more parameters.

[0053] According to a further aspect, in operation, an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, and in operation, inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, and in operation, an addition unit that adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, and in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit that generates a prediction of the current block based on a decoded reference block in a current picture, and an image decoding apparatus including the prediction control unit connected to a memory are provided. The entropy decoding unit reads, from the bitstream, one or more parameters including a partition parameter indicating that an image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition in operation, divides the image block into the plurality of partitions based on the partition parameter, and decodes the first partition and the second partition.

[0054] According to another aspect, an image decoding method including a partition syntax operation is provided. The method generally includes three steps: decoding one or more parameters including partition parameters indicating to divide an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition from a bit stream, dividing the image block into the plurality of partitions based on the partition parameters, and decoding the first partition and the second partition.

[0055] In the drawings, the same reference numerals indicate the same components. The sizes and relative positions of the components in the drawings do not necessarily follow a scale ratio.

[0056] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that the embodiments described below are all illustrative or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are only examples and are not intended to limit the scope of the claims. Therefore, components not described in the independent claims defining the broadest inventive concept, although disclosed in the following embodiments, may be understood as arbitrary components.

[0057] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, for example, any of the following may be implemented.

[0058] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0059] (2) In the encoding device or decoding device of the embodiments, arbitrary changes such as addition, replacement, deletion, etc. may be made to the functions or processes performed by some of the multiple components of the encoding device or decoding device. For example, any function or process may be replaced or combined with other functions or processes described in any of the aspects of the present disclosure.

[0060] (3) In the method performed by the encoding device or decoding device of the embodiments, arbitrary changes such as addition, replacement, and deletion may be made to some of the multiple processes included in the method. For example, any process in the method may be replaced or combined with other processes described in any of the aspects of the present disclosure.

[0061] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiments may be combined with the components described in any of the aspects of the present disclosure, or may be combined with the components having a part of the functions described in any of the aspects of the present disclosure, or may be combined with the components that perform a part of the processes performed by the components described in each aspect of the present disclosure.

[0062] (5) The components having a part of the functions of the encoding device or decoding device of the embodiments, or the components that perform a part of the processes of the encoding device or decoding device of the embodiments may be combined or replaced with the components described in any of the aspects of the present disclosure, the components having a part of the functions described in any of the aspects of the present disclosure, or the components that perform a part of the processes described in any of the aspects of the present disclosure.

[0063] (6) In the method performed by the encoding device or decoding device according to the embodiment, any one of the plurality of processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.

[0064] (7) Some of the plurality of processes included in the method performed by the encoding device or decoding device according to the embodiment may be combined with the process described in any of the aspects of the present disclosure.

[0065] (8) The manner of implementing the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device according to the embodiment. For example, the processes and / or configurations may be implemented in a device used for a purpose different from the moving image encoding or moving image decoding disclosed in the embodiment.

[0066] [Overview of Encoding Device] First, the overview of the encoding device according to the embodiment will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 according to the embodiment. The encoding device 100 is a moving image encoding device that encodes moving images in block units.

[0067] As shown in FIG. 1, the encoding device 100 is a device that encodes images in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0068] The symbolization device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. Further, the symbolization device 100 may be realized as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0069] Each component included in the symbolization device 100 will be described below.

[0070] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These blocks of fixed size are sometimes called coding tree units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These blocks of variable size are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be processing units for CUs, PUs, and TUs.

[0071] FIG. 2 is a diagram showing an example of block splitting in the embodiment. In FIG. 2, the solid line represents the block boundary by quadtree block splitting, and the broken line represents the block boundary by binary tree block splitting.

[0072] Here, block 10 is a square block of 128 x 128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quad-tree block division).

[0073] The upper left 64x64 block is further vertically divided into two rectangular 32x64 blocks, and the left 32x64 block is further vertically divided into two rectangular 16x64 blocks (binary-tree block division). As a result, the upper left 64x64 block is divided into two 16x64 blocks 11, 12 and a 32x64 block 13.

[0074] The upper right 64x64 block is horizontally divided into two rectangular 64x32 blocks 14, 15 (binary-tree block division).

[0075] The lower left 64x64 block is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper left block and the lower right block are further divided. The upper left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the lower left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0076] The lower right 64x64 block 23 is not divided.

[0077] As described above, in FIG. 2, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0078] In FIG. 2, one block was divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to these. For example, one block may be divided into three blocks (ternary-tree block division). Such a division including ternary-tree block division may be called MBT (multi type tree) division.

[0079] [Subtraction unit] The subtraction unit 104 receives an input from the division unit 102 and subtracts a prediction signal (prediction samples input from the prediction control unit 128 shown below) from the original signal (original samples) in units of blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also called the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.

[0080] The original signal is an input signal of the encoding device 100 and is a signal representing the image of each picture constituting the moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0081] [Conversion unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.

[0082] Note that the conversion unit 106 may adaptively select a conversion type from a plurality of conversion types and convert the prediction error into conversion coefficients using a transform basis function corresponding to the selected conversion type. Such a conversion may be called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).

[0083] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the transform basis functions corresponding to each transform type. In FIG. 3, N indicates the number of input pixels. The selection of a transform type from among these plurality of transform types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.

[0084] Information indicating whether to apply such EMT or AMT (for example, called an EMT flag or an AMT flag) and information indicating the selected transform type are usually signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level, and may be at other levels (for example, the bit sequence level, the picture level, the slice level, the tile level, or the CTU level).

[0085] Also, the conversion unit 106 may re-convert the conversion coefficients (conversion results). Such re-conversion may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (for example, 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are usually signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level, and may be at other levels (for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level).

[0086] The conversion unit 106 may be applied with separable conversion and non-separable conversion. Separable conversion is a method of performing multiple conversions by separating for each direction by the number of dimensions of the input, and non-separable conversion is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing the conversion collectively.

[0087] For example, as an example of non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.

[0088] In a further example of non-separable conversion, after regarding a 4×4 input block as an array having 16 elements, a conversion (for example, Hypercube Givens Transform) that performs a plurality of Givens rotations on the array may be performed.

[0089] [Quantization unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order, and quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0090] The predetermined scanning order is an order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0091] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the quantization error increases.

[0092] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) based on the quantized coefficients input from the quantization unit 108. Specifically, for example, the entropy encoding unit 110 binarizes the quantized coefficients, arithmetic-encodes the binary signal, and outputs a compressed bit stream or sequence.

[0093] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114.

[0094] [Inverse Transform Unit] The inverse transform unit 114 restores the prediction error (residual) by inverse-transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients to restore the prediction error of the current block. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.

[0095] Note that since information is usually lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error usually includes a quantization error.

[0096] [Addition Unit] The adder 116 reconstructs the current block by adding the prediction error input from the inverse conversion unit 114 and the prediction sample input from the prediction control unit 128. Then, the adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.

[0097] [Block Memory] The block memory 118 is a storage unit for storing blocks within the coded picture (referred to as the "current picture") that are referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the adder 116.

[0098] [Loop Filter Unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop and includes, for example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF).

[0099] In the case of the ALF, a least-squares error filter for removing coding distortion is applied, and for example, for each 2x2 sub-block within the current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.

[0100] Specifically, first, sub-blocks (for example, 2x2 sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed based on the direction and activity of the gradient. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes.

[0101] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantizing the addition result.

[0102] Based on the results of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0103] As the shape of the filter used in ALF, for example, a circularly symmetric shape is used. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Note that the signaling of the information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (for example, sequence level, slice level, tile level, CTU level, or CU level).

[0104] The on / off of ALF may be determined at the picture level or the CU level. For example, whether to apply ALF at the CU level may be determined for luminance, and whether to apply ALF at the picture level may be determined for chrominance. Information indicating the on / off of ALF is usually signaled at the picture level or the CU level. Note that the signaling of the information indicating the on / off of ALF is not necessarily limited to the picture level or the CU level and may be at other levels (for example, sequence level, slice level, tile level, or CTU level).

[0105] A set of coefficients for a plurality of selectable filters (e.g., filters up to 15 or 25) is usually signaled at the picture level. Note that the signaling of the coefficient set does not necessarily have to be limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0106] [Frame memory] Frame memory 122 is a storage unit for storing reference pictures used for inter prediction and may also be referred to as a frame buffer, for example. Specifically, frame memory 122 stores the reconstructed blocks filtered by loop filter unit 120.

[0107] [Intra prediction unit] Intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also referred to as in-picture prediction) of a current block with reference to a block within the current picture as stored in block memory 118. Specifically, intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block and outputs the intra prediction signal to prediction control unit 128.

[0108] For example, intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes usually include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0109] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC standard.

[0110] The plurality of directional prediction modes include, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions.

[0111] FIG. 5A is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent 32 additional directions (the 2 "non-directional" prediction modes are not shown in FIG. 5A).

[0112] In various processing examples, in the intra prediction of a chrominance block, a luminance block may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. This intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0113] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction. Such intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called a PDPC flag) is usually signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level or CTU level).

[0114] [Inter Prediction Unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also referred to as inter-picture prediction) of the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 block) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or current sub-block, and finds a reference block or sub-block within the reference picture that most matches the current block or sub-block. Then, the inter prediction unit 126 acquires motion information (e.g., motion vector) that compensates for (or predicts) the motion or change from the reference block or sub-block to the current block or sub-block. Then, the inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, and generates an inter prediction signal for the current block or sub-block. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0115] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms. For example, the motion vector may be signaled. As another example, the difference between the motion vector and the predicted motion vector may be signaled.

[0116] Note that the inter prediction signal may be generated using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks. Specifically, an inter prediction signal may be generated in units of sub-blocks within the current block by weighted addition of a prediction signal based on the motion information obtained by motion search (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture). Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation).

[0117] In the OBMC mode, information indicating the size of sub-blocks for OBMC (e.g., called OBMC block size) may be signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (e.g., called OBMC flag) may be signaled at the CU level. Note that the signaling level of these pieces of information does not have to be limited to the sequence level and the CU level, and may be at other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0118] The OBMC mode will be described in more detail. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the prediction image correction process by OBMC processing.

[0119] Referring to FIG. 5C, first, a prediction image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded (current). In FIG. 5C, the arrow “MV” points to the reference picture, indicating what the current block in the current picture refers to in order to obtain the prediction image.

[0120] Next, the motion vector (MV_L) already derived for the encoded left adjacent block is applied (reused) to the block to be coded (current) to obtain a prediction image (Pred_L). The motion vector (MV_L) is indicated by the arrow “MV_L” pointing from the current block to the reference picture. Then, the first correction of the prediction image is performed by superimposing the two prediction images Pred and Pred_L. This has the effect of mixing the boundaries between adjacent blocks.

[0121] Similarly, the motion vector (MV_U) already derived for the encoded upper adjacent block is applied (reused) to the block to be encoded (current) block to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. Then, the predicted image Pred_U is superposed on the predicted image after the first correction (i.e., Pred and Pred_L) to perform the second correction of the predicted image. This has the effect of mixing the boundaries between adjacent blocks in one aspect. The predicted image obtained by the second correction is the final predicted image of the current block with the boundaries mixed (smoothed) with adjacent blocks.

[0122] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described, but it is also possible to configure to perform more corrections than two stages using the right adjacent block or the lower adjacent block.

[0123] Note that the region for superposition may be only a partial region near the block boundary, rather than the entire pixel region of the block.

[0124] Here, the prediction image correction process of OBMC for obtaining one predicted image Pred by superposing additional predicted images Pred_L and Pred_U based on one reference picture has been described. However, when the predicted image is corrected based on a plurality of reference pictures, the same process may be applied to each of the plurality of reference pictures. In such a case, by performing OBMC image correction based on a plurality of reference pictures, after obtaining the corrected predicted images from each reference picture, the finally obtained predicted image is obtained by further superposing the plurality of obtained corrected predicted images.

[0125] Note that in OBMC, the unit of the target block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.

[0126] As a method for determining whether to apply OBMC processing, for example, there is a method of using an obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the value "1" is set as the obmc_flag and OBMC processing is applied for encoding. If it does not belong to a region with complex motion, the value "0" is set as the obmc_flag and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding the obmc_flag described in the stream (i.e., the compressed sequence), decoding is performed by switching whether to apply OBMC processing according to the value.

[0127] Note that the motion information may be derived on the decoding device side without being signaled from the encoding device side. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also, for example, the motion information may be derived by performing motion search on the decoding device side. In this case, on the decoding device side, the motion search may be performed without using the pixel values of the current block.

[0128] Here, the mode of performing motion search on the decoding device side will be described. This mode of performing motion search on the decoding device side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.

[0129] An example of FRUC processing is shown in FIG. 5D. First, by referring to the motion vectors of encoded blocks that are spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector (MV) is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, the evaluation value of each candidate MV included in the candidate list is calculated, and one candidate MV is selected based on the evaluation value.

[0130] Based on the motion vectors of the selected candidates, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, a motion vector for the current block may be derived. That is, for the region around the best candidate MV, search is performed using pattern matching and evaluation values in the reference picture. If there is an MV with a better evaluation value, the best candidate MV is updated to the MV, and it may be used as the final MV of the current block. It is also possible to configure not to perform the process of updating to an MV with a better evaluation value.

[0131] When processing is performed in sub-block units, it may be the same process.

[0132] Note that the evaluation value may be calculated in various ways. For example, the reconstructed image of the region in the reference picture corresponding to the motion vector is compared with the reconstructed image of a predetermined region (for example, as described later, it may be the region of another reference picture or the region of an adjacent block of the current picture), and the difference in pixel values between the two reconstructed images is calculated and used as the evaluation value of the motion vector. In addition to the difference value, other information may be used to calculate the evaluation value.

[0133] Next, an example of pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (for example, the merge list) is selected as the starting point of the search by pattern matching. For example, as the pattern matching, the first pattern matching or the second pattern matching may be used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0134] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, for the area in the reference picture, as the predetermined area for calculating the evaluation value of the above-mentioned candidate, the area in another reference picture along the motion trajectory of the current block is used.

[0135] FIG. 6 is a diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the pair that most matches among pairs of two blocks in two different reference pictures (Ref0, Ref1), which are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV, which can bring good results.

[0136] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, two mirror-symmetric bidirectional motion vectors are derived.

[0137] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent blocks)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined region for calculating the evaluation value of the candidate described above.

[0138] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the coded region of both or either of the left adjacent and upper adjacent regions and the reconstructed image at the equivalent position in the coded reference picture (Ref0) specified by the candidate MV is derived, and an evaluation value is calculated using the obtained difference value, and it is possible to select the candidate MV with the best evaluation value among the plurality of candidate MVs as the best candidate MV.

[0139] Information indicating whether or not to apply such an FRUC mode (e.g., called an FRUC flag) may be signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the pattern matching method (e.g., the first pattern matching or the second pattern matching) (e.g., called an FRUC mode flag) may be signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0140] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called the BIO (bi - directional optical flow) mode.

[0141] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x , v y ) represents a velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.

[0142] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0143]

Equation

[0144] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the merge list or the like may be corrected in pixel units.

[0145] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks.

[0146] Next, a mode of deriving a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.

[0147] FIG. 9A is a diagram for explaining the derivation of a motion vector in units of sub-blocks based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v x ,v y ) of each sub-block within the current block is derived by the following formula (2).

[0148] [Equation]

[0149] Here, x and y respectively indicate the horizontal position and the vertical position of the sub-block, and w indicates a predetermined weight coefficient.

[0150] This affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right corner control points are different. Information indicating this affine motion compensation prediction mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode need not be limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0151] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal (a signal output from the intra prediction unit 124) or an inter prediction signal (a signal output from the inter prediction unit 126), and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0152] As shown in FIG. 1, in various processing examples, the prediction control unit 128 may output prediction parameters input to the entropy encoding unit 110. The entropy encoding unit 110 may generate an encoded bit stream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantized coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bit stream and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used in the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or indicating the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0153] FIG. 9B shows an example of a process for deriving the motion vector of a current picture in the merge mode.

[0154] First, a prediction MV list in which candidates for the predicted MV are registered is generated. As candidates for the predicted MV, there are a spatial adjacent prediction MV which is the MV of a plurality of encoded blocks spatially adjacent to the target block, a temporal adjacent prediction MV which is the MV of a neighboring block obtained by projecting the position of the target block in the encoded reference picture, a combined prediction MV which is an MV generated by combining the MV values of the spatial adjacent prediction MV and the temporal adjacent prediction MV, a zero prediction MV which is an MV with a value of zero, and the like.

[0155] Next, one predicted MV is selected from among the plurality of predicted MVs registered in the prediction MV list, and thus determined as the MV of the target block.

[0156] Furthermore, in the variable length encoding section, merge_idx, which is a signal indicating which predicted MV has been selected, is described in the stream and encoded.

[0157] Note that the predicted MVs registered in the prediction MV list described in FIG. 9B are just examples, and there may be a configuration in which the number is different from that in the figure, a configuration that does not include some types of the predicted MVs in the figure, or a configuration in which predicted MVs other than the types of the predicted MVs in the figure are added.

[0158] The final MV may be determined by performing a DMVR (decoder motion vector refinement) process, which will be described later, using the MV of the target block derived in the merge mode.

[0159] FIG. 9C is a conceptual diagram for explaining an example of the DMVR process for determining the MV.

[0160] First, take the optimal MVP set for the current block (e.g., in the merge mode) as the candidate MV. Then, according to the candidate MV (L0), identify reference pixels from the first reference picture (L0), which is the encoded picture in the L0 direction. Similarly, according to the candidate MV (L1), identify reference pixels from the second reference picture (L1), which is the encoded picture in the L1 direction. Generate a template by taking the average of these reference pixels.

[0161] Next, use the template to search the peripheral regions of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) respectively, and determine the MV with the minimum cost as the final MV. Note that the cost value may be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search region, and the candidate MV values.

[0162] Typically, in the encoder and the decoder described later, the configurations and operations of the processes described here are basically common.

[0163] Even if it is not the process example described here, any process can be used as long as it can search the periphery of the candidate MV to derive the final MV.

[0164] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.

[0165] FIG. 9D is a conceptual diagram for explaining an example of a predicted image generation method using luminance correction processing by LIC processing.

[0166] First, derive an MV from the encoded reference picture and obtain a reference image corresponding to the current block.

[0167] Next, for the current block, information indicating how the luminance values change between the reference picture and the current picture is extracted. This extraction is performed based on the luminance pixel values in the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV. Then, using the information indicating how the luminance values change, luminance correction parameters are calculated.

[0168] By performing a luminance correction process of applying the luminance correction parameters to the reference image in the reference picture specified by the MV, a predicted image for the current block is generated.

[0169] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0170] Also, although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from a plurality of reference pictures. Luminance correction processing may be performed on the reference images obtained from each reference picture in the same manner as described above, and then a predicted image may be generated.

[0171] As a method for determining whether to apply the LIC process, for example, there is a method using a lic_flag which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value "1" is set as the lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value "0" is set as the lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding may be performed by switching whether to apply the LIC process according to the value.

[0172] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the current block is in the merge mode, it is determined whether the encoded blocks around the ones selected during the derivation of the MV in the merge mode process are encoded by applying the LIC process. Encoding is performed by switching whether to apply the LIC process according to the result. Note that, also in this example, the same process is applied to the process on the decoder side.

[0173] [Outline of Decoder] Next, an outline of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing the functional configuration of a decoder 200 according to an embodiment. The decoder 200 is a moving image decoder that decodes moving images in block units.

[0174] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0175] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0176] Each component included in the decoder 200 will be described below.

[0177] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bitstream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bitstream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. The entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may output prediction parameters included in the encoded bitstream (see FIG. 1) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can execute the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoder side.

[0178] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the decoding target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, for each quantization coefficient of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0179] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error (residual) by inverse-transforming the transform coefficients input from the inverse quantization unit 204.

[0180] For example, when the information read from the encoded bitstream indicates that EMT or AMT is to be applied (for example, the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the read transform type.

[0181] For example, when the information decoded from the encoded bitstream indicates that the NSST is to be applied, the inverse conversion unit 206 applies an inverse reconversion to the conversion coefficients.

[0182] [Addition unit] The addition unit 208 reconstructs the current block by adding the prediction error input from the inverse conversion unit 206 and the prediction sample input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0183] [Block memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture) that are blocks referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the addition unit 208.

[0184] [Loop filter unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.

[0185] When the information indicating the on / off of the ALF decoded from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0186] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.

[0187] [Intra prediction unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bitstream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0188] In addition, when an intra prediction mode that refers to the luminance block is selected for the intra prediction of the chrominance difference block, the intra prediction unit 216 may predict the chrominance difference component of the current block based on the luminance component of the current block.

[0189] Also, when the information decoded from the encoded bitstream (e.g., the prediction parameter output from the entropy decoding unit 202) indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical directions.

[0190] [Inter Prediction Unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bitstream (e.g., the prediction parameter output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.

[0191] When the information decoded from the encoded bitstream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.

[0192] Also, when the information decoded from the encoded bit stream indicates that the FRUC mode is applicable, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0193] Also, when the BIO mode is applicable, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bit stream indicates that the affine motion compensation prediction mode is applicable, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.

[0194] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the addition unit 208 as the prediction signal. Overall, the configuration, function, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoder side may correspond to the configuration, function, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoder side.

[0195] [Non-rectangular partitioning] Also in the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126 of the symbolization device (see FIG. 1), and also in the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218 of the decoding device (see FIG. 10), conventionally, a plurality of partitions (or a plurality of variable-size blocks, or a plurality of sub-blocks) obtained from the division of each block and from which motion information (for example, a plurality of motion vectors) can be obtained are always rectangular as shown in FIG. 2. The inventors have discovered that generating a plurality of partitions having a non-rectangular shape such as a triangular shape can lead to an improvement in image quality and coding efficiency according to the content of the image in the picture in various implementations. Hereinafter, various embodiments in which at least one partition divided from an image block for the purpose of prediction has a non-rectangular shape will be described. Note that these embodiments are equally applicable to the encoding device side (the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126) and to the decoding device side (the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218), and may be implemented in the encoding device of FIG. 1 or the decoding device of FIG. 10 or the like.

[0196] FIG. 11 is a flowchart showing an example of a process of dividing an image block into a plurality of partitions including at least a first partition having a non-rectangular shape (for example, a triangle) and a second partition, and further performing a process of encoding (or decoding) the image block as a reconfigured combination of the first partition and the second partition.

[0197] In step S1001, the image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. For example, as shown in FIG. 12, the image block may be divided from the upper left corner to the lower right corner of the image block to create a first partition and a second partition, both having a non-rectangular shape (e.g., a triangle). Alternatively, the image block may be divided from the upper right corner to the lower left corner of the image block to create a first partition and a second partition, both having a non-rectangular shape (e.g., a triangle). With reference to FIGS. 12 and 17 to 19, various examples of non-rectangular division will be described later.

[0198] In step S1002, the process predicts a first motion vector for the first partition and a second motion vector for the second partition. For example, predicting the first motion vector and the second motion vector may include selecting the first motion vector from a set of first motion vector candidates and selecting the second motion vector from a set of second motion vector candidates.

[0199] In step S1003, motion compensation processing is performed to obtain the first partition using the first motion vector derived in step S1002 above and to obtain the second partition using the second motion vector derived in step S1002 above.

[0200] In step S1004, prediction processing is performed on the image block as the (reconfigured) combination of the first partition and the second partition. The prediction processing includes boundary smoothing processing for smoothing the boundary between the first partition and the second partition. For example, the boundary smoothing processing involves weighting a plurality of first values of a plurality of boundary pixels predicted based on the first partition and a plurality of second values of a plurality of boundary pixels predicted based on the second partition. Various implementations of the boundary smoothing processing will be described later with reference to FIGS. 13, 14, 20, and 21A to 21D.

[0201] In step S1005, the processing encodes or decodes the image block using one or more parameters including a first partition having a non-rectangular shape and a partition parameter indicating dividing the image block into the second partition. As summarized in the table of FIG. 15, for example, the partition parameter ("first index value") may encode, for example, the division direction applied to the division (e.g., from top left to bottom right or from top right to bottom left as shown in FIG. 12) and the first motion vector and the second motion vector derived in step S1002 described above, together. Details of such partition syntax operations with one or more parameters including the partition parameter will be described in detail later with reference to FIGS. 15, 16, and 22 to 25.

[0202] FIG. 17 is a flowchart showing a process 2000 of dividing an image block. In step S2001, the process divides an image into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. As shown in FIG. 12, the image block is divided into a first partition having a triangular shape and a second partition also having a triangular shape. There are many other examples in which an image block is divided into a plurality of partitions including a first partition and a second partition, where at least the first partition has a non-rectangular shape. The non-rectangular shape may be a triangle, a trapezoid, or a polygon having at least five sides and angles.

[0203] For example, as shown in FIG. 18, the image block may be divided into two triangular-shaped partitions. The image block may be divided into more than two triangular-shaped partitions (e.g., three triangular-shaped partitions). The image block may be divided into a combination of one or more triangular-shaped partitions and one or more rectangular-shaped partitions. Alternatively, the image block may be divided into a combination of one or more triangular-shaped partitions and one or more polygon-shaped partitions.

[0204] Furthermore, as shown in FIG. 19, the image block may be divided into an L-shaped (polygon-shaped) partition and a rectangular-shaped partition. The image block may be divided into a pentagon (polygon) - shaped partition and a triangular-shaped partition. The image block may be divided into a hexagon (polygon) - shaped partition and a pentagon (polygon) - shaped partition. Alternatively, the image block may be divided into a plurality of polygon-shaped partitions.

[0205] Referring again to FIG. 17, in step S2002, for the first partition, the process predicts the first motion vector by, for example, selecting the first partition from a first set of motion vector candidates, and for the second partition, the process predicts the second motion vector by, for example, selecting the second partition from a second set of motion vector candidates. For example, the first set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the first partition, and the second set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the second partition. The plurality of adjacent partitions may be one or both of a plurality of spatially adjacent partitions and a plurality of temporally adjacent partitions. Some examples of the plurality of spatially adjacent partitions include partitions located to the left, lower left, lower, lower right, right, upper right, upper, or upper left of the partition being processed. Some examples of the plurality of temporally adjacent partitions include a plurality of co-located partitions in a plurality of reference pictures of the image block.

[0206] In various implementations, the plurality of partitions adjacent to the first partition and the plurality of partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The first set of motion vector candidates may be the same as or different from the second set of motion vector candidates. Further, at least one of the first set of motion vector candidates and the second set of motion vector candidates may be the same as another third set of motion vector candidates prepared for the image block.

[0207] In some implementations, in step S2002, in response to a determination that the second partition, like the first partition, also has a non-rectangular shape (e.g., triangular), process 2000 creates (for the non-rectangular shaped second partition) a second set of motion vector candidates that includes the motion vectors of a plurality of partitions adjacent to the second partition, excluding the first partition (i.e., excluding the motion vector of the first partition). On the other hand, in response to a determination that the second partition has a rectangular shape and is not the same as the first partition, process 2000 creates (for the rectangular shaped second partition) a second set of motion vector candidates that includes the motion vectors of a plurality of partitions adjacent to the second partition, including the first partition.

[0208] In step S2003, the process encodes or decodes the first partition using the first motion vector derived in step S2002 described above, and encodes or decodes the second partition using the second motion vector derived in step S2002 described above.

[0209] An image block splitting process such as process 2000 in FIG. 17 may be performed by an image encoding apparatus including, for example, a circuit and a memory connected to the circuit as shown in FIG. 1. The circuit, in operation, splits an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0210] According to another embodiment, as shown in FIG. 1, an image encoding apparatus includes a splitting unit 102 that receives an original image in operation and splits it into a plurality of blocks, an adding unit 104 that receives a plurality of blocks from the splitting unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the adding unit 104 in operation and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bitstream, and an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, and an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in a current picture in operation, and a prediction control unit 128 connected to memories 118 and 122, and is provided. The prediction control unit 128 splits a plurality of blocks into a plurality of partitions including a first partition having a non-rectangular shape and a second partition in operation (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0211] According to another embodiment, an image decoding apparatus including a circuit and a memory connected to the circuit as shown in FIG. 10, for example, is provided. The circuit performs, in operation, splitting an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (FIG. 17, step S2001), predicting a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decoding the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0212] According to a further embodiment, as shown in FIG. 10, the image decoding apparatus receives a coded bitstream in operation, decodes it to obtain a plurality of quantized transform coefficients, and an entropy decoding unit 202. In operation, an inverse quantization unit 204 that inverse quantizes a plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and an inverse transform unit 206 that inverse transforms the plurality of transform coefficients to obtain a plurality of residuals. In operation, an addition unit 208 that adds a plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and a plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, and an inter prediction unit 218 that generates a prediction of a current block based on a reference block in a decoded reference picture in operation, and an intra prediction unit 216 that generates a prediction of the current block based on a decoded reference block in the current picture in operation, and a prediction control unit 220 connected to memories 210 and 214 are provided. The prediction control unit 220 divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition in operation (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0213] [Boundary Smoothing] As described above in FIG. 11, step S1004, according to various embodiments, performing prediction processing on an image block as a (reconstructed) combination of a first partition having a non-rectangular shape and a second partition may involve applying a boundary smoothing process along the boundary between the first partition and the second partition.

[0214] For example, FIG. 21B shows an example of a boundary smoothing process involving weighting a plurality of first values of a plurality of first predicted boundary pixels based on a first partition and a plurality of second values of a plurality of second predicted boundary pixels based on a second partition.

[0215] Figure 20 is a flowchart showing an overall boundary smoothing process 3000 involving weighting a plurality of first values of a plurality of first predicted boundary pixels based on a first partition and a plurality of second values of a plurality of second predicted boundary pixels based on a second partition. In step S3001, as shown in FIG. 21A or as shown in FIGS. 12, 18, and 19 described above, an image block is divided along a boundary into a first partition and a second partition, where at least the first partition has a non-rectangular shape.

[0216] In step S3002, a plurality of first values (e.g., color, luminance, transparency, etc.) of a pixel set of the first partition (the "plurality of boundary pixels" in FIG. 21A) along the boundary are first predicted using the information of the first partition. In step S3003, a plurality of second values of the (same) pixel set of the first partition along the boundary are second predicted using the information of the second partition. In some embodiments, at least one of the first prediction and the second prediction is an inter-prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in the encoded reference picture. Referring to FIG. 21D, in some implementations, the prediction process predicts a plurality of first values of all pixels of a first partition (the "first sample set") including a pixel set where the first partition and the second partition overlap, and predicts the second values only of a pixel set (the "second sample set") where the first partition and the second partition overlap. In other implementations, at least one of the first prediction and the second prediction is an intra-prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in the current picture. In some implementations, the prediction method used for the first prediction is different from the prediction method used for the second prediction. For example, the first prediction may include an inter-prediction process, or the second prediction may include an intra-prediction process. The information used for the first prediction of the plurality of first values or the second prediction of the plurality of second values may be a plurality of motion vectors, a plurality of intra-prediction directions, etc. of the first partition or the second partition.

[0217] In step S3004, the plurality of first values predicted using the first partition and the plurality of second values predicted using the second partition are weighted. In step S3005, the first partition is encoded or decoded using the weighted plurality of first values and the plurality of second values.

[0218] FIG. 21B shows an example of a boundary smoothing operation in which the first partition and the second partition overlap by (at most) 5 pixels per row or per row. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 5. FIG. 21C shows another example of a boundary smoothing operation in which the first partition and the second partition overlap by (at most) 3 pixels per row or per column. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 3.

[0219] FIG. 13 shows another example of a boundary smoothing operation in which the first partition and the second partition overlap by (at most) 4 pixels per row or per column. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 4. In the example shown, a plurality of weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the plurality of first values of the 4 pixels in the set, respectively, and a plurality of weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the plurality of second values of the 4 pixels in the set, respectively.

[0220] FIG. 14 further shows a plurality of examples of boundary smoothing operations where the first partition and the second partition overlap at 0 pixels per row or column (i.e., they do not overlap), overlap at (maximum) 1 pixel per row or column, and overlap at (maximum) 2 pixels per row or column. In an example where the first partition and the second partition do not overlap, a plurality of zero weights are applied. In an example where the first partition and the second partition overlap at 1 pixel per row or column, a weight of 1 / 2 may be applied to a plurality of first values of a plurality of pixels in a set predicted based on the first partition, and a weight of 1 / 2 may be applied to a plurality of second values of a plurality of pixels in a set predicted based on the second partition. In an example where the first partition and the second partition overlap at 2 pixels per row or column, weights of 1 / 3 and 2 / 3 may be applied respectively to a plurality of first values of two pixels in a set predicted based on the first partition, and weights of 2 / 3 and 1 / 3 may be applied respectively to a plurality of second values of two pixels in a set predicted based on the second partition.

[0221] According to the plurality of embodiments described above, the number of a plurality of pixels in a set where the first partition and the second partition overlap is an integer. In other implementations, the number of overlapping pixels in the set may be, for example, non-integer or fractional. The plurality of weights applied to the plurality of first values and the plurality of second values of the pixel set may also be fractional or integer according to each application.

[0222] A boundary smoothing process such as process 3000 in FIG. 20 may be performed by, for example, an image encoding apparatus including a circuit and a memory connected to the circuit as shown in FIG. 1. In operation, the circuit performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block (FIG. 20, step S3001). The boundary smoothing operation includes predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition (step S3002), second-predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).

[0223] According to another embodiment, as shown in FIG. 1, an image encoding apparatus includes a splitting unit 102 that receives an original image in operation and splits it into a plurality of blocks, an adding unit 104 that receives a plurality of blocks from the splitting unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the adding unit 104 in operation and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bitstream, an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in a current picture in operation, and a prediction control unit 128 connected to memories 118 and 122. The prediction control unit 128 performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition that is split from an image block in operation (FIG. 20, step S3001). The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).

[0224] According to another embodiment, for example, an image decoding apparatus including a circuit and a memory connected to the circuit, as shown in FIG. 10, is provided. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition (FIG. 20, step S3001). The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).

[0225] According to another embodiment, the image decoding apparatus shown in FIG. 10 includes an entropy decoding unit 202 that receives and decodes an encoded bit stream in operation to obtain a plurality of quantized transform coefficients, an inverse quantization unit 204 that inverse quantizes the plurality of quantized transform coefficients in operation to obtain a plurality of transform coefficients, and inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, an inverse transform unit 206, an addition unit 208 that adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and the plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter prediction unit 218 that generates a prediction of a current block based on a reference block in a decoded reference picture in operation, an intra prediction unit 216 that generates a prediction of the current block based on a decoded reference block in the current picture in operation, and a prediction control unit 220 connected to memories 210 and 214. The prediction control unit 220 performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition that are divided from an image block in operation (FIG. 20, step S3001). The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).

[0226] [Entropy Encoding and Decoding Using Partition Parameter Syntax] As shown in FIG. 11 and step S1005, according to various embodiments, the first partition having a non-rectangular shape and the image block divided into the second partition may be encoded or decoded using one or more parameters including partition parameters indicating the non-rectangular division of the image block. In various embodiments, such partition parameters may, as will be described in more detail later, for example, encode the division direction applied to the division (e.g., from the upper left to the lower right, or from the upper right to the lower left, see FIG. 12) and the first motion vector and the second motion vector predicted in step S1002 as a whole.

[0227] FIG. 15 is a table of a plurality of sample partition parameters ("first index values") and a plurality of information sets each encoded as a whole by the plurality of partition parameters. The plurality of partition parameters ("first index values") range from 0 to 6 and encode as a whole the direction of dividing the image block into the first partition and the second partition, both of which are triangular (see FIG. 12), the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002). In particular, partition parameter 0 encodes that the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0228] Partition parameter 1 encodes that the partitioning direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the partitioning direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the partitioning direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the partitioning direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the partitioning direction is from the upper left corner to the lower right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 6 encodes that the partitioning direction is from the upper left corner to the lower right corner, the first motion vector is the "fourth" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0229] FIG. 22 is a flowchart showing a method 4000 performed on the encoder side. In step S4001, the process divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the division. For example, as shown in FIG. 15 described above, the partition parameter may indicate the direction of dividing the image block (e.g., from the upper right corner to the lower left corner, or from the upper left corner to the lower right corner). In step S4002, the process encodes the first partition and the second partition. In step S4003, the process writes one or more parameters including the partition parameter into a bitstream that can be received and decoded so that the decoder side can obtain one or more parameters and perform the same prediction process on the first partition and the second partition on the decoder side (as performed on the encoder side). One or more parameters including the partition parameter encode various information such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used for dividing the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., either integrally or separately.

[0230] FIG. 23 is a flowchart showing a method 5000 performed on the decoder side. In step S5001, the process reads one or more parameters from a bitstream including partition parameters indicating that an image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition. One or more parameters including the partition parameters read from the bitstream may encode various information such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used for dividing the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., which are required for the decoder side to perform the same prediction process as that performed on the encoder side, either integrally or separately. In step S5002, the process 5000 divides the image block into a plurality of partitions based on the partition parameters read from the bitstream. In step S5003, the process decodes the first partition and the second partition, which are divided from the image block.

[0231] FIG. 24 is a table of a sample table whose characteristics are similar to those described above in FIG. 15, and a plurality of information sets integrally encoded by a plurality of sample partition parameters ("first index values") and a plurality of partition parameters respectively. In FIG. 24, the partition parameter ("first index value") ranges from 0 to 6, and encodes the shapes of the first partition and the second partition divided from the image block, the direction of dividing the image block into the first partition and the second partition, the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002) integrally. In particular, partition parameter 0 encodes that neither the first partition nor the second partition has a triangular shape, and thus the division direction information is "N / A", the first motion vector information is "N / A", and the second motion vector information is "N / A".

[0232] Partition parameter 1 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition.The partition parameter 6 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0233] According to some implementations, multiple partition parameters (multiple index values) may be binarized according to a binarization method selected according to at least one or the values of one or more parameters. FIG. 16 shows an example of a binarization method for binarizing multiple index values (multiple partition parameter values).

[0234] FIG. 25 is a table of an example combination of a first parameter and a second parameter, where one of the first parameter and the second parameter is a partition parameter indicating dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition. In this example, the partition parameter may be used to indicate dividing the image block without integrally encoding other information encoded by one or more of the other multiple parameters.

[0235] In the first example in FIG. 25, the first parameter is used to indicate the image block size, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the plurality of partitions divided from the image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that there is no triangular-shaped partition when the image block size is larger than 64×64, or 2) to indicate that there is no triangular-shaped partition when the ratio of the width to the height of the image block is larger than 4 (for example, 64×4).

[0236] In the second example of FIG. 25, the first parameter is used to indicate a prediction mode, and the second parameter is used as a partition parameter (flag) to indicate that at least one of a plurality of partitions divided from an image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that there is no triangular partition when the image block is encoded in an intra mode.

[0237] In the third example of FIG. 25, the first parameter is used as a partition parameter (flag) to indicate that at least one of a plurality of partitions divided from an image block has a triangular shape, and the second parameter is used to indicate a prediction mode. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the image block must be inter-encoded when at least one of a plurality of partitions divided from the image block has a triangular shape.

[0238] In the fourth example of FIG. 25, the first parameter indicates a motion vector of an adjacent block, and the second parameter is used as a partition parameter indicating a direction in which the image block is divided into two triangles. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the direction in which the image block is divided into two triangles is from the upper left corner to the lower right corner when the motion vector of the adjacent block is in an oblique direction.

[0239] In the fifth example of FIG. 25, the first parameter indicates an intra prediction direction of an adjacent block, and the second parameter is used as a partition parameter indicating a direction in which the image block is divided into two triangles. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the direction in which the image block is divided into two triangles is from the upper right corner to the lower left corner when the intra prediction direction of the adjacent block is in an inverse oblique direction.

[0240] Multiple tables of one or more parameters including partition parameters, and as shown in FIGS. 15, 24, and 25, which information is encoded together or separately is presented only as multiple examples, and it should be understood that numerous other ways of encoding various information together or separately as part of the partition syntax operations described above are within the scope of this disclosure. For example, the partition parameter may indicate that the first partition is a triangle, trapezoid, or polygon having at least five sides and angles. The partition parameter may indicate that the second partition has a non-rectangular shape such as a triangle, trapezoid, and polygon having at least five sides and angles. The partition parameter may indicate one or more pieces of information about the partition, such as the non-rectangular shape of the first partition, the shape of the second partition (which may be non-rectangular or rectangular), the partitioning direction applied to divide the image block into multiple partitions (e.g., from the upper left corner to the lower right corner of the image block, and from the upper right corner to the lower left corner of the image block). The partition parameter encodes additional information such as the first motion vector of the first partition, the second motion vector of the second partition, the image block size, the prediction mode, the motion vectors of adjacent blocks, the intra prediction direction of adjacent blocks, etc. together. Alternatively, any of the additional information may be encoded separately by one or more parameters other than the partition parameter.

[0241] A partition syntax operation such as process 4000 of FIG. 22 may be performed by an image encoding apparatus including a circuit and a memory connected to the circuit, as shown in FIG. 1 for example. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on partition parameters indicating the division (FIG. 22, step S4001), encodes the first partition and the second partition (S4002), and performs a partition syntax operation including writing one or more parameters including the partition parameters into a bitstream (S4003).

[0242] According to another embodiment, as shown in FIG. 1, an image encoding apparatus includes a division unit 102 that receives an original image in operation and divides it into a plurality of blocks, an addition unit 104 that receives a plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the addition unit 104 in operation and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bitstream, and an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in the current picture in operation, and a prediction control unit 128 connected to memories 118 and 122, and is provided. In operation, the prediction control unit 128 divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on partition parameters indicating the division (FIG. 22, step S4001), and encodes the first partition and the second partition (step S4002). The entropy encoding unit 110 writes one or more parameters including the partition parameters into a bitstream in operation (step S4003).

[0243] According to another embodiment, for example, an image decoding apparatus including a circuit and a memory connected to the circuit, as shown in FIG. 10, is provided. The circuit, in operation, reads from the bitstream one or more parameters including a partition parameter indicating to divide an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (FIG. 23, step S5001), divides the image block into the plurality of partitions based on the partition parameter (S5002), and performs a partition syntax operation including decoding the first partition and the second partition (S5003).

[0244] According to a further embodiment, the image decoding apparatus shown in FIG. 10 includes, in operation, an entropy decoding unit 202 that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, an inverse quantization unit 204 that, in operation, inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and an inverse transform unit 206 that inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, an addition unit 208 that, in operation, adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and the plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, and an inter prediction unit 218 that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit 216 that, in operation, generates a prediction of a current block based on a decoded reference block in a current picture, and a prediction control unit 220 connected to memories 210 and 214. The entropy decoding unit 202, in operation, in some implementations, cooperates with the prediction control unit 220 to read, from the bitstream, one or more parameters including partition parameters indicating that an image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (FIG. 23, step S5001), divides the image block into a plurality of partitions based on the partition parameters (S5002), and decodes the first partition and the second partition (S5003).

[0245] [Implementation and Application] In each of the above embodiments, each of the functional or operative blocks can usually be realized by an MPU (micro processing unit), a memory, and the like. Further, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (program) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as a semiconductor memory. It is also possible to realize each functional block by hardware (a dedicated circuit). Various combinations of hardware and software can be adopted.

[0246] The processes described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing may be performed, or distributed processing may be performed.

[0247] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible, and these are also included within the scope of the aspects of the present disclosure.

[0248] Furthermore, here, application examples of the moving image encoding method (image encoding method) or moving image decoding method (image decoding method) shown in each of the above embodiments, and various systems for implementing such application examples will be described. Such a system may be characterized by having an image encoding device using an image encoding method, an image decoding device using an image decoding method, or an image encoding / decoding device having both. Other configurations of such a system can be appropriately changed as the case may be.

[0249] [Usage Example] FIG. 26 is a diagram showing the overall configuration of a suitable content supply system ex100 for realizing a content distribution service. The communication service providing area is divided into a desired size, and in each cell, base stations ex106, ex107, ex108, ex109, ex110, which are fixed radio stations in the illustrated example, are installed.

[0250] In this content supply system ex100, devices such as computer ex111, game console ex112, camera ex113, home appliances ex114, and smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may be connected by combining any of the above devices. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network or short-range wireless without passing through base stations ex106 to ex110. Further, the streaming server ex103 may be connected to devices such as computer ex111, game console ex112, camera ex113, home appliances ex114, and smartphone ex115 via the Internet ex101 or the like. Also, the streaming server ex103 may be connected to terminals within a hotspot in an airplane ex117 via a satellite ex116.

[0251] Note that a wireless access point or hotspot or the like may be used instead of base stations ex106 to ex110. Also, the streaming server ex103 may be directly connected to the communication network ex104 without passing through the Internet ex101 or Internet service provider ex102, or may be directly connected to the airplane ex117 without passing through the satellite ex116.

[0252] The camera ex113 is a device capable of taking still images and videos such as a digital camera. Also, the smartphone ex115 is a smartphone device, mobile phone, or PHS (Personal Handy-phone System) or the like corresponding to the mobile communication system methods called 2G, 3G, 3.9G, 4G, and in the future 5G.

[0253] The home appliances ex114 are devices included in a refrigerator or a household fuel cell cogeneration system or the like.

[0254] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 through a base station ex106 or the like, enabling live distribution and the like. In live distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) may perform the encoding process described in each of the above embodiments on a still image or moving image content photographed by the user using the terminal. The video data obtained by encoding and the audio data obtained by encoding the audio corresponding to the video may be multiplexed, and the obtained data may be transmitted to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0255] On the other hand, the streaming server ex103 stream-distributes the content data transmitted to the requested client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117 that is capable of decoding the encoded data. Each device that has received the distributed data may decode and reproduce the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0256] [Distributed processing] In addition, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, record, and deliver data. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world and between the edge servers. In a CDN, a physically closer edge server can be dynamically assigned according to the client. By caching and delivering the content to the edge server, the delay can be reduced. Also, when several types of errors occur or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the delivery entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the delivery, so that high-speed and stable delivery can be realized.

[0257] Moreover, not only the distributed processing of the delivery itself, but also the encoding process of the captured data may be performed on each terminal, on the server side, or shared between them. As an example, generally in the encoding process, the processing loop is performed twice. In the first loop, the complexity of the image in units of frames or scenes, or the amount of codes is detected. Also, in the second loop, processing is performed to improve the encoding efficiency while maintaining the image quality. For example, by having the terminal perform the first encoding process and the server side that receives the content perform the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode in almost real time, the already encoded data processed by the terminal can also be received and played back by other terminals, enabling more flexible real-time delivery.

[0258] As another example, cameras such as camera ex113 extract feature amounts (amounts of features or characteristics) from images, compress data related to the feature amounts as metadata, and transmit the data to a server. The server performs compression according to the meaning (or importance of content) of the image, for example, by determining the importance of an object from the feature amounts and switching the quantization accuracy. Feature amount data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.

[0259] As yet another example, in a stadium, a shopping mall, or a factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using the plurality of terminals that have performed the shooting, and other terminals and a server that have not performed the shooting as necessary, encoding processes are respectively assigned and distributed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing a picture. Thereby, delay can be reduced and more real-time performance can be realized.

[0260] Since the plurality of video data are of substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can refer to each other. Also, the server may receive the encoded data from each terminal, change the reference relationship between the plurality of data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.

[0261] Furthermore, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based (for example, VP9) method, or convert H.264 to H.265, etc.

[0262] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" will be used as the entity performing the process. However, part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.

[0263] [3D, Multi - angle] There is an increasing trend to integrate and utilize different scenes captured by a plurality of cameras ex113 and / or terminals such as smartphone ex115 that are substantially synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal can be integrated based on the relative positional relationship between the terminals obtained separately or the regions where the feature points included in the videos match.

[0264] The server may not only encode a two - dimensional moving image, but also automatically encode a still image based on scene analysis of the moving image or at a time specified by the user, and transmit it to the receiving terminal. When the server can obtain the relative positional relationship between the shooting terminals, it can generate the three - dimensional shape of the scene based not only on the two - dimensional moving image but also on videos of the same scene captured from different angles. The server may separately encode the three - dimensional data generated by a point cloud or the like, or generate the video to be transmitted to the receiving terminal by selecting or reconstructing from the videos captured by a plurality of terminals based on the results of recognizing or tracking a person or an object using the three - dimensional data.

[0265] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or can also enjoy the content obtained by cutting out the video from the selected viewpoint from the three - dimensional data reconstructed using a plurality of images or videos. Furthermore, sounds are also collected from a plurality of different angles together with the video, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0266] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become widespread. In the case of VR images, the server may create viewpoint images for the right eye and the left eye, respectively, and perform encoding that allows reference between each viewpoint video by means of Multi-View Coding (MVC) or the like, or may perform encoding as separate streams without referring to each other. At the time of decoding the separate streams, they may be played back in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0267] In the case of AR images, the server may superimpose virtual object information in the virtual space on the camera information of the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to the request for virtual object information. The server may create superimposed data in accordance with the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data typically has an α value indicating transparency in addition to RGB, and the server may encode it with the α value of the portion other than the object created from the three-dimensional data set to 0 or the like so that the portion is in a transparent state. Alternatively, the server may generate data in which an RGB value of a predetermined value is set as the background like chroma key, and the portion other than the object is the background color. The RGB value of the predetermined value may be predetermined.

[0268] The decoding process of the data delivered in the same way may be performed on the client (e.g., a terminal), on the server side, or they may be shared between each other. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request is received by another terminal for decoding, and the decoded signal may be transmitted to a device having a display. By dispersing the processing regardless of the performance of the communicable terminals themselves and selecting appropriate content, it is possible to reproduce data with good image quality. As another example, while receiving large-size image data on a TV or the like, only some areas such as tiles in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to check at hand the area in one's own field of responsibility or the area to be confirmed in more detail.

[0269] Under a situation where multiple short-range, medium-range, or long-range wireless communications inside and outside the house are available, it may be possible to receive content seamlessly using a delivery system standard such as MPEG-DASH. The user may freely select a decoding device or a display device such as the user's terminal or a display arranged inside and outside the house and switch them in real time. Also, using the user's own location information or the like, decoding can be performed while switching the terminal to be decoded and the terminal to be displayed. Thereby, while the user is moving to the destination, it becomes possible to map and display information on a part of the wall surface or the ground of the adjacent building in which a displayable device is embedded. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time or being copied to an edge server in a content delivery service, it is also possible to switch the bit rate of the received data.

[0270] [Scalable Encoding] Regarding content switching, an explanation will be given using a scalable stream compressed and encoded by applying the moving image encoding method shown in each of the above embodiments and shown in FIG. 27. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of the time / spatial scalable stream realized by encoding in layers as shown in the figure, it may be configured to switch content. That is, by determining which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth on the decoding side, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when a user wants to watch the continuation of a video that was being viewed on a smartphone ex115 while moving on a device such as an Internet TV after returning home, the device only needs to decode the same stream to a different layer, thus reducing the burden on the server side.

[0271] Furthermore, as described above, pictures are encoded for each layer. In addition to the configuration that realizes scalability in the enhancement layer above the base layer, the enhancement layer may include meta information based on statistical information of the image, etc. The decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. Super-resolution may improve the signal-to-noise ratio while maintaining and / or expanding the resolution. The meta information includes information for specifying linear or non-linear filter coefficients for use in super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used in super-resolution processing.

[0272] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of an object or the like in the image. The decoding side decodes only a part of the area by selecting the tile to be decoded. Further, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of a desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 28, the meta information may be stored using a data storage structure different from the pixel data, such as an SEI (supplemental enhancement information) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0273] The meta information may be stored in a unit composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. The decoding side can obtain the time when a specific person appears in the video, etc., and by combining the picture unit information and the time information, can specify the picture in which the object exists and determine the position of the object in the picture.

[0274] [Optimization of Web Page] FIG. 29 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 30 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 29 and 30, a web page may include a plurality of link images that are links to image contents, and the appearance thereof may be different depending on the device for browsing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) may display a still image or an I picture that each content has as a link image, or may display a video like a gif animation with a plurality of still images or I pictures, etc., or may receive only the base layer, decode and display the video.

[0275] When a user selects a linked image, the display device decodes, for example, with the base layer having the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Furthermore, in order to ensure real-time performance, before selection or when the communication bandwidth is extremely tight, the display device can reduce the delay (the delay from the start of content decoding to the start of display) between the decoding time and the display time of the leading picture by decoding and displaying only forward-reference pictures (I pictures, P pictures, B pictures with only forward reference). Additionally, the display device may deliberately ignore the reference relationships of the pictures, perform rough decoding with all B pictures and P pictures as forward references, and perform normal decoding as the received pictures increase over time.

[0276] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for the autonomous driving or driving assistance of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.

[0277] In this case, since a vehicle, drone, airplane, etc. including the receiving terminal moves, the receiving terminal can realize seamless reception and decoding by transmitting the position information of the receiving terminal while switching between base stations ex106 to ex110. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0278] In the content supply system ex100, the client can receive, in real time, the encoded information transmitted by the user, decode it, and play it back.

[0279] [Distribution of Personal Content] In addition, in the content supply system ex100, not only high-quality and long-duration content by video distributors but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. It is considered that such individual content will increase in the future. In order to make individual content into better content, the server may perform encoding processing after performing editing processing. This can be realized, for example, by using the following configuration.

[0280] At the time of shooting in real time or storing and shooting later, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection from the original picture data or encoded data. Then, based on the recognition result, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of the object, or changes the color tone. The server encodes the edited data based on the editing result. It is also known that the viewing rate decreases if the shooting time is too long. The server may automatically clip not only less important scenes but also scenes with little movement within a specific time range according to the shooting time so as to become content within that range, based on the image processing result. Or, the server may generate and encode a digest based on the result of semantic analysis of the scene.

[0281] Personal content may contain elements that, as they are, would infringe copyrights, moral rights of the author, or the right of portrait, etc., and there may be inconvenient situations for individuals, such as the sharing scope exceeding the intended scope. Therefore, for example, the server may deliberately change the image so that the face of a person in the peripheral part of the screen or the inside of a house is out of focus and then encode it. Furthermore, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Alternatively, as pre-processing or post-processing of encoding, the user may specify a person or a background area that the user wants to process the image from the perspective of copyright, etc. The server may perform processing such as replacing the specified area with another video or blurring the focus. In the case of a person, the person can be tracked in a moving image, and the video of the face part of the person can be replaced.

[0282] Since the viewing of personal content with a small data volume has a strong requirement for real-time performance, depending on the bandwidth, the decoding device may first receive the base layer with the highest priority and perform decoding and playback. During this period, the decoding device may receive the enhancement layer and, when the playback is looped or played back two or more times, play back a high-quality video including the enhancement layer. For a stream with scalable encoding like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played back for the first time and a second stream encoded with reference to the first video are configured as one stream.

[0283] [Other Implementation and Application Examples] Also, these encoding or decoding processes are generally processed in the LSIex500 possessed by each terminal. The LSI (large scale integration circuitry) ex500 (see FIG. 26) may be a one-chip configuration or a configuration consisting of multiple chips. In addition, software for moving image encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by a computer ex111 or the like, and encoding or decoding processing may be performed using the software. Further, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time may be data encoded by the LSIex500 possessed by the smartphone ex115.

[0284] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal may download a codec or application software and then acquire and play the content.

[0285] Also, not limited to the content supply system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above embodiments can be incorporated into a digital broadcast system. In order to transmit and receive multiplexed data in which video and audio are multiplexed on a broadcast radio wave using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the unicast-friendly configuration of the content supply system ex100, but the same application is possible for encoding and decoding processes.

[0286] [Hardware Configuration] FIG. 31 is a diagram showing further details of the smartphone ex115 shown in FIG. 26. Also, FIG. 32 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying data obtained by decoding videos captured by the camera unit ex465 and videos received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured videos or still images, recorded audio, received videos or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with the SIM ex468 for identifying the user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.

[0287] A main control unit ex460 capable of comprehensively controlling the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a synchronous bus ex470.

[0288] When the power key is turned on by the user's operation, the power supply circuit unit ex461 activates the smartphone ex115 to an operable state and supplies power to each unit from the battery pack.

[0289] The smartphone ex115 performs processes such as calls and data communication based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, subjected to spread spectrum processing by the modulation / demodulation unit ex452, and subjected to digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, and the resulting signal is transmitted via the antenna ex450. Also, received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, subjected to inverse spread spectrum processing by the modulation / demodulation unit ex452, converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. In the data communication mode, text, still images, or video data can be sent under the control of the main control unit ex460 via the operation input control unit ex462 based on operations of the operation unit ex466 of the main body unit or the like. Similar transmission and reception processes are performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while the camera unit ex465 is imaging a video or a still image, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450. The predetermined manner may be determined in advance.

[0290] When receiving a video attached to an email or chat, or a video linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and the video or still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. Since real-time streaming is becoming increasingly popular, depending on the user's situation, it may not be socially appropriate to play the audio. Therefore, as an initial value, it is desirable to have a configuration that plays only the video data without playing the audio signal, and the audio may be played synchronously only when the user performs an operation such as clicking on the video data.

[0291] Also, although the smartphone ex115 has been described as an example here, as the terminal, in addition to the transceiver type terminal having both an encoder and a decoder, other implementation forms such as a transmitting terminal having only an encoder and a receiving terminal having only a decoder are conceivable. In the digital broadcast system, it has been described as receiving or transmitting multiplexed data in which audio data is multiplexed with video data. However, in the multiplexed data, character data related to the video etc. may be multiplexed in addition to the audio data. Also, instead of the multiplexed data, the video data itself may be received or transmitted.

[0292] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, many types of terminals often have a GPU. Therefore, a configuration may be adopted in which a wide area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so that it can be commonly used. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures using the GPU instead of the CPU.

Claims

1. A circuit, and a memory connected to the circuit, wherein, in operation, the circuit performs processing along a boundary between a first partition having a non-rectangular shape and a second partition in an image block, and the processing includes: selecting a first motion vector of the first partition from a first set of motion vector candidates; using the first motion vector to perform a first prediction of a plurality of first values of a pixel set of the first partition; selecting a second motion vector of the second partition from the first set of motion vector candidates; using the second motion vector to perform a second prediction of a plurality of second values of the pixel set of the first partition; weighting the plurality of first values and the plurality of second values; wherein when a ratio of a width to a height of the image block is greater than 4, or when a ratio of the height to the width of the image block is greater than 4, the circuit invalidates the processing, An image encoding device.

2. A circuit, and a memory connected to the circuit, wherein, in operation, the circuit performs processing along a boundary between a first partition having a non-rectangular shape and a second partition in an image block, and the processing includes: selecting a first motion vector of the first partition from a first set of motion vector candidates; using the first motion vector to perform a first prediction of a plurality of first values of a pixel set of the first partition; selecting a second motion vector of the second partition from the first set of motion vector candidates; using the second motion vector to perform a second prediction of a plurality of second values of the pixel set of the first partition; weighting the plurality of first values and the plurality of second values; wherein when a ratio of a width to a height of the image block is greater than 4, or when a ratio of the height to the width of the image block is greater than 4, the circuit invalidates the processing, An image decoding device.

3. A bitstream generation device including a circuit, and a memory connected to the circuit, wherein, in operation, the circuit generates information regarding a size of the image block for causing a decoding device to execute partition processing along a boundary between a first partition having a non-rectangular shape and a second partition in the image block, includes the information in a bitstream, and the partition processing includes The first motion vector of the first partition in the image block is selected from a first set of motion vector candidates; A plurality of first values of a pixel set of the first partition are calculated using the first motion vector; The second motion vector of the second partition is selected from the first set of motion vector candidates; A plurality of second values of the pixel set are calculated using the second motion vector; A plurality of third values of the pixel set are calculated by weighting the plurality of first values and the plurality of second values, and When a ratio of a width of the image block to a height of the image block is greater than 4, or when a ratio of the height to the width is greater than 4, the bit stream generation device invalidates the partition processing. Bit stream generation device.

Citation Information

Patent Citations

  • Image processing device, and image processing method

    EP2592834A1

  • Image processing device and image processing method

    JP2012019490A

  • Smoothing of overlapping regions resulting from geometric motion subdivision.

    JP2013520877A

  • Adaptive transform size selection for geometric motion partitioning

    US20110200097A1

Cited By

  • Image encoder, image decoder, and bit stream generating device

    JP2025133798A