Image encoding device, image decoding device, and bitstream generation device

JP7866120B2Active Publication Date: 2026-05-26PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-06-30
Publication Date
2026-05-26

Smart Images

  • Figure 0007866120000003
    Figure 0007866120000003
  • Figure 0007866120000004
    Figure 0007866120000004
  • Figure 0007866120000005
    Figure 0007866120000005
Patent Text Reader

Abstract

To improve video coding technology.SOLUTION: An image encoder 100 comprises a circuit and a memory. The circuit performs, in operation, a process along a boundary between a first partition and a second partition. The process includes: encoding a partition parameter; selecting a first motion vector from a first set of motion vector candidates; first-predicting first values of a set of pixels, using the first motion vector; selecting a second motion vector from the first set of motion vector candidates; second-predicting second values of the set of pixels, using the second motion vector; and weighting the first values and the second values. When the ratio of width and height of an image block is larger than 4 or when the ratio of the height and width is larger than 4, the circuit disables the process.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to video coding, for example, systems, components, and methods for encoding and decoding moving images, which construct a current block based on a reference frame by performing interpretation, or construct a current block based on an encoded / decoded reference block in the current frame by performing intrapretation. [Background technology]

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this advancement, there is a constant need to provide improvements and optimizations to video coding technology to handle the ever-increasing volume of digital video data in various applications. This disclosure further relates to advancements, improvements, and optimizations in video coding, and more particularly to interpretation or intraprediction of dividing an image block into multiple partitions, each partition including at least a first partition having a non-rectangular shape (e.g., a triangle) and a second partition. [Overview of the Initiative]

[0003] According to one embodiment, an image coding device is provided comprising a circuit and a memory connected to the circuit. The circuit, in operation, performs processing along the boundary between a first partition having a non-rectangular shape and a second partition in an image block, the processing including coding partition parameters relating to the shapes of the first and second partitions; selecting a first motion vector of the first partition from a first motion vector candidate set; first predicting a plurality of first values ​​of the pixel set of the first partition using the first motion vector; selecting a second motion vector of the second partition from the first motion vector candidate set; second predicting a plurality of second values ​​of the pixel set of the first partition using the second motion vector; and weighting the plurality of first values ​​and the plurality of second values, wherein the circuit disables the processing if the ratio of the width to the height of the image block is greater than 4, or if the ratio of the height to the width is greater than 4.

[0004] In some embodiments of the embodiments of this disclosure, coding efficiency may be improved, coding / decoding processes may be simplified, coding / decoding speeds may be accelerated, and appropriate components / operations used in coding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, and reference blocks, may be efficiently selected.

[0005] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. Benefits and / or advantages may be obtained individually by various embodiments and features of the specification and drawings, and it is not necessary to provide all of the various embodiments and features of the specification and drawings in order to obtain one or more such benefits and / or advantages.

[0006] Furthermore, comprehensive or specific embodiments may be implemented as systems, methods, integrated circuits, computer programs, storage media, or any combination thereof. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 is a block diagram showing the functional configuration of an encoding device according to an embodiment. [Figure 2] Figure 2 shows an example of block division. [Figure 3] Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. [Figure 4A] Figure 4A shows an example of the filter shape used in an ALF (Adaptive Loop Filter). [Figure 4B] Figure 4B shows another example of the filter shape used in ALF. [Figure 4C] Figure 4C shows another example of the filter shape used in ALF. [Figure 5A] Figure 5A shows the 67 intra-prediction modes in intra-prediction. [Figure 5B] Figure 5B is a flowchart illustrating the overview of the predictive image correction process using OBMC (overlapped block motion compensation). [Figure 5C] Figure 5C is a conceptual diagram illustrating the overview of the predictive image correction process using OBMC processing. [Figure 5D] Figure 5D shows an example of FRUC (frame rate up-conversion). [Figure 6] Figure 6 illustrates pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] Figure 7 illustrates pattern matching (template matching) between a template in the current picture and a block in the referenced picture. [Figure 8] Figure 8 is a diagram illustrating a model that assumes uniform linear motion. [Figure 9A]Figure 9A is a diagram illustrating the derivation of subblock-level motion vectors based on the motion vectors of multiple adjacent blocks. [Figure 9B] Figure 9B is a diagram illustrating the overview of the motion vector derivation process using merge mode. [Figure 9C] Figure 9C is a conceptual diagram illustrating the overview of DMVR (dynamic motion vector refreshing) processing. [Figure 9D] Figure 9D is a diagram illustrating the outline of a predictive image generation method using brightness correction processing by LIC (local illumination compensation). [Figure 10] Figure 10 is a block diagram showing the functional configuration of a decoding device according to an embodiment. [Figure 11] Figure 11 is a flowchart showing the overall processing flow for dividing an image block into multiple partitions, including at least a first partition and a second partition having a non-rectangular shape (e.g., a triangle), according to one embodiment, and then performing further processing. [Figure 12] Figure 12 shows two exemplary methods for dividing an image block into a first partition having a non-rectangular shape (e.g., a triangle) and a second partition (which also has a non-rectangular shape in the examples shown). [Figure 13] Figure 13 shows an example of boundary smoothing, which involves weighting multiple primary values ​​of multiple boundary pixels predicted based on the first partition with multiple secondary values ​​of multiple boundary pixels predicted based on the second partition. [Figure 14] Figure 14 shows three further examples of boundary smoothing, which involves weighting multiple primary values ​​of multiple boundary pixels predicted based on the first partition with multiple secondary values ​​of multiple boundary pixels predicted based on the second partition. [Figure 15]FIG. 15 is a table diagram showing a plurality of parameters (a plurality of first index values) and a plurality of information sets respectively encoded by the plurality of parameters. [Figure 16] FIG. 16 is a table diagram showing the binarization of a plurality of parameters (a plurality of index values). [Figure 17] FIG. 17 is a flowchart showing a process of dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition. [Figure 18] FIG. 18 is a diagram showing a plurality of examples of dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape which is a triangle in the plurality of shown examples and a second partition. [Figure 19] FIG. 19 is a diagram showing a further plurality of examples of dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape which is a polygon with at least five sides and angles in the plurality of shown examples and a second partition. [Figure 20] FIG. 20 is a flowchart showing a boundary smoothing process involving weighting a plurality of first values of a plurality of boundary images predicted based on a first partition and a plurality of second values of a plurality of boundary images predicted based on a second partition. [Figure 21A] FIG. 21A is a diagram showing an example of a boundary smoothing process in which a plurality of first weighted values of a plurality of boundary pixels are predicted based on a first partition and a plurality of second weighted values of a plurality of boundary pixels are predicted based on a second partition. [Figure 21B] FIG. 21B is a diagram showing an example of a boundary smoothing process in which a plurality of first weighted values of a plurality of boundary pixels are predicted based on a first partition and a plurality of second weighted values of a plurality of boundary pixels are predicted based on a second partition. [Figure 21C]Figure 21C shows an example of boundary smoothing, in which multiple weighted first values ​​for multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​for multiple boundary pixels are predicted based on a second partition. [Figure 21D] Figure 21D shows an example of boundary smoothing, in which multiple weighted primary values ​​for multiple boundary pixels are predicted based on a primary partition, and multiple weighted secondary values ​​for multiple boundary pixels are predicted based on a secondary partition. [Figure 22] Figure 22 is a flowchart illustrating a method performed by an encoding device to divide an image block into multiple partitions, including a first partition and a second partition, which have a non-rectangular shape, based on partition parameters indicating the division, and to write one or more parameters, including the partition parameters, to the bitstream during entropy coding. [Figure 23] Figure 23 is a flowchart illustrating a method performed by a decoding device that reads one or more parameters from a bitstream, including partition parameters that indicate dividing an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape, divides the image block into multiple partitions based on the partition parameters, and decodes the first and second partitions. [Figure 24] Figure 24 is a table diagram showing multiple partition parameters ("multiple first index values") that indicate dividing an image block into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape, and multiple information sets that can be encoded together by the multiple partition parameters. [Figure 25] Figure 25 is a table of several combinations of the first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is divided into multiple partitions, including a first partition and a second partition, both having a non-rectangular shape. [Figure 26] Figure 26 is a block diagram showing the overall configuration of the content supply system that realizes the content distribution service. [Figure 27] Figure 27 is a conceptual diagram showing an example of an encoding structure during scalable encoding. [Figure 28] Figure 28 is a conceptual diagram showing an example of an encoding structure during scalable encoding. [Figure 29] Figure 29 is a conceptual diagram showing an example of a web page display screen. [Figure 30] Figure 30 is a conceptual diagram showing an example of a web page display screen. [Figure 31] Figure 31 is a block diagram showing an example of a smartphone. [Figure 32] Figure 32 is a block diagram showing an example of a smartphone configuration. [Modes for carrying out the invention]

[0008] According to one embodiment, an image encoding device is provided comprising a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicts a first motion vector for the first partition and a second motion vector for the second partition; encodes the first partition using the first motion vector and encodes the second partition using the second motion vector.

[0009] In a further embodiment, the second partition has a non-rectangular shape. In another embodiment, the non-rectangular shape is triangular. In a further embodiment, the non-rectangular shape is selected from the group consisting of triangles, trapezoids, and polygons having at least five sides and angles.

[0010] In another embodiment, the prediction includes selecting the first motion vector from a first motion vector candidate set and selecting the second motion vector from a second motion vector candidate set. For example, the first motion vector candidate set may include multiple motion vectors of multiple partitions adjacent to the first partition, and the second motion vector candidate set may include multiple motion vectors of multiple partitions adjacent to the second partition. The multiple partitions adjacent to the first partition and the multiple partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The multiple adjacent partitions may be either or both of spatially adjacent partitions and / or temporally adjacent partitions. The first motion vector candidate set may be the same as or different from the second motion vector candidate set.

[0011] In another embodiment, the prediction includes selecting a first motion vector candidate from a first motion vector candidate set and deriving the first motion vector by adding a first difference motion vector to the first motion vector candidate, and selecting a second motion vector candidate from a second motion vector candidate set and deriving the second motion vector by adding a second difference motion vector to the second motion vector candidate.

[0012] In another embodiment, the present invention provides an image encoding device comprising: a division unit that, in operation, receives a source image and divides it into a plurality of blocks; an addition unit that, in operation, receives the plurality of blocks from the division unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a conversion unit that, in operation, performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients; a quantization unit that, in operation, quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients; an entropy encoding unit that, in operation, encodes the plurality of quantized conversion coefficients to generate a bitstream; an inter-prediction unit that, in operation, generates a prediction for the current block based on a reference block in a pre-encoded reference picture; an intra-prediction unit that, in operation, generates a prediction for the current block based on a pre-encoded reference block in the current picture; and the prediction control unit connected to memory. In operation, the prediction control unit divides the plurality of blocks into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicts a first motion vector for the first partition; predicts a second motion vector for the second partition; encodes the first partition using the first motion vector; and encodes the second partition using the second motion vector.

[0013] In another embodiment, an image encoding method is provided which includes dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicting a first motion vector for the first partition and predicting a second motion vector for the second partition; encoding the first partition using the first motion vector and encoding the second partition using the second motion vector.

[0014] According to one embodiment, an image decoding device is provided comprising a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicts a first motion vector for the first partition and a second motion vector for the second partition; decodes the first partition using the first motion vector and decodes the second partition using the second motion vector.

[0015] In a further embodiment, the second partition has a non-rectangular shape. In another embodiment, the non-rectangular shape is triangular. In a further embodiment, the non-rectangular shape is selected from the group consisting of triangles, trapezoids, and polygons having at least five sides and angles.

[0016] In another embodiment, an image decoding device is provided, comprising: an entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantization conversion coefficients; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantization of the plurality of quantization conversion coefficients to obtain a plurality of conversion coefficients and inverse transform the plurality of conversion coefficients to obtain a plurality of residuals; an addition unit that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; an inter-prediction unit that, in operation, generates a prediction of the current block based on a reference block in a decoded reference picture; an intra-prediction unit that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and the prediction control unit connected to memory. In operation, the prediction control unit divides an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicts a first motion vector for the first partition; predicts a second motion vector for the second partition; decodes the first partition using the first motion vector; and decodes the second partition using the second motion vector.

[0017] In another embodiment, an image decoding method is provided which includes dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predicting a first motion vector for the first partition and a second motion vector for the second partition; decoding the first partition using the first motion vector and decoding the second partition using the second motion vector.

[0018] According to one embodiment, an image coding device is provided comprising a circuit and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape and a second partition, which is divided from an image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and coding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0019] In a further embodiment, the non-rectangular shape is a triangle. In another embodiment, the non-rectangular shape is selected from the group consisting of triangles, trapezoids, and polygons having at least five sides and angles. In yet another embodiment, the second partition has a non-rectangular shape.

[0020] In another embodiment, at least one of the first and second predictions is an interpretation process that predicts the plurality of first values ​​and the plurality of second values ​​based on a reference partition in an encoded reference picture. The interpretation process may predict a plurality of first values ​​for a plurality of pixels in the first partition which includes the set of pixels, or it may predict a plurality of second values ​​for only the set of pixels in the first partition.

[0021] In another embodiment, at least one of the first and second predictions is an intra-prediction process that predicts the plurality of first and plurality of second values ​​based on encoded reference partitions in the current picture.

[0022] According to another embodiment, the prediction method used for the first prediction differs from the prediction method used for the second prediction.

[0023] In a further embodiment, the number of pixel sets in each row or column that predict the plurality of first values ​​and the plurality of second values ​​is an integer. For example, if the number of pixel sets in each row or column is four, then multiple weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the plurality of first values ​​of the four pixels in the pixel set, respectively, and multiple weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the plurality of second values ​​of the four pixels in the pixel set, respectively. As another example, if the number of pixel sets in each row or column is two, then multiple weights of 1 / 3 and 2 / 3 may be applied to the plurality of first values ​​of the two pixels in the pixel set, respectively, and multiple weights of 2 / 3 and 1 / 3 may be applied to the plurality of second values ​​of the two pixels in the pixel set, respectively.

[0024] In other embodiments, the weights may be integer values ​​or fractional values.

[0025] In another embodiment, an image encoding device is provided, comprising: a division unit that, in operation, receives a source image and divides it into a plurality of blocks; an addition unit that, in operation, receives the plurality of blocks from the division unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a conversion unit that, in operation, performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients; a quantization unit that, in operation, quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients; an entropy encoding unit that, in operation, encodes the plurality of quantized conversion coefficients to generate a bitstream; an inter-prediction unit that, in operation, generates a prediction for the current block based on a reference block in a pre-encoded reference picture; an intra-prediction unit that, in operation, generates a prediction for the current block based on a pre-encoded reference block in the current picture; and the prediction control unit connected to memory. The prediction control unit, in operation, performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape and a second partition, which are divided from the image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0026] In another embodiment, an image coding method is provided that performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition. The method typically includes four steps: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and coding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0027] In a further embodiment, an image decoding device is provided comprising a circuit and a memory connected to the circuit. The circuit, in operation, performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape and a second partition, which is divided from an image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0028] In another embodiment, the non-rectangular shape is a triangle. In yet another embodiment, the non-rectangular shape is selected from the group consisting of triangles, trapezoids, and polygons having at least five sides and angles. In yet another embodiment, the second partition has a non-rectangular shape.

[0029] In another embodiment, at least one of the first and second predictions is an interpretation process that predicts the plurality of first values ​​and the plurality of second values ​​based on a reference partition in an encoded reference picture. The interpretation process may predict a plurality of first values ​​for a plurality of pixels in the first partition which includes the set of pixels, or it may predict a plurality of second values ​​for only the set of pixels in the first partition.

[0030] In another embodiment, at least one of the first and second predictions is an intra-prediction process that predicts the plurality of first and plurality of second values ​​based on encoded reference partitions in the current picture.

[0031] In another embodiment, an image decoding device is provided, comprising: an entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantization conversion coefficients; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantization of the plurality of quantization conversion coefficients to obtain a plurality of conversion coefficients and inverse transform the plurality of conversion coefficients to obtain a plurality of residuals; an addition unit that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; an inter-prediction unit that, in operation, generates a prediction of the current block based on a reference block in a decoded reference picture; an intra-prediction unit that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and the prediction control unit connected to memory. The prediction control unit, in operation, performs boundary smoothing operations along the boundary between a first partition having a non-rectangular shape and a second partition, which are separated from an image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0032] In another embodiment, an image decoding method is provided that performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition. The method typically includes four steps: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0033] According to one embodiment, an image encoding device is provided comprising a circuit and a memory connected to the circuit. In operation, the circuit performs a partition syntax operation, which includes dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on partition parameters indicating the division; encoding the first partition and the second partition; and writing one or more parameters, including the partition parameters, to a bitstream.

[0034] In a further embodiment, the partition parameter indicates that the first partition has a triangular shape.

[0035] In another embodiment, the partition parameter indicates that the second partition has a non-rectangular shape.

[0036] In another embodiment, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0037] In another embodiment, the partition parameter collectively encodes the partitioning direction applied to divide the image block into the plurality of partitions. For example, the partitioning direction may include the upper left corner to the lower right corner of the image block, and the upper right corner to the lower left corner of the image block. The partition parameter collectively encodes at least the first motion vector of the first partition.

[0038] In another embodiment, one or more parameters other than the partition parameter encode the division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode at least the first motion vector of the first partition together with the other parameters.

[0039] In another embodiment, the partition parameter may encode at least the first motion vector of the first partition together with the partition parameter. The partition parameter may encode the second motion vector of the second partition together with the partition parameter.

[0040] In another embodiment, one or more parameters other than the partition parameters may encode at least the first motion vector of the first partition.

[0041] In another embodiment, the one or more parameters are binarized according to a binarization scheme selected according to the value of at least one of the one or more parameters.

[0042] In a further embodiment, an image encoding device is provided, comprising: a division unit that, in operation, receives a source image and divides it into a plurality of blocks; an addition unit that, in operation, receives the plurality of blocks from the division unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a conversion unit that, in operation, performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients; a quantization unit that, in operation, quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients; an entropy encoding unit that, in operation, encodes the plurality of quantized conversion coefficients to generate a bitstream; an inter-prediction unit that, in operation, generates a prediction for the current block based on a reference block in a pre-encoded reference picture; an intra-prediction unit that, in operation, generates a prediction for the current block based on a pre-encoded reference block in the current picture; and the prediction control unit connected to memory. The prediction control unit, in operation, divides an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on a partition parameter indicating a division, and encodes the first partition and the second partition. The entropy encoding unit writes one or more parameters, including the partition parameter, to the bitstream during its operation.

[0043] In another embodiment, an image encoding method is provided which includes partition syntax operation. The method typically includes three steps: dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on partition parameters indicating a division; encoding the first partition and the second partition; and writing one or more parameters, including the partition parameters, to a bitstream.

[0044] In another embodiment, an image decoding device is provided comprising a circuit and a memory connected to the circuit. The circuit, in operation, performs a partition syntax operation, which includes reading one or more parameters from a bitstream, including partition parameters indicating that an image block is to be divided into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape; dividing the image block into the plurality of partitions based on the partition parameters; and decoding the first partition and the second partition.

[0045] In a further embodiment, the partition parameter indicates that the first partition has a triangular shape.

[0046] In another embodiment, the partition parameter indicates that the second partition has a non-rectangular shape.

[0047] In another embodiment, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.

[0048] In another embodiment, the partition parameter collectively encodes the partitioning direction applied to divide the image block into the plurality of partitions. For example, the partitioning direction includes from the upper-left corner to the lower-right corner of the image block, and from the upper-right corner to the lower-left corner of the image block. The partition parameter may also collectively encode at least the first motion vector of the first partition.

[0049] In another embodiment, one or more parameters other than the partition parameter encode the division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode at least the first motion vector of the first partition together with the other parameters.

[0050] In another embodiment, the partition parameter may encode at least the first motion vector of the first partition together with the partition parameter. The partition parameter may encode the second motion vector of the second partition together with the partition parameter.

[0051] In another embodiment, one or more parameters other than the partition parameters may encode at least the first motion vector of the first partition.

[0052] In another embodiment, the one or more parameters are binarized according to a binarization scheme selected according to the value of at least one of the one or more parameters.

[0053] In a further embodiment, an image decoding device is provided, comprising: an entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantization conversion coefficients; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantization of the plurality of quantization conversion coefficients to obtain a plurality of conversion coefficients and inverse transform the plurality of conversion coefficients to obtain a plurality of residuals; an addition unit that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; an inter-prediction unit that, in operation, generates a prediction of the current block based on a reference block in a decoded reference picture; an intra-prediction unit that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and the prediction control unit connected to memory. The entropy decoding unit reads one or more parameters from the bitstream, including a partition parameter indicating that the image block is divided into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape, during operation. Based on the partition parameter, the unit divides the image block into the plurality of partitions and decodes the first partition and the second partition.

[0054] In another embodiment, an image decoding method is provided which includes partition syntax operation. The method typically includes three steps: reading one or more parameters from a bitstream, including partition parameters indicating that an image block is to be divided into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape; dividing the image block into the plurality of partitions based on the partition parameters; and decoding the first partition and the second partition.

[0055] In drawings, the same reference numeral indicates the same component. The size and relative position of components in drawings do not necessarily follow the scale ratio.

[0056] The embodiments will be described below in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, relationships and sequences of steps, etc., shown in the following embodiments are examples and are not intended to limit the scope of the claims. Therefore, components disclosed in the following embodiments but not described in the independent claims defining the broadest inventive concept may be understood as any component.

[0057] Embodiments of encoding and decoding devices are described below. These embodiments are examples of encoding and decoding devices to which the processes and / or configurations described in each aspect of this disclosure can be applied. The processes and / or configurations can also be implemented in encoding and decoding devices different from those in the embodiments. For example, with respect to the processes and / or configurations applicable to the embodiments, one of the following may be implemented:

[0058] (1) Any of the multiple components of the encoding or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0059] (2) In the encoding or decoding device of the embodiment, any modifications such as addition, replacement, or deletion of functions or processes performed by some of the multiple components of the encoding or decoding device may be made. For example, any of the functions or processes may be replaced or combined with other functions or processes described in any of the embodiments of this disclosure.

[0060] (3) In the methods performed by the encoding or decoding apparatus of the embodiment, any modifications, such as additions, replacements, and deletions, may be made to some of the processes included in the method. For example, any of the processes in the method may be replaced with or combined with other processes described in any of the embodiments of this disclosure.

[0061] (4) Some of the multiple components constituting the encoding or decoding device of the embodiment may be combined with components described in any of the embodiments of this disclosure, or with components that have some of the functions described in any of the embodiments of this disclosure, or with components that perform some of the processing performed by the components described in any of the embodiments of this disclosure.

[0062] (5) Components that provide some of the functions of the encoding or decoding device of the embodiment, or components that perform some of the processing of the encoding or decoding device of the embodiment, may be combined with or replaced with components described in any of the aspects of the disclosure and components that provide some of the functions described in any of the aspects of the disclosure, or components that perform some of the processing described in any of the aspects of the disclosure.

[0063] (6) In a method performed by an encoding or decoding device of an embodiment, any of the processes included in the method may be replaced or combined with any of the processes described in any of the embodiments of the present disclosure.

[0064] (7) Some of the processes included in the methods performed by the encoding or decoding device of the embodiment may be combined with the processes described in any of the embodiments of this disclosure.

[0065] (8) The methods of carrying out the processes and / or configurations described in each aspect of the present disclosure are not limited to the encoding or decoding devices of the embodiments. For example, the processes and / or configurations may be carried out in devices used for purposes other than the video encoding or video decoding disclosed in the embodiments.

[0066] [Overview of the coding device] First, an overview of the encoding device according to the embodiment will be described. Figure 1 is a block diagram showing the functional configuration of the encoding device 100 according to the embodiment. The encoding device 100 is a video encoding device that encodes video in block units.

[0067] As shown in Figure 1, the encoding device 100 is a device that encodes an image in block units and comprises a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0068] The encoding device 100 can be implemented, for example, by a general-purpose processor and memory. In this case, when a software program stored in memory is executed by the processor, the processor functions as a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. Alternatively, the encoding device 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0069] The following describes each component included in the encoding device 100.

[0070] [Divided part] The splitting unit 102 divides each picture contained in the input video into multiple blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first divides the picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks are sometimes called coding tree units (CTUs). Then, based on recursive quadtree and / or binary tree block partitioning, the splitting unit 102 divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less). These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs). In various processing examples, CUs, PUs, and TUs do not need to be distinguished, and some or all of the blocks in the picture may be processing units of CUs, PUs, or TUs.

[0071] Figure 2 shows an example of block partitioning in the embodiment. In Figure 2, solid lines represent block boundaries due to quadtree block partitioning, and dashed lines represent block boundaries due to binary tree block partitioning.

[0072] Here, block 10 is a 128x128 pixel square block (128x128 block). This 128x128 block 10 is first divided into four 64x64 square blocks (quadtree block partitioning).

[0073] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block partitioning). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0074] The 64x64 block in the upper right is horizontally divided into two rectangular 64x32 blocks, 14 and 15 (binary tree block division).

[0075] The bottom-left 64x64 block is divided into four square 32x32 blocks (quadrutree block division). Of the four 32x32 blocks, the top-left and bottom-right blocks are further divided. The top-left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the rightmost 16x32 block is further horizontally divided into two 16x16 blocks (binary tree block division). The bottom-right 32x32 block is horizontally divided into two 32x16 blocks (binary tree block division). As a result, the bottom-left 64x64 block is divided into 16x32 block 16, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.

[0076] The 64x64 block 23 in the bottom right will not be divided.

[0077] As described above, in Figure 2, block 10 is divided into 13 variable-sized blocks 11-23 based on recursive quad-tree and binary tree block partitioning. Such partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.

[0078] In Figure 2, one block was divided into four or two blocks (quadrutree or binary tree block partitioning), but partitioning is not limited to these. For example, one block may be divided into three blocks (ternary tree block partitioning). Partitioning that includes such ternary tree block partitioning is sometimes called MBT (multi-type tree) partitioning.

[0079] [Subtraction Unit] The subtraction unit 104 subtracts the predicted signal (predicted samples input from the prediction control unit 128, shown below) from the original signal (original sample) in block units that are input from the division unit 102 and divided by the division unit 102. In other words, the subtraction unit 104 calculates the prediction error (also called residual) of the block to be encoded (hereinafter referred to as the current block). The subtraction unit 104 then outputs the calculated prediction error (residual) to the conversion unit 106.

[0080] The source signal is the input signal to the encoding device 100, and is a signal representing the image of each picture that makes up the moving image (for example, a luminance (luma) signal and two chroma (chroma) signals). In the following, the signal representing the image may also be called a sample.

[0081] [Conversion section] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example.

[0082] The transformation unit 106 may also adaptively select a transformation type from among several transformation types and use a transformation basis function corresponding to the selected transformation type to convert the prediction error into transformation coefficients. Such a transformation is sometimes called an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform).

[0083] Multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 3 is a table showing the transformation basis functions corresponding to each transformation type. In Figure 3, N represents the number of input pixels. The selection of a transformation type from among these multiple transformation types may depend, for example, on the type of prediction (intra-prediction and inter-prediction) or on the intra-prediction mode.

[0084] Information indicating whether or not to apply EMT or AMT (e.g., called an EMT flag or AMT flag) and information indicating the selected conversion type are typically signaled at the CU level. However, the signaling of this information is not limited to the CU level and may be at other levels (e.g., bit sequence level, picture level, slice level, tile level, or CTU level).

[0085] Furthermore, the transformation unit 106 may retransform the transformation coefficients (transformation results). Such retransformation is sometimes called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the transformation unit 106 performs retransformation for each subblock (e.g., 4x4 subblock) contained in the block of transformation coefficients corresponding to the intra-prediction error. Information indicating whether or not to apply NSST and information regarding the transformation matrix used for NSST are usually signaled at the CU level. However, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0086] The transformation unit 106 may be subjected to either a separable transformation or a non-separable transformation. A separable transformation is a method in which the input is separated into directions equal to the number of dimensions and transformed multiple times, while a non-separable transformation is a method in which, when the input is multidimensional, two or more dimensions are treated as one dimension and transformed together.

[0087] For example, one example of a non-separable transformation is to treat a 4x4 block as a single array with 16 elements and then perform the transformation on that array using a 16x16 transformation matrix.

[0088] Another example of a non-separable transformation is one in which a 4x4 input block is treated as a single array with 16 elements, and then multiple Givens rotations are performed on that array (for example, the Hypercube Givens Transform).

[0089] [Quantization section] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order and quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. The quantization unit 108 then outputs the quantized conversion coefficients of the current block (hereinafter referred to as quantization coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0090] A predetermined scan order is the order for quantization / inverse quantization of the transformation coefficients. For example, a predetermined scan order is defined as ascending frequency (from low to high frequency) or descending frequency (from high to low frequency).

[0091] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0092] [Entropy coding unit] The entropy coding unit 110 generates an encoded signal (encoded bitstream) based on the quantization coefficients input from the quantization unit 108. Specifically, for example, the entropy coding unit 110 binarizes the quantization coefficients, arithmetically encodes the binary signal, and outputs a compressed bitstream or sequence.

[0093] [Dequantization section] The inverse quantization unit 112 inversely quantizes the quantization coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inversely quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inversely quantized conversion coefficients of the current block to the inverse conversion unit 114.

[0094] [Inverse Transformation Section] The inverse transform unit 114 restores the prediction error (residual) by performing an inverse transform on the transformation coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform on the transformation coefficients corresponding to the transformation by the transformation unit 106. The inverse transform unit 114 then outputs the restored prediction error to the adder unit 116.

[0095] Furthermore, the recovered prediction error usually does not match the prediction error calculated by the subtraction unit 104 because information is typically lost due to quantization. In other words, the recovered prediction error usually includes quantization errors.

[0096] [Addition section] The adder 116 reconstructs the current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called the local decoded block.

[0097] [Block memory] The block memory 118 is a storage unit for storing blocks within the picture to be encoded ("current picture") that are referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0098] [Loop Filter Section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder unit 116 and outputs the filtered reconstructed block to the frame memory 122. A loop filter is a filter used within the encoding loop (in-loop filter), and includes, for example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF).

[0099] In ALF, a least-squares error filter is applied to remove coding distortion. For example, for each 2x2 subblock within the current block, one filter selected from several filters is applied based on the direction and activity of the local gradient.

[0100] Specifically, first, subblocks (e.g., 2x2 subblocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of subblocks is based on the direction and activity of the gradient. For example, a classification value C (e.g., C = 5D + A) is calculated using the gradient direction value D (e.g., 0-2 or 0-4) and the gradient activity value A (e.g., 0-4). Then, based on the classification value C, the subblocks are classified into multiple classes.

[0101] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). The gradient activation value A is derived, for example, by adding the gradients in multiple directions and quantizing the sum.

[0102] Based on the results of this classification, a filter for the subblock is determined from among multiple filters.

[0103] For example, a circularly symmetric shape is used as the filter shape in ALF. Figures 4A to 4C show several examples of filter shapes used in ALF. Figure 4A shows a 5x5 diamond-shaped filter, Figure 4B shows a 7x7 diamond-shaped filter, and Figure 4C shows a 9x9 diamond-shaped filter. Information indicating the filter shape is usually signaled at the picture level. However, the signaling of information indicating the filter shape is not limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0104] The ALF (Automatic Laser Level) on / off status may be determined at the picture level or the CU (Camera Unit) level. For example, the decision to apply ALF to luminance may be made at the CU level, while the decision to apply ALF to color difference may be made at the picture level. Information indicating whether ALF is on or off is usually signaled at the picture level or the CU level. However, the signaling of information indicating whether ALF is on or off is not limited to the picture level or the CU level, but may be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0105] The coefficient sets for multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level. However, signaling of the coefficient sets is not limited to the picture level; it may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or subblock level).

[0106] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used for interpretation, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0107] [Intra Prediction Unit] The intra-prediction unit 124 generates a prediction signal (intra-prediction signal) by performing intra-prediction (also called in-screen prediction) of the current block by referring to blocks in the current picture, such as those stored in the block memory 118. Specifically, the intra-prediction unit 124 generates an intra-prediction signal by performing intra-prediction by referring to samples (e.g., luminance values, color difference values) of blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 128.

[0108] For example, the intra-prediction unit 124 performs intra-prediction using one of a predetermined set of intra-prediction modes. The set of intra-prediction modes typically includes one or more non-directional prediction modes and multiple directional prediction modes.

[0109] One or more non-directional prediction modes include, for example, the Planar prediction mode and DC prediction mode as defined in the H.265 / HEVC standard.

[0110] Multiple directional prediction modes include, for example, the 33 directional prediction modes defined in the H.265 / HEVC standard. Alternatively, multiple directional prediction modes may include an additional 32 directional prediction modes (a total of 65 directional prediction modes).

[0111] Figure 5A is a conceptual diagram showing all 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra-prediction. Solid arrows represent the 33 directions specified in the H.265 / HEVC standard, and dashed arrows represent the additional 32 directions (the two "non-directional" prediction modes are not shown in Figure 5A).

[0112] In various processing examples, the luminance block may be referenced in the intra-prediction of the chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. This intra-prediction is sometimes called CCLM (cross-component linear model) prediction. Such an intra-prediction mode for a chrominance block that references a luminance block (e.g., called the CCLM mode) may be added as one of the intra-prediction modes for a chrominance block.

[0113] The intra-prediction unit 124 may correct the pixel values ​​after intra-prediction based on the gradient of the horizontal / vertical reference pixels. Intra-prediction with such correction is sometimes called PDPC (position dependent intra-prediction combination). Information indicating whether or not PDPC is applied (for example, called a PDPC flag) is usually signaled at the CU level. However, the signaling of this information is not limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0114] [International Prediction Department] The inter-prediction unit 126 generates a prediction signal (inter-prediction signal) by performing inter-prediction (also called inter-screen prediction) of the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture. Inter-prediction is performed in units of the current block or the current sub-block within the current block (e.g., a 4x4 block). For example, the inter-prediction unit 126 performs motion estimation within the reference picture for the current block or current sub-block and finds the reference block or sub-block within the reference picture that best matches that current block or sub-block. Then, the inter-prediction unit 126 acquires motion information (e.g., a motion vector) that compensates for (or predicts) the movement or change from the reference block or sub-block to the current block or sub-block. Then, based on that motion information, the inter-prediction unit 126 performs motion compensation (or motion prediction) and generates an inter-prediction signal for the current block or sub-block. Finally, the inter-prediction unit 126 outputs the generated inter-prediction signal to the prediction control unit 128.

[0115] The motion information used for motion compensation may be signaled as an interprediction signal in various forms. For example, the motion vector may be signaled. As another example, the difference between the motion vector and the predicted motion vector may be signaled.

[0116] Furthermore, an inter-prediction signal may be generated using not only the motion information of the current block obtained through motion search, but also the motion information of adjacent blocks. Specifically, an inter-prediction signal may be generated for each sub-block within the current block by weighting and adding together a prediction signal based on motion information obtained through motion search (in the reference picture) and a prediction signal based on motion information of adjacent blocks (in the current picture). Such inter-prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).

[0117] In OBMC mode, information indicating the size of subblocks for OBMC (e.g., called OBMC block size) may be signaled at the sequence level. Information indicating whether or not OBMC mode is applied (e.g., called OBMC flag) may also be signaled at the CU level. Note that the signaling levels for this information are not limited to sequence and CU levels, but may be other levels (e.g., picture level, slice level, tile level, CTU level, or subblock level).

[0118] Let's explain the OBMC mode in more detail. Figures 5B and 5C are flowcharts and conceptual diagrams illustrating the predictive image correction process using OBMC processing.

[0119] Referring to Figure 5C, first, a predicted image (Pred) is obtained using normal motion compensation with the motion vector (MV) assigned to the current block to be encoded. In Figure 5C, the arrow "MV" points to the reference picture, indicating what the current block within the current picture is referencing in order to obtain the predicted image.

[0120] Next, the motion vector (MV_L) already derived for the encoded left-adjacent block is applied (reused) to the current block to be encoded to obtain the predicted image (Pred_L). The motion vector (MV_L) is indicated by an arrow "MV_L" pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by superimposing the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0121] Similarly, the motion vector (MV_U) already derived for the encoded upper adjacent block is applied (reused) to the current block to be encoded to obtain the predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. Then, the predicted image Pred_U is superimposed on the predicted image that has undergone the first correction (i.e., Pred and Pred_L) to perform a second correction of the predicted image. In one embodiment, this has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block, with the boundaries with adjacent blocks blended (smoothed).

[0122] While this explanation describes a two-stage correction method using left-adjacent blocks and upper-adjacent blocks, it is also possible to use right-adjacent blocks and lower-adjacent blocks to perform corrections more than two times.

[0123] Furthermore, the area to be superimposed does not have to be the entire pixel area of ​​the block, but rather only a portion of the area near the block boundary.

[0124] Here, we have described the OBMC predictive image correction process for obtaining a single predictive image Pred by superimposing additional predictive images Pred_L and Pred_U based on a single reference picture. However, if the predictive image is corrected based on multiple reference pictures, the same process may be applied to each of the multiple reference pictures. In such cases, by performing OBMC image correction based on multiple reference pictures, corrected predictive images are obtained from each reference picture, and then these multiple corrected predictive images are further superimposed to obtain the final predictive image.

[0125] In OBMC, the unit of the target block may be the prediction block unit, or it may be a subblock unit obtained by further dividing the prediction block.

[0126] One method for determining whether or not to apply OBMC processing is to use an obmc_flag signal that indicates whether or not to apply OBMC processing. Specifically, in an encoding device, it is determined whether or not the block to be encoded belongs to a region with complex motion. If it belongs to a region with complex motion, the value of obmc_flag is set to "1" and OBMC processing is applied to perform encoding. If it does not belong to a region with complex motion, the value of obmc_flag is set to "0" and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, the obmc_flag written in the stream (i.e., the compressed sequence) is decoded, and the device switches whether or not to apply OBMC processing depending on its value and performs decoding.

[0127] Furthermore, motion information may be derived at the decoding device side without being signaled at the encoding device side. For example, the merge mode specified in the H.265 / HEVC standard may be used. Alternatively, motion information may be derived by performing a motion search at the decoding device side. In this case, the decoding device side may perform the motion search without using the pixel values ​​of the current block.

[0128] Here, we will explain the mode in which motion detection is performed on the decoding device side. This mode in which motion detection is performed on the decoding device side is sometimes called PMMVD (pattern matched motion vector derivation) mode or FRUC (frame rate up-conversion) mode.

[0129] An example of FRUC processing is shown in Figure 5D. First, a list of multiple candidates (which may be the same as the merge list) is generated by referencing the motion vectors of encoded blocks spatially or temporally adjacent to the current block, each having a predicted motion vector (MV). Next, the best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value is calculated for each candidate MV included in the candidate list, and one candidate MV is selected based on the evaluation value.

[0130] Then, based on the motion vector of the selected candidate, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Alternatively, for example, the motion vector for the current block may be derived by performing pattern matching in the area surrounding the position in the reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed in the area surrounding the best candidate MV using pattern matching in the reference picture and evaluation values, and if an MV with a better evaluation value is found, the best candidate MV may be updated to that MV and made the final MV for the current block. It is also possible to configure the system so that it does not perform the process of updating to an MV with a better evaluation value.

[0131] The same processing method can be used when processing at the sub-block level.

[0132] The evaluation value may be calculated using various methods. For example, the reconstructed image of a region in a reference picture corresponding to the motion vector may be compared with the reconstructed image of a predetermined region (for example, a region in another reference picture, or a region in an adjacent block of the current picture, as described later), and the difference in pixel values ​​between the two reconstructed images may be calculated and used as the evaluation value for the motion vector. In addition to the difference value, other information may also be used to calculate the evaluation value.

[0133] Next, we will explain in detail an example of pattern matching. First, one candidate MV included in the candidate MV list (e.g., merge list) is selected as the starting point for the search using pattern matching. For example, first pattern matching or second pattern matching may be used. First pattern matching and second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0134] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are aligned with the motion trajectory of the current block. Therefore, in the first pattern matching, a region in the other reference picture that is aligned with the motion trajectory of the current block is used as a predetermined region for calculating the evaluation value of the candidate mentioned above, relative to the region in the reference picture.

[0135] Figure 6 illustrates an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in Figure 6, in first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the best-matching pair of blocks in two different reference pictures (Ref0, Ref1) that are along the motion trajectory of the current block. Specifically, for the current block, the difference between the reconstructed image at a specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference picture (Ref1) specified by a symmetric MV scaled by the display time interval of the candidate MV is derived, and an evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, which can yield good results.

[0136] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, then the first pattern matching derives two mirror-symmetric bidirectional motion vectors.

[0137] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (blocks adjacent to the current block in the current picture (e.g., blocks above and / or to the left)) and the blocks in the reference picture. Therefore, in the second pattern matching, the blocks adjacent to the current block in the current picture are used as a predetermined area for calculating the evaluation value of the candidates mentioned above.

[0138] Figure 7 illustrates an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in Figure 7, in the second pattern matching, the motion vector of the current block is derived by searching in the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or either of the left adjacent and upper adjacent regions is derived, and the reconstructed image of the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is calculated using the obtained difference value, and the candidate MV with the best evaluation value among multiple candidate MVs is selected as the best candidate MV.

[0139] Information indicating whether or not to apply such a FRUC mode (e.g., called a FRUC flag) may be signaled at the CU level. Furthermore, if FRUC mode is applied (e.g., the FRUC flag is true), information indicating the pattern matching method (e.g., first pattern matching or second pattern matching) (e.g., called a FRUC mode flag) may be signaled at the CU level. Note that the signaling of this information is not limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0140] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming a constant velocity linear motion will be described. This mode is sometimes called the BIO (bi - directional optical flow) mode.

[0141] FIG. 8 is a diagram for explaining a model assuming a constant velocity linear motion. In FIG. 8, (v x , v y ) indicates a velocity vector, τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the motion vector corresponding to the reference picture Ref1.

[0142] At this time, under the assumption of a constant velocity linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.

[0143]

Equation

[0144] Here, I (k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the merge list or the like may be corrected in pixel units.

[0145] Furthermore, motion vectors may be derived on the decoding side using a method different from that used for deriving motion vectors based on a model that assumes uniform linear motion. For example, motion vectors may be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0146] Next, we will describe a mode in which motion vectors are derived at the sub-block level based on the motion vectors of multiple adjacent blocks. This mode is sometimes called the affine motion compensation prediction mode.

[0147] Figure 9A is a diagram illustrating the derivation of subblock-level motion vectors based on the motion vectors of multiple adjacent blocks. In Figure 9A, the current block contains 16 4x4 subblocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent subblocks. Then, using the two motion vectors v0 and v1, the motion vector (v) of each subblock within the current block is derived by the following equation (2). x ,v y ) is derived.

[0148]

number

[0149] Here, x and y represent the horizontal and vertical positions of the subblock, respectively, and w represents a predetermined weighting coefficient.

[0150] This affine motion compensation prediction mode may include several modes that differ in the method of deriving the motion vectors of the upper-left and upper-right corner control points. Information indicating this affine motion compensation prediction mode (e.g., called an affine flag) may be signaled at the CU level. However, the signaling of information indicating this affine motion compensation prediction mode is not limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0151] [Prediction Control Unit] The prediction control unit 128 selects either the intra-prediction signal (a signal output from the intra-prediction unit 124) or the inter-prediction signal (a signal output from the inter-prediction unit 126), and outputs the selected signal as the prediction signal to the subtraction unit 104 and the addition unit 116.

[0152] As shown in Figure 1, in various processing examples, the prediction control unit 128 may output prediction parameters that are input to the entropy coding unit 110. The entropy coding unit 110 may generate a coded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by a decoder. The decoder may receive and decode the coded bitstream and perform the same processing as the prediction processing performed in the intra-prediction unit 124, inter-prediction unit 126, and prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, prediction type, or prediction mode used in the intra-prediction unit 124 or inter-prediction unit 126), or any index, flag, or value that is based on or indicates the prediction processing performed in the intra-prediction unit 124, inter-prediction unit 126, and prediction control unit 128.

[0153] Figure 9B shows an example of the process for deriving the motion vector of the current picture in merge mode.

[0154] First, a list of predicted MVs is generated, containing registered candidates for predicted MVs. Candidates for predicted MVs include spatially adjacent predicted MVs, which are the MVs of multiple encoded blocks located spatially around the target block; temporally adjacent predicted MVs, which are the MVs of nearby blocks projected onto the target block's position in the encoded reference picture; combined predicted MVs, which are generated by combining the MV values ​​of spatially adjacent predicted MVs and temporally adjacent predicted MVs; and zero predicted MVs, which are MVs with a value of zero.

[0155] Next, one predicted MV is selected from the multiple predicted MVs registered in the predicted MV list to determine it as the MV for the target block.

[0156] Furthermore, the variable-length coding unit encodes the merge_idx signal, which indicates which predicted MV was selected, by writing it to a stream.

[0157] Note that the predicted MVs registered in the predicted MV list explained in Figure 9B are just an example, and the number of predicted MVs may differ from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or it may include predicted MVs other than those shown in the figure.

[0158] The final MV may be determined by performing the DMVR (decoder motion vector refinement) process described later, using the MV of the target block derived by merge mode.

[0159] Figure 9C is a conceptual diagram illustrating an example of DMVR processing for determining MV.

[0160] First, the optimal MVP set for the current block (for example, in merge mode) is designated as the candidate MV. Then, according to the candidate MV(L0), reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 direction. Similarly, according to the candidate MV(L1), reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 direction. A template is generated by taking the average of these reference pixels.

[0161] Next, using the template, the surrounding regions of candidate MVs in the first reference picture (L0) and the second reference picture (L1) are searched, and the MV with the minimum cost is determined as the final MV. The cost value may be calculated, for example, using the difference between each pixel value of the template and each pixel value of the search region, as well as the candidate MV value.

[0162] Typically, the encoding device and the decoding device (described later) share the same basic configuration and operation for the processing described here.

[0163] Any process that can explore the vicinity of a candidate MV and derive the final MV is acceptable, even if it is not the exact process described here.

[0164] Next, we will describe an example of a mode that generates a predicted image (prediction) using LIC (local illumination compensation) processing.

[0165] Figure 9D is a conceptual diagram illustrating an example of a predictive image generation method using brightness correction processing by LIC processing.

[0166] First, the MV is derived from the encoded reference picture to obtain the reference image corresponding to the current block.

[0167] Next, information is extracted showing how the luminance values ​​have changed between the reference picture and the current picture for the current block. This extraction is based on the luminance pixel values ​​of the encoded left adjacent reference region (peripheral reference region) and encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the derived MV. Then, the luminance correction parameter is calculated using the information showing how the luminance values ​​have changed.

[0168] A predicted image for the current block is generated by applying the brightness correction parameters to the reference image within the reference picture specified in MV.

[0169] Note that the shape of the surrounding reference region in Figure 9D is just one example, and other shapes may be used.

[0170] Furthermore, although this explanation describes the process of generating a predicted image from a single reference picture, the process is similar when generating predicted images from multiple reference pictures. Alternatively, the brightness correction process may be applied to each reference picture obtained from the reference picture in the same manner as described above before generating the predicted image.

[0171] One method for determining whether or not to apply LIC processing is to use a signal called lic_flag, which indicates whether or not to apply LIC processing. For example, in an encoding device, it is determined whether the current block belongs to a region where brightness changes are occurring. If it belongs to a region where brightness changes are occurring, the value of lic_flag is set to "1" and LIC processing is applied and encoding is performed. If it does not belong to a region where brightness changes are occurring, the value of lic_flag is set to "0" and encoding is performed without applying LIC processing. On the other hand, in a decoding device, the lic_flag written in the stream may be decoded, and the device may switch whether or not to apply LIC processing depending on its value and perform decoding.

[0172] Another way to determine whether to apply LIC processing is, for example, by checking whether LIC processing was applied to surrounding blocks. A specific example is, if the current block is in merge mode, the system checks whether the surrounding encoded blocks selected during the MV derivation in merge mode were encoded with LIC processing. Based on this result, the system switches whether to apply LIC processing and then performs the encoding. Note that in this example, the same process is applied to the decoding device.

[0173] [Overview of the decryption device] Next, an overview of a decoding device capable of decoding the encoded signal (encoded bitstream) output from the above-mentioned encoding device 100 will be described. Figure 10 is a block diagram showing the functional configuration of the decoding device 200 according to the embodiment. The decoding device 200 is a video decoding device that decodes video in block units.

[0174] As shown in Figure 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0175] The decoding device 200 can be implemented, for example, by a general-purpose processor and memory. In this case, when the software program stored in memory is executed by the processor, the processor functions as an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a loop filter unit 212, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220. Alternatively, the decoding device 200 may be implemented as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transformation unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0176] The following describes each component included in the decoding device 200.

[0177] [Entropy Decoder] The entropy decoding unit 202 entropically decodes the encoded bitstream. Specifically, the entropy decoding unit 202 arithmetically decodes the encoded bitstream into a binary signal, for example. Then, the entropy decoding unit 202 debinarizes the binary signal. The entropy decoding unit 202 outputs the quantization coefficients in block units to the inverse quantization unit 204. The entropy decoding unit 202 may also output prediction parameters included in the encoded bitstream (see Figure 1) to the intra-prediction unit 216, inter-prediction unit 218, and prediction control unit 220 in the embodiment. The intra-prediction unit 216, inter-prediction unit 218, and prediction control unit 220 can perform the same prediction processing as the intra-prediction unit 124, inter-prediction unit 126, and prediction control unit 128 on the encoding device side.

[0178] [Dequantization section] The inverse quantization unit 204 inversely quantizes the quantization coefficients of the block to be decoded (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, for each quantization coefficient of the current block, the inverse quantization unit 204 inversely quantizes the quantization coefficient based on the quantization parameter corresponding to that quantization coefficient. The inverse quantization unit 204 then outputs the inversely quantized quantization coefficients (i.e., transformation coefficients) of the current block to the inverse transformation unit 206.

[0179] [Inverse Transformation Section] The inverse transform unit 206 recovers the prediction error (residual) by inversely transforming the transformation coefficients input from the inverse quantization unit 204.

[0180] For example, if the information decoded from the encoded bitstream indicates that EMT or AMT should be applied (e.g., the AMT flag is true), the inverse transform unit 206 inversely transforms the transformation coefficients of the current block based on the information indicating the decoded transformation type.

[0181] For example, if the information decoded from the encoded bitstream indicates that NSST should be applied, the inverse transform unit 206 applies inverse retransformation to the transformation coefficients.

[0182] [Addition section] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transformer 206 and the prediction sample input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0183] [Block memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are located within the decoded picture (hereinafter referred to as the current picture). Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208.

[0184] [Loop Filter Section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208 and outputs the filtered reconstructed block to the frame memory 214 and the display device, etc.

[0185] If the information interpreted from the encoded bitstream indicating ALF on / off indicates ALF is on, one filter is selected from among several filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstruction block.

[0186] [Frame memory] The frame memory 214 is a memory unit for storing reference pictures used for interpretation, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0187] [Intra Prediction Unit] The intra-prediction unit 216 generates a prediction signal (intra-prediction signal) by performing intra-prediction based on the intra-prediction mode decoded from the encoded bitstream, and by referring to the blocks in the current picture stored in the block memory 210. Specifically, the intra-prediction unit 216 generates an intra-prediction signal by performing intra-prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra-prediction signal to the prediction control unit 220.

[0188] Furthermore, if an intra-prediction mode that references a luminance block is selected in the intra-prediction of a color difference block, the intra-prediction unit 216 may predict the color difference component of the current block based on the luminance component of the current block.

[0189] Furthermore, if the information decoded from the encoded bitstream (for example, prediction parameters output from the entropy decoding unit 202) indicates the application of PDPC, the intra-prediction unit 216 corrects the pixel value after intra-prediction based on the gradient of the reference pixels in the horizontal / vertical directions.

[0190] [International Prediction Department] The inter-prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. Prediction is performed in units of the current block or sub-blocks within the current block (e.g., 4x4 blocks). For example, the inter-prediction unit 218 generates an inter-prediction signal for the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) decoded from the encoded bitstream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter-prediction signal to the prediction control unit 220.

[0191] If the information decoded from the encoded bitstream indicates that the OBMC mode should be applied, the interpretation unit 218 generates an interpretation prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent block.

[0192] Furthermore, if the information decoded from the encoded bitstream indicates that FRUC mode should be applied, the interpretation unit 218 derives motion information by performing a motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the interpretation unit 218 performs motion compensation (prediction) using the derived motion information.

[0193] Furthermore, when the BIO mode is applied, the inter-prediction unit 218 derives motion vectors based on a model that assumes uniform linear motion. Also, if the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode should be applied, the inter-prediction unit 218 derives motion vectors on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0194] [Prediction Control Unit] The prediction control unit 220 selects either the intra-prediction signal or the inter-prediction signal and outputs the selected signal as the prediction signal to the summer unit 208. Overall, the configuration, functions, and processing of the prediction control unit 220, intra-prediction unit 216, and inter-prediction unit 218 on the decoding device side may correspond to the configuration, functions, and processing of the prediction control unit 128, intra-prediction unit 124, and inter-prediction unit 126 on the encoding device side.

[0195] [Non-rectangular division] In the prediction control unit 128 connected to the intra-prediction unit 124 and inter-prediction unit 126 of the encoding device (see Figure 1), and in the prediction control unit 220 connected to the intra-prediction unit 216 and inter-prediction unit 218 of the decoding device (see Figure 10), conventionally, the multiple partitions (or multiple variable-size blocks, or multiple sub-blocks) obtained from the division of each block and from which motion information (e.g., multiple motion vectors) is obtained are always rectangular, as shown in Figure 2. The inventors have found that generating multiple partitions having a non-rectangular shape, such as a triangular shape, leads to improvements in image quality and encoding efficiency in various embodiments, depending on the image content in the picture. Various embodiments in which at least one partition divided from an image block for prediction purposes has a non-rectangular shape are described below. These embodiments are equally applicable to the encoding device side (prediction control unit 128 connected to the intra-prediction unit 124 and the inter-prediction unit 126) and the decoding device side (prediction control unit 220 connected to the intra-prediction unit 216 and the inter-prediction unit 218), and may be implemented in the encoding device shown in Figure 1, or the decoding device shown in Figure 10.

[0196] Figure 11 is a flowchart illustrating an example of a process that divides an image block into multiple partitions, including at least a first partition and a second partition having a non-rectangular shape (e.g., a triangle), and then encodes (or decodes) the image block as a reconstructed combination of the first and second partitions.

[0197] In step S1001, the image block is divided into multiple partitions, including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. For example, as shown in Figure 12, the image block may be divided from the upper left corner to the lower right corner to create a first and second partition that both have a non-rectangular shape (e.g., a triangle). Alternatively, the image block may be divided from the upper right corner to the lower left corner to create a first and second partition that both have a non-rectangular shape (e.g., a triangle). Various examples of non-rectangular divisions will be described later with reference to Figures 12 and 17-19.

[0198] In step S1002, the process predicts a first motion vector for the first partition and a second motion vector for the second partition. For example, predicting the first and second motion vectors may include selecting a first motion vector from a set of candidate first motion vectors and selecting a second motion vector from a set of candidate second motion vectors.

[0199] In step S1003, motion compensation processing is performed to obtain the first partition using the first motion vector derived in step S1002, and to obtain the second partition using the second motion vector derived in step S1002.

[0200] In step S1004, prediction processing is performed on the image block as a (reconstructed) combination of the first partition and the second partition. The prediction processing includes boundary smoothing processing to smooth the boundary between the first partition and the second partition. For example, boundary smoothing processing involves weighting multiple first values ​​of multiple boundary pixels predicted based on the first partition and multiple second values ​​of multiple boundary pixels predicted based on the second partition. Various implementations of boundary smoothing processing will be described later with reference to Figures 13, 14, 20 and 21A-21D.

[0201] In step S1005, the process encodes or decodes an image block using one or more parameters, including a partition parameter indicating that the image block is divided into a first partition and a second partition having a non-rectangular shape. As summarized in the table in Figure 15, for example, the partition parameter ("first index value") may encode together the division direction applied to the division (for example, from top left to bottom right or top right to bottom left, as shown in Figure 12) and the first and second motion vectors derived in step S1002 described above. Details of such partition syntax operation with one or more parameters, including the partition parameter, will be described in detail later with reference to Figures 15, 16 and 22-25.

[0202] Figure 17 is a flowchart of the process 2000 for dividing an image block. In step S2001, the process divides the image into multiple partitions, including a first partition having a non-rectangular shape and a second partition which may or may not have a non-rectangular shape. As shown in Figure 12, the image block is divided into a first partition having a triangular shape and a second partition which also has a triangular shape. There are many other examples in which the image block is divided into multiple partitions, including a first partition and a second partition, where at least the first partition has a non-rectangular shape. The non-rectangular shape may be a triangle, a trapezoid, or a polygon having at least five sides and angles.

[0203] For example, as shown in Figure 18, an image block may be divided into two triangular partitions. An image block may be divided into more than two triangular partitions (e.g., three triangular partitions). An image block may be divided into a combination of one or more triangular partitions and one or more rectangular partitions. Alternatively, an image block may be divided into a combination of one or more triangular partitions and one or more polygonal partitions.

[0204] Furthermore, as shown in Figure 19, the image block may be divided into L-shaped (polygonal) partitions and rectangular partitions. The image block may be divided into pentagonal (polygonal) partitions and triangular partitions. The image block may be divided into hexagonal (polygonal) partitions and pentagonal (polygonal) partitions. Alternatively, the image block may be divided into multiple polygonal partitions.

[0205] Referring again to Figure 17, in step S2002, the process predicts a first motion vector for the first partition by, for example, selecting the first partition from a first motion vector candidate set, and predicts a second motion vector for the second partition by, for example, selecting the second partition from a second motion vector candidate set. For example, the first motion vector candidate set may include multiple motion vectors for multiple partitions adjacent to the first partition, and the second motion vector candidate set may include multiple motion vectors for multiple partitions adjacent to the second partition. The adjacent partitions may be either or both spatially adjacent partitions and / or temporally adjacent partitions. Some examples of spatially adjacent partitions include partitions located to the left, lower left, bottom, lower right, right, upper right, top, or upper left of the partition being processed. Some examples of temporally adjacent partitions include multiple co-located partitions in multiple reference pictures of an image block.

[0206] In various implementations, multiple partitions adjacent to the first partition, and multiple partitions adjacent to the second partition, may be outside the image block divided into the first and second partitions. The first motion vector candidate set may be the same as or different from the second motion vector candidate set. Furthermore, the first motion vector candidate set and at least one of the second motion vector candidate sets may be the same as another third motion vector candidate set prepared for the image block.

[0207] In some implementations, in step S2002, if it is determined that the second partition, like the first partition, has a non-rectangular shape (e.g., a triangle), process 2000 creates a second motion vector candidate set (for the non-rectangular second partition) that includes multiple motion vectors of multiple partitions adjacent to the second partition, excluding the first partition (i.e., excluding the motion vector of the first partition). On the other hand, if it is determined that the second partition has a rectangular shape, unlike the first partition, process 2000 creates a second motion vector candidate set (for the rectangular second partition) that includes the first partition and multiple motion vectors of multiple partitions adjacent to the second partition.

[0208] In step S2003, the process encodes or decodes the first partition using the first motion vector derived in step S2002 described above, and encodes or decodes the second partition using the second motion vector derived in step S2002 described above.

[0209] Image block division processing, such as process 2000 in Figure 17, may be performed by an image encoding device including a circuit and memory connected to the circuit, as shown in Figure 1. In operation, the circuit divides an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape (step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0210] In another embodiment, an image encoding device, as shown in Figure 1, is provided, comprising: a splitting unit 102 that receives a source image in operation and divides it into multiple blocks; an adding unit 104 that receives multiple blocks from the splitting unit and multiple predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block and outputs a residual; a conversion unit 106 that performs a conversion on the multiple residuals output from the adder unit 104 and outputs multiple conversion coefficients; a quantization unit 108 that quantizes the multiple conversion coefficients in operation and generates multiple quantized conversion coefficients; an entropy encoding unit 110 that encodes the multiple quantized conversion coefficients in operation and generates a bitstream; an inter-prediction unit 126 that generates a prediction for the current block based on a reference block in an encoded reference picture in operation; an intra-prediction unit 124 that generates a prediction for the current block based on an encoded reference block in the current picture in operation; and a prediction control unit 128 connected to memories 118 and 122. In operation, the prediction control unit 128 divides the block into multiple partitions, including a first partition and a second partition, which have a non-rectangular shape (Figure 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0211] In other embodiments, an image decoding device is provided that includes a circuit and a memory connected to the circuit, as shown in Figure 10, for example. The circuit operates by dividing an image block into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape (Figure 17, step S2001), predicting a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decoding the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0212] Furthermore, according to the embodiment, the image decoding device shown in Figure 10 includes an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantization conversion coefficients, an inverse quantization unit 204 and an inverse conversion unit 206 that in operation inverse quantization conversion coefficients to obtain a plurality of conversion coefficients and inversely convert the plurality of conversion coefficients to obtain a plurality of residuals, an addition unit 208 that in operation adds a plurality of residuals output from the inverse quantization unit 204 and the inverse conversion unit 206 and a plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter-prediction unit 218 that in operation generates a prediction of the current block based on the reference block in the decoded reference picture, an intra-prediction unit 216 that in operation generates a prediction of the current block based on the decoded reference block in the current picture, and a prediction control unit 220 connected to memories 210 and 214. In operation, the prediction control unit 220 divides an image block into a plurality of partitions, including a first partition and a second partition, which have a non-rectangular shape (Figure 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0213] [Boundary smoothing] As described above in Figure 11, step S1004, predictive processing on an image block as a (reconstructed) combination of a first partition and a second partition having a non-rectangular shape may be accompanied by the application of boundary smoothing processing along the boundary between the first partition and the second partition, according to various embodiments.

[0214] For example, Figure 21B shows an example of boundary smoothing, which involves weighting multiple first values ​​of multiple boundary pixels predicted first based on the first partition with multiple second values ​​of multiple boundary pixels predicted second based on the second partition.

[0215] Figure 20 is a flowchart illustrating an overall boundary smoothing process 3000, according to one embodiment, which involves weighting a plurality of first values ​​of a plurality of boundary pixels first predicted based on a first partition with a plurality of second values ​​of a plurality of boundary pixels second predicted based on a second partition. In step S3001, as shown in Figure 21A, or in Figures 12, 18, and 19 described above, the image block is divided along its boundary into a first partition and a second partition, where at least the first partition has a non-rectangular shape.

[0216] In step S3002, multiple primary values ​​(e.g., color, brightness, transparency, etc.) of a set of pixels in the first partition along the boundary ("multiple boundary pixels" in Figure 21A) are first predicted using information from the first partition. In step S3003, multiple secondary values ​​of the (same) set of pixels in the first partition along the boundary are second predicted using information from the second partition. In some embodiments, at least one of the first and second predictions is an inter-prediction process that predicts multiple primary and secondary values ​​based on a reference partition in an encoded reference picture. Referring to Figure 21D, in some implementations, the prediction process predicts multiple primary values ​​for all pixels in the first partition ("first sample set") that includes the set of pixels where the first and second partitions overlap, and predicts secondary values ​​only for the set of pixels where the first and second partitions overlap ("second sample set"). In other implementations, at least one of the first and second predictions is an intra-prediction process that predicts multiple primary and secondary values ​​based on an encoded reference partition in the current picture. In some implementations, the prediction method used for the first prediction is different from the prediction method used for the second prediction. For example, the first prediction may include an intra-prediction process, and the second prediction may include an intra-prediction process. The information used for the first prediction of multiple primary values, or the second prediction of multiple secondary values, may include multiple motion vectors of the first or second partition, multiple intra-prediction directions, etc.

[0217] In step S3004, the multiple first values ​​predicted using the first partition and the multiple second values ​​predicted using the second partition are weighted. In step S3005, the first partition is encoded or decoded using the weighted multiple first values ​​and multiple second values.

[0218] Figure 21B shows an example of boundary smoothing operation where the first and second partitions overlap by up to 5 pixels in each row or row. That is, the maximum number of pixel sets in each row or column where multiple first values ​​are predicted based on the first partition and multiple second values ​​are predicted based on the second partition is 5. Figure 21C shows another example of boundary smoothing operation where the first and second partitions overlap by up to 3 pixels in each row or column. That is, the maximum number of pixel sets in each row or column where multiple first values ​​are predicted based on the first partition and multiple second values ​​are predicted based on the second partition is 3.

[0219] Figure 13 shows another example of boundary smoothing operation where the first and second partitions overlap by (up to) four pixels in each row or column. That is, there are at most four sets of pixels in each row or column where multiple primary values ​​are predicted based on the first partition and multiple secondary values ​​are predicted based on the second partition. In the example shown, multiple weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to multiple primary values ​​of four pixels in a set, respectively, and multiple weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to multiple secondary values ​​of four pixels in a set, respectively.

[0220] Figure 14 further shows several examples of boundary smoothing operations where the first and second partitions overlap by 0 pixels in each row or column (i.e., they do not overlap), overlap by (up to) 1 pixel in each row or column, and overlap by (up to) 2 pixels in each row or column. In the example where the first and second partitions do not overlap, multiple zero weights are applied. In the example where the first and second partitions overlap by 1 pixel in each row or column, a weight of 1 / 2 may be applied to multiple primary values ​​of multiple pixels in the set predicted based on the first partition, and a weight of 1 / 2 may be applied to multiple secondary values ​​of multiple pixels in the set predicted based on the second partition. In an example where the first and second partitions overlap by two pixels in each row or column, weights of 1 / 3 and 2 / 3 may be applied to multiple primary values ​​of two pixels in the set predicted based on the first partition, and weights of 2 / 3 and 1 / 3 may be applied to multiple secondary values ​​of two pixels in the set predicted based on the second partition.

[0221] In the embodiments described above, the number of pixels in the set where the first and second partitions overlap is an integer. In other implementations, the number of overlapping pixels in a set may be, for example, a non-integer or a fraction. The weights applied to the multiple first and multiple second values ​​of the pixel set may also be fractions or integers, depending on the application.

[0222] Boundary smoothing operations, such as process 3000 in Figure 20, may be performed by an image encoding device including a circuit and memory connected to the circuit, as shown in Figure 1. In operation, the circuit performs boundary smoothing operations along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition (Figure 20, step S3001). The boundary smoothing operations include predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0223] In another embodiment, an image encoding device, as shown in Figure 1, is provided, comprising: a splitting unit 102 that receives a source image in operation and divides it into multiple blocks; an adding unit 104 that receives multiple blocks from the splitting unit and multiple predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block and outputs a residual; a conversion unit 106 that performs a conversion on the multiple residuals output from the adder unit 104 and outputs multiple conversion coefficients; a quantization unit 108 that quantizes the multiple conversion coefficients in operation and generates multiple quantized conversion coefficients; an entropy encoding unit 110 that encodes the multiple quantized conversion coefficients in operation and generates a bitstream; an inter-prediction unit 126 that generates a prediction for the current block based on a reference block in an encoded reference picture in operation; an intra-prediction unit 124 that generates a prediction for the current block based on an encoded reference block in the current picture in operation; and a prediction control unit 128 connected to memories 118 and 122. In operation, the prediction control unit 128 performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition (Figure 20, step S3001). The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition (step S3002); second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition (step S3003); weighting the plurality of first values ​​and the plurality of second values ​​(step S3004); and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0224] In other embodiments, for example, an image decoding device is provided that includes a circuit and a memory connected to the circuit, as shown in Figure 10. The circuit performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape and a second partition, which are separated from an image block in operation (Figure 20, step S3001). The boundary smoothing operation includes first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0225] In another embodiment, the image decoding device shown in Figure 10 includes an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantization conversion coefficients, an inverse quantization unit 204 and an inverse transform unit 206 that in operation inverse quantization conversion coefficients to obtain a plurality of conversion coefficients and inverse transform the plurality of conversion coefficients to obtain a plurality of residuals, an addition unit 208 that in operation adds a plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and a plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter-prediction unit 218 that in operation generates a prediction of the current block based on a reference block in a decoded reference picture, an intra-prediction unit 216 that in operation generates a prediction of the current block based on a decoded reference block in the current picture, and a prediction control unit 220 connected to memories 210 and 214. In operation, the prediction control unit 220 performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition (Figure 20, step S3001). The boundary smoothing operation includes: first predicting a plurality of first values ​​of the pixel set of the first partition along the boundary using information of the first partition (step S3002); second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition (step S3003); weighting the plurality of first values ​​and the plurality of second values ​​(step S3004); and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0226] [Entropy coding and decoding using partition parameter syntax] As shown in Figure 11, step S1005, according to various embodiments, an image block divided into a first partition and a second partition having a non-rectangular shape may be encoded or decoded using one or more parameters including a partition parameter indicating the non-rectangular division of the image block. In various embodiments, such a partition parameter may be encoded together, for example, the division direction applied to the division (e.g., from top left to bottom right, or from top right to bottom left, see Figure 12) and the first and second motion vectors predicted in step S1002, as will be described in more detail later.

[0227] Figure 15 is a table of multiple information sets, each encoded as a whole by multiple sample partition parameters ("first index values") and multiple partition parameters. The multiple partition parameters ("first index values") range from 0 to 6 and encode as a whole the direction of dividing the image block into a first and second partition, both of which are triangles (see Figure 12), the predicted first motion vector for the first partition (Figure 11, step S1002), and the predicted second motion vector for the second partition (Figure 11, step S1002). In particular, partition parameter 0 encodes that the division direction is from the upper left corner to the lower right corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0228] Partition parameter 1 encodes that the partitioning direction is from the upper right corner to the lower left corner, that the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the partitioning direction is from the upper right corner to the lower left corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the partitioning direction is from the upper left corner to the lower right corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the division direction is from the upper right corner to the lower left corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the division direction is from the upper left corner to the lower right corner, that the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 6 encodes that the division direction is from the upper left corner to the lower right corner, that the first motion vector is the "fourth" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0229] Figure 22 is a flowchart of method 4000 performed on the encoding device side. In step S4001, the process divides an image block into multiple partitions, including a first partition having a non-rectangular shape and a second partition, based on a partition parameter indicating the division. For example, as shown in Figure 15 above, the partition parameter may indicate the direction in which the image block is divided (e.g., from the upper right corner to the lower left corner, or from the upper left corner to the lower right corner). In step S4002, the process encodes the first and second partitions. In step S4003, the process writes one or more parameters, including the partition parameter, to a bitstream that can be received and decoded so that the decoding device side can acquire one or more parameters and perform the same prediction processing on the first and second partitions on the decoding device side (as performed on the encoding device side). The one or more parameters, including the partition parameter, encode various information together or separately, such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used to divide the image block to acquire the first and second partitions, the first motion vector of the first partition, the second motion vector of the second partition, etc.

[0230] Figure 23 is a flowchart of method 5000 performed on the decoding device side. In step S5001, the process reads one or more parameters from a bitstream in which one or more parameters indicate that the image block is divided into multiple partitions, including a first partition and a second partition having a non-rectangular shape. The one or more parameters, including the partition parameters, read from the bitstream may encode together or separately various information necessary for the decoding device to perform the same prediction processing as performed on the encoding device side, such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used to divide the image block to obtain the first and second partitions, the first motion vector of the first partition, the second motion vector of the second partition, etc. In step S5002, the process 5000 divides the image block into multiple partitions based on the partition parameters read from the bitstream. In step S5003, the process decodes the first and second partitions that are divided from the image block.

[0231] Figure 24 is a table of multiple information sets, each encoded as a whole by multiple sample partition parameters ("first index value") and multiple partition parameters, with characteristics similar to the sample table described above in Figure 15. In Figure 24, the partition parameters ("first index value") range from 0 to 6 and encode as a whole the shapes of the first and second partitions separated from the image block, the direction in which the image block is divided into the first and second partitions, the predicted first motion vector for the first partition (Figure 11, step S1002), and the predicted second motion vector for the second partition (Figure 11, step S1002). In particular, partition parameter 0 encodes that neither the first nor the second partition has a triangular shape, and therefore the division direction information is "N / A", the first motion vector information is "N / A", and the second motion vector information is "N / A".

[0232] Partition parameter 1 encodes that the first and second partitions are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the first and second partitions are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the first and second partitions are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the first and second partitions are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the first and second partitions are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition.Partition parameter 6 encodes that the first and second partitions are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0233] According to some implementations, multiple partition parameters (multiple index values) may be binarized according to a binarization scheme selected depending on the values ​​of at least one or more parameters. Figure 16 shows an example of a binarization scheme for binarizing multiple index values ​​(multiple partition parameter values).

[0234] Figure 25 is a table of example combinations of the first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is divided into multiple partitions, including a first partition and a second partition, both having a non-rectangular shape. In this example, the partition parameter may also be used to indicate that the image block is divided without encoding other information encoded by one or more of the other parameters as a whole.

[0235] In the first example in Figure 25, the first parameter is used to indicate the image block size, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the multiple partitions divided from the image block has a triangular shape. Such a combination of the first and second parameters may be used, for example, to indicate that 1) there are no triangular partitions when the image block size is greater than 64 × 64, or 2) there are no triangular partitions when the ratio of the width to height of the image block is greater than 4 (e.g., 64 × 4).

[0236] In the second example of Figure 25, the first parameter is used to indicate the prediction mode, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the multiple partitions divided from the image block has a triangular shape. Such a combination of the first and second parameters may be used, for example, to indicate that there are no triangular partitions when the image block is encoded in intra mode.

[0237] In the third example in Figure 25, the first parameter is used as a partition parameter (flag) to indicate that at least one of the multiple partitions divided from the image block has a triangular shape, and the second parameter is used to indicate the prediction mode. Such a combination of the first and second parameters may be used, for example, to indicate that 1) the image block must be intercoded if at least one of the multiple partitions divided from the image block has a triangular shape.

[0238] In the fourth example in Figure 25, the first parameter indicates the motion vector of adjacent blocks, and the second parameter is used as a partition parameter indicating the direction in which the image block is divided into two triangles. Such a combination of the first and second parameters may be used, for example, to indicate that the direction in which the image block is divided into two triangles is from the upper left corner to the lower right corner when the motion vector of adjacent blocks is diagonal.

[0239] In the fifth example in Figure 25, the first parameter indicates the intra-prediction direction of the adjacent block, and the second parameter is used as a partition parameter indicating the direction in which the image block is divided into two triangles. Such a combination of the first and second parameters may be used, for example, to indicate that if the intra-prediction direction of the adjacent block is in the opposite diagonal direction, the direction in which the image block is divided into two triangles is from the upper right corner to the lower left corner.

[0240] Multiple tables of one or more parameters, including partition parameters, and which information is encoded together or separately, as shown in Figures 15, 24, and 25, are presented only as examples, and it should be understood that many other ways of encoding various information together or separately as part of the partition syntax operation described above are within the scope of this disclosure. For example, a partition parameter may indicate that the first partition is a triangle, trapezoid, or polygon having at least five sides and angles. A partition parameter may indicate that the second partition has a non-rectangular shape, such as a triangle, trapezoid, or polygon having at least five sides and angles. The partition parameter may indicate one or more pieces of information about the division, such as the non-rectangular shape of the first partition, the shape of the second partition (which may be non-rectangular or rectangular), and the division direction applied to divide an image block into multiple partitions (e.g., from the top-left corner to the bottom-right corner of the image block, and from the top-right corner to the bottom-left corner of the image block). The partition parameters encode further information as a whole, such as the first motion vector of the first partition, the second motion vector of the second partition, image block size, prediction mode, motion vector of adjacent blocks, and intra-prediction direction of adjacent blocks. Alternatively, any of this further information may be encoded separately by one or more parameters other than the partition parameters.

[0241] A partition syntax operation like the one shown in Figure 22, process 4000, may be performed by an image encoding device including a circuit and memory connected to the circuit, for example, as shown in Figure 1. The circuit performs a partition syntax operation that includes dividing an image block into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape, based on partition parameters indicating the division (Figure 22, step S4001), encoding the first and second partitions (S4002), and writing one or more parameters, including the partition parameters, as a bitstream (S4003).

[0242] In another embodiment, an image encoding device, as shown in Figure 1, is provided, comprising: a splitting unit 102 that receives a source image in operation and divides it into multiple blocks; an adding unit 104 that receives multiple blocks from the splitting unit and multiple predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block and outputs a residual; a conversion unit 106 that performs a conversion on the multiple residuals output from the adder unit 104 and outputs multiple conversion coefficients; a quantization unit 108 that quantizes the multiple conversion coefficients in operation and generates multiple quantized conversion coefficients; an entropy encoding unit 110 that encodes the multiple quantized conversion coefficients in operation and generates a bitstream; an inter-prediction unit 126 that generates a prediction for the current block based on a reference block in an encoded reference picture in operation; an intra-prediction unit 124 that generates a prediction for the current block based on an encoded reference block in the current picture in operation; and a prediction control unit 128 connected to memories 118 and 122. In operation, the prediction control unit 128 divides an image block into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape, based on partition parameters indicating the division (Figure 22, step S4001), and encodes the first partition and the second partition (step S4002). In operation, the entropy encoding unit 110 writes one or more parameters, including partition parameters, as a bitstream (step S4003).

[0243] In other embodiments, for example, an image decoding device is provided that includes a circuit and a memory connected to the circuit, as shown in Figure 10. The circuit performs a partition syntax operation which includes reading one or more parameters from a bitstream, including partition parameters that indicate that an image block is divided into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape (Figure 23, step S5001); dividing the image block into a plurality of partitions based on the partition parameters (S5002); and decoding the first partition and the second partition (S5003).

[0244] Furthermore, according to the embodiment, the image decoding device shown in Figure 10 includes an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantization conversion coefficients, an inverse quantization unit 204 and an inverse conversion unit 206 that in operation inverse quantization conversion coefficients to obtain a plurality of conversion coefficients and inversely convert the plurality of conversion coefficients to obtain a plurality of residuals, an addition unit 208 that in operation adds a plurality of residuals output from the inverse quantization unit 204 and the inverse conversion unit 206 and a plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter-prediction unit 218 that in operation generates a prediction of the current block based on the reference block in the decoded reference picture, an intra-prediction unit 216 that in operation generates a prediction of the current block based on the decoded reference block in the current picture, and a prediction control unit 220 connected to memories 210 and 214. In operation, the entropy decoding unit 202, in cooperation with the prediction control unit 220 in some implementations, reads one or more parameters from the bitstream, including partition parameters that indicate that the image block is divided into multiple partitions, including a first partition and a second partition having a non-rectangular shape (Figure 23, step S5001). Based on the partition parameters, it divides the image block into multiple partitions (S5002) and decodes the first and second partitions (S5003).

[0245] [Implementation and Application] In each of the above embodiments, each functional or operational block can typically be implemented by an MPU (micro processing unit) and memory, etc. Furthermore, the processing performed by each functional block may be implemented as a program execution unit, such as a processor, that reads and executes software (programs) recorded on a recording medium such as ROM. This software may be distributed. This software may be recorded on various recording media such as semiconductor memory. It is also possible to implement each functional block using hardware (dedicated circuits). Various combinations of hardware and software can be employed.

[0246] The processing described in each embodiment may be implemented by centralized processing using a single device (system), or by distributed processing using multiple devices. Furthermore, the processor executing the above program may be one or multiple. In other words, centralized processing may be performed, or distributed processing may be performed.

[0247] The embodiments of this disclosure are not limited to those described above, and various modifications are possible, which are also included within the scope of the embodiments of this disclosure.

[0248] Furthermore, here we will describe application examples of the video encoding method (image encoding method) or video decoding method (image decoding method) shown in each of the above embodiments, and various systems for implementing these application examples. Such systems may be characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding and decoding device that includes both. Other configurations of such systems can be appropriately modified as needed.

[0249] [Usage example] Figure 26 shows the overall configuration of a suitable content supply system ex100 for realizing a content distribution service. The service area for the communication service is divided into cells of a desired size, and within each cell, there are base stations ex106, ex107, ex108, ex109, and ex110, which are fixed radio stations in the illustrated example.

[0250] In this content supply system ex100, devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 are connected to the Internet ex101 via Internet service provider ex102 or communication network ex104 and base stations ex106 - ex110. The content supply system ex100 may be connected by combining any of the above devices. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network or short - range wireless etc. without going through base stations ex106 - ex110. Further, streaming server ex103 may be connected to devices such as computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via Internet ex101 etc. Also, streaming server ex103 may be connected to terminals etc. within a hotspot in an airplane ex117 via satellite ex116.

[0251] Note that instead of base stations ex106 - ex110, a wireless access point or hotspot etc. may be used. Also, streaming server ex103 may be directly connected to communication network ex104 without going through Internet ex101 or Internet service provider ex102, or may be directly connected to airplane ex117 without going through satellite ex116.

[0252] Camera ex113 is a device capable of still - image shooting and video shooting such as a digital camera etc. Also, smartphone ex115 is a smartphone device, mobile phone, or PHS (Personal Handy - phone System) etc. corresponding to the mobile communication system methods called 2G, 3G, 3.9G, 4G, and in the future 5G.

[0253] Home appliance ex114 is a refrigerator or a device included in a household fuel cell cogeneration system etc.

[0254] In the content supply system ex100, live streaming becomes possible when a terminal with a shooting function is connected to the streaming server ex103 via a base station ex106 or the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal inside an airplane ex117) may perform the encoding process described in each of the above embodiments on still images or video content captured by a user using the terminal, or it may multiplex the video data obtained by encoding with sound data encoded from the sound corresponding to the video, and then transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of this disclosure.

[0255] Meanwhile, the streaming server ex103 streams the content data sent to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117, etc., that is capable of decoding the encoded data. Each device that receives the distributed data may decode and play back the received data. That is, each device may function as an image decoding device according to one aspect of this disclosure.

[0256] [Distributed Processing] Furthermore, the streaming server ex103 may consist of multiple servers or computers that distribute data processing, recording, and distribution. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), where content delivery is achieved through a network connecting numerous edge servers distributed worldwide. In a CDN, the physically closest edge server can be dynamically assigned depending on the client. Latency can be reduced by caching and delivering content to the edge server. In addition, if several types of errors occur or the communication state changes due to increased traffic, processing can be distributed among multiple edge servers, the delivery entity can be switched to another edge server, or delivery can be continued by bypassing the failed part of the network, thus enabling high-speed and stable delivery.

[0257] Furthermore, beyond the distributed processing of the distribution itself, the encoding process of the captured data can be performed on each terminal, on the server side, or shared among them. For example, encoding generally involves two processing loops. In the first loop, the complexity or code amount of the image at the frame or scene level is detected. In the second loop, processing is performed to improve encoding efficiency while maintaining image quality. For example, if the terminal performs the first encoding process and the server that receives the content performs the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode near real time, the first encoded data from the terminal can be received and played back on other terminals, enabling more flexible real-time distribution.

[0258] Another example is the camera ex113, which extracts features (quantities of features or characteristics) from an image, compresses the data related to the features as metadata, and sends it to the server. The server performs compression according to the meaning (or importance of content) of the image, for example, by determining the importance of objects from the features and switching the quantization precision. Feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction during further compression on the server. Alternatively, a simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and a more computationally intensive encoding such as CABAC (Context-Adaptive Binary Arithmetic Coding) may be performed on the server.

[0259] Another example is a scenario in a stadium, shopping mall, or factory where multiple video data sets of nearly identical scenes may exist, captured by multiple terminals. In such cases, the encoding process is distributed among the multiple terminals that captured the footage, along with other terminals and servers as needed, by assigning encoding tasks to each unit, for example, at the Group of Picture (GOP) level, picture level, or tile level (a division of a picture). This reduces latency and enables more real-time performance.

[0260] Since multiple video data sets depict essentially the same scene, the server may manage and / or instruct the video data captured by each terminal to reference each other. Alternatively, the server may receive the encoded data from each terminal, change the reference relationships between the multiple data sets, or correct or replace the pictures themselves and re-encode them. This allows for the creation of a stream with improved quality and efficiency for each individual data set.

[0261] Furthermore, the server may transcode the video data to change its encoding method before distributing it. For example, the server may convert an MPEG-based encoding to a VP-based encoding (e.g., VP9), or convert H.264 to H.265, etc.

[0262] Thus, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, the terms "server" or "terminal" will be used to refer to the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0263] [3D, Multi-angle] It is becoming increasingly common to integrate and utilize images or videos of different scenes, or the same scene, captured from different angles, by multiple cameras ex113 and / or smartphones ex115, which are nearly synchronized with each other. Videos captured by each device can be integrated based on the relative positional relationship between the devices, or on areas where feature points contained in the video coincide.

[0264] The server may not only encode two-dimensional video but also encode still images automatically based on scene analysis of the video, or at a time specified by the user, and transmit them to the receiving terminal. Furthermore, if the server can obtain the relative positional relationship between the shooting terminals, it can generate a three-dimensional shape of the scene based not only on two-dimensional video but also on video of the same scene taken from different angles. The server may separately encode three-dimensional data generated by a point cloud or the like, or it may select or reconstruct video from video taken by multiple terminals to transmit to the receiving terminal based on the results of recognizing or tracking a person or object using the three-dimensional data.

[0265] In this way, users can enjoy scenes by arbitrarily selecting each video corresponding to each shooting terminal, or they can enjoy content in which a video from a selected viewpoint is extracted from 3D data reconstructed using multiple images or videos. Furthermore, along with the video, sound is also collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0266] In recent years, content that links the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create separate viewpoint images for the right and left eyes and perform encoding that allows referencing between the viewpoint images using Multi-View Coding (MVC), or it may encode them as separate streams without referencing each other. When decoding the separate streams, it is advisable to synchronize playback so that the virtual 3D space is reproduced according to the user's viewpoint.

[0267] In the case of AR images, the server may superimpose virtual object information from the virtual space onto camera information from the real space, based on its three-dimensional position or the user's viewpoint movement. The decoding device may acquire or retain the virtual object information and three-dimensional data, generate a two-dimensional image according to the user's viewpoint movement, and create superimposed data by smoothly stitching them together. Alternatively, the decoding device may send the user's viewpoint movement to the server in addition to the request for virtual object information. The server may create superimposed data according to the viewpoint movement received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Typically, superimposed data has an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the three-dimensional data to 0, and encode the data in a state where those parts are transparent. Alternatively, the server may set predetermined RGB values ​​as the background, like in chroma keying, and generate data where parts other than the object are the background color. The predetermined RGB values ​​may be predetermined.

[0268] Similarly, the decryption process of the distributed data can be performed by the client (e.g., a terminal), the server, or a shared task between them. For example, one terminal may send a reception request to the server, another terminal may receive the content corresponding to that request, decrypt it, and then transmit the decrypted signal to a device with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-capable terminals themselves, it is possible to play back data with high image quality. Another example is that while receiving large image data on a TV or similar device, a portion of the picture, such as tiles, may be decrypted and displayed on the viewer's personal terminal. This allows for sharing the overall picture while simultaneously allowing users to check their own area of ​​responsibility or areas they wish to examine in more detail.

[0269] In situations where multiple short-range, medium-range, or long-range wireless communication networks are available both indoors and outdoors, it may be possible to seamlessly receive content using distribution system standards such as MPEG-DASH. Users may freely select and switch in real time between decoding devices or display devices, such as their own terminals or displays located indoors or outdoors. Furthermore, decoding can be performed while switching between the decoding terminal and the display terminal using the user's location information. This makes it possible to map and display information on a part of the wall or ground of an adjacent building with a displayable device embedded, while the user is moving to their destination. It is also possible to switch the bitrate of the received data based on the ease of access to the encoded data on the network, such as when the encoded data is cached on a server that can be accessed quickly from the receiving terminal, or copied to an edge server in the content delivery service.

[0270] [Scalable encoding] Regarding content switching, we will explain using a scalable stream compressed and encoded using the video encoding method described in each of the embodiments above, as shown in Figure 27. The server may have multiple streams with the same content but different qualities as individual streams, but it may also be configured to switch content by taking advantage of the characteristics of a temporally / spatially scalable stream realized by encoding it in layers, as shown in the figure. In other words, the decoding side can freely switch between decoding low-resolution and high-resolution content by deciding which layer to decode according to internal factors such as performance and external factors such as the state of the communication bandwidth. For example, if a user wants to continue watching a video they were watching on their smartphone ex115 while on the go, for example, on an internet TV or other device after returning home, the device only needs to decode the same stream up to a different layer, thus reducing the burden on the server.

[0271] Furthermore, as described above, in addition to a configuration where each layer encodes a picture and scalability is achieved by an enhancement layer above the base layer, the enhancement layer may also include metadata based on image statistics. The decoding side may generate high-quality content by super-resolution the picture in the base layer based on the metadata. Super-resolution may improve the signal-to-noise ratio while maintaining and / or increasing the resolution. The metadata may include information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values ​​in the filtering process, machine learning, or least-squares operation used in the super-resolution process.

[0272] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of an object or the like in the image. The decoding side decodes only a partial area by selecting the tile to be decoded. Further, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 28, the meta information may be stored using a data storage structure different from pixel data, such as an SEI (supplemental enhancement information) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0273] The meta information may be stored in a unit composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. The decoding side can obtain the time when a specific person appears in the video, etc., and by combining the picture unit information and the time information, can specify the picture in which the object exists and determine the position of the object in the picture.

[0274] [Optimization of Web Page] FIG. 29 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 30 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 29 and 30, a web page may include a plurality of link images that are links to image contents, and their appearances may be different depending on the device for browsing. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until a link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) may display a still image or an I picture that each content has as a link image, or may display a video like a gif animation with a plurality of still images or I pictures, etc., or may receive only the base layer, decode and display the video.

[0275] When a linked image is selected by the user, the display device performs decoding, prioritizing the base layer, for example. If the HTML of the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device can decode and display only forward-referenced pictures (I pictures, P pictures, and B pictures that only use forward references), thereby reducing the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display). In addition, the display device may deliberately ignore the reference relationships of the pictures and roughly decode all B pictures and P pictures using forward references, performing normal decoding as time passes and more pictures are received.

[0276] [Autonomous driving] Furthermore, when transmitting and receiving still image or video data, such as 2D or 3D map information, for autonomous driving or driving assistance of a vehicle, the receiving terminal may receive metadata such as weather or construction information in addition to image data belonging to one or more layers, and decode these in association with each other. The metadata may belong to a layer, or it may simply be multiplexed with the image data.

[0277] In this case, since the vehicle, drone, or airplane containing the receiving terminal is moving, the receiving terminal can transmit its own location information, enabling seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much metadata is received or how much map information is updated, depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0278] The content delivery system ex100 allows the client to receive, decode, and play back encoded information transmitted by the user in real time.

[0279] [Distribution of personal content] Furthermore, the ex100 content delivery system allows for unicast or multicast distribution of not only high-definition, long-duration content from video distribution companies, but also low-definition, short-duration content from individuals. It is expected that the amount of such individual content will continue to increase. To improve the quality of individual content, the server may perform editing before encoding. This can be achieved, for example, using a configuration like the following.

[0280] During shooting, or after shooting, the server performs recognition processing such as detecting shooting errors, searching for scenes, analyzing semantics, and detecting objects from the original image data or encoded data in real time. Based on the recognition results, the server manually or automatically edits the images, correcting out-of-focus or shaky images, deleting less important scenes such as those with lower brightness or out of focus compared to other pictures, emphasizing object edges, and altering color tones. The server then encodes the edited data based on the editing results. It is also known that viewership decreases if the shooting time is too long, so the server may automatically clip scenes with little movement, as well as less important scenes, based on the image processing results, to ensure that the content falls within a specific time range according to the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.

[0281] Personal content may contain elements that infringe on copyright, moral rights, or portrait rights, and the scope of sharing may exceed the intended scope, which can be inconvenient for the individual. Therefore, for example, the server may intentionally change the image to one that is out of focus, such as the faces of people at the edges of the screen or the interior of a house, before encoding. Furthermore, the server may recognize whether the face of a person other than those previously registered is visible in the image to be encoded, and if so, it may apply a mosaic effect to the face. Alternatively, as a pre-processing or post-processing step before encoding, the user may specify a person or background area that they want to process from a copyright perspective. The server may replace the specified area with another image or blur the focus. In the case of a person, the server can track the person in a video and replace the image of the person's face.

[0282] Viewing personal content with small data volumes requires real-time processing, so depending on the bandwidth, the decoder may prioritize receiving the base layer first and then decode and play it back. During this time, the decoder may receive the enhancement layer, and if playback is looped or if the content is played back more than once, it may play back the high-quality video including the enhancement layer. With a stream that uses this scalable encoding, it is possible to provide an experience where the video is rough when unselected or at the beginning of viewing, but gradually the stream becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if the rough stream played back the first time and the second stream encoded by referencing the first video are configured as a single stream.

[0283] [Other examples of practical applications] Furthermore, these encoding or decoding processes are generally performed in the LSIex500 present in each terminal. The LSI (large scale integration circuitry) ex500 (see Figure 26) may be a single chip or a configuration consisting of multiple chips. Alternatively, video encoding or decoding software may be embedded in some recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by a computer ex111, and the encoding or decoding process may be performed using that software. In addition, if the smartphone ex115 has a camera, video data acquired by that camera may be transmitted. This video data may be data encoded by the LSIex500 present in the smartphone ex115.

[0284] The LSIex500 may also be configured to be activated by downloading application software. In this case, the terminal first determines whether it supports the content encoding method or whether it has the capability to perform the specific service. If the terminal does not support the content encoding method or does not have the capability to perform the specific service, the terminal may download a codec or application software and then acquire and play the content.

[0285] Furthermore, not only the content supply system ex100 via the Internet ex101, but also digital broadcasting systems can incorporate at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments. While the content supply system ex100 has a configuration that is more suited to multicast than unicast, as it transmits and receives multiplexed data with video and sound multiplexed onto broadcast radio waves using satellites, etc., the encoding and decoding processes are similar and can be applied in the same way.

[0286] [Hardware configuration] Figure 31 shows further details of the smartphone ex115 shown in Figure 26. Figure 32 shows an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves with the base station ex110, a camera unit ex465 capable of taking video and still images, and a display unit ex458 that displays video captured by the camera unit ex465 and data decoded from video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466, which is a touch panel, an audio output unit ex457, which is a speaker for outputting voice or sound, an audio input unit ex456, which is a microphone for inputting voice, a memory unit ex467 capable of storing captured video or still images, recorded audio, received video or still images, encoded data such as emails, or decoded data, and a slot unit ex464, which is an interface unit with SIM ex468 for identifying the user and authenticating access to various data, including the network. External memory may be used instead of the memory unit ex467.

[0287] The main control unit ex460, which can comprehensively control the display unit ex458 and the operation unit ex466, is connected to the power supply circuit unit ex461, the operation input control unit ex462, the video signal processing unit ex455, the camera interface unit ex463, the display control unit ex459, the modulation / demodulation unit ex452, the multiplexing / decompression unit ex453, the audio signal processing unit ex454, the slot unit ex464, and the memory unit ex467 via the synchronization bus ex470.

[0288] The power supply circuit unit ex461, when the power key is turned on by the user, starts up the smartphone ex115 into an operational state and supplies power to each component from the battery pack.

[0289] The smartphone ex115 performs tasks such as voice calls and data communication based on the control of the main control unit ex460, which has a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, spread spectrum processing is performed by the modulation / demodulation unit ex452, digital-to-analog conversion and frequency conversion processing are performed by the transmission / reception unit ex451, and the resulting signal is transmitted via the antenna ex450. Received data is amplified, subjected to frequency conversion and analog-to-digital conversion processing, despread spectrum processing is performed by the modulation / demodulation unit ex452, converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. In data communication mode, text, still images, or video data can be transmitted via the operation input control unit ex462 under the control of the main control unit ex460 based on operations such as those performed by the operation unit ex466 on the main unit. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the video encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / decoding unit ex453. The audio signal processing unit ex454 encodes the audio signal picked up by the audio input unit ex456 while the camera unit ex465 is capturing video or still images, and sends the encoded audio data to the multiplexing / decoding unit ex453. The multiplexing / decoding unit ex453 multiplexes the encoded video data and encoded audio data in a predetermined manner, performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits via the antenna ex450. The predetermined manner may be set in advance.

[0290] When receiving video attached to an email or chat, or video linked to a webpage, etc., the multiplexing / decomposition unit ex453 separates the multiplexed data received via antenna ex450, dividing it into a video data bitstream and an audio data bitstream. It then supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and the video or still image contained in the linked video file is displayed from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. As real-time streaming is becoming increasingly widespread, audio playback may be socially inappropriate depending on the user's situation. Therefore, as an initial setting, it is preferable to play only the video data without playing the audio signal, and to synchronize the audio playback only when the user performs an action such as clicking on the video data.

[0291] While the smartphone ex115 was used as an example here, other implementation forms are possible for terminals, such as a transmitting terminal with only an encoder and a receiving terminal with only a decoder, in addition to a transmitting / receiving terminal that has both an encoder and a decoder. In the explanation for digital broadcasting systems, multiplexed data in which audio data is multiplexed with video data is received or transmitted. However, the multiplexed data may also contain text data related to the video in addition to audio data. Furthermore, the video data itself may be received or transmitted instead of multiplexed data.

[0292] Although it was explained that the main control unit ex460, including the CPU, controls the encoding or decoding process, various terminals often have a GPU. Therefore, a configuration that leverages the GPU's performance to process a wide area at once using memory shared by the CPU and GPU, or memory whose addresses are managed so that it can be used in common, is also possible. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform motion detection, deblocking filters, SAO (Sample Adaptive Offset), and transformation / quantization processes at once on the GPU, rather than on the CPU, in units such as pictures.

Claims

1. Circuits and, The circuit comprises a memory connected to the aforementioned circuit, The circuit, in its operation, performs processing along the boundary between a first partition having a non-rectangular shape in the image block and a second partition, and the processing is as follows: The partition parameters relating to the shapes of the first partition and the second partition are to be coded, Selecting the first motion vector of the first partition from the first motion vector candidate set, Using the first motion vector, a first prediction is made of multiple first values ​​of the pixel set of the first partition, Selecting the second motion vector of the second partition from the first motion vector candidate set, Using the second motion vector, a second prediction is made of multiple second values ​​of the pixel set of the first partition, This includes weighting the plurality of first values ​​and the plurality of second values, If the ratio of the width of the image block to the height of the image block is greater than 4, or if the ratio of the height to the width is greater than 4, the circuit disables the process. Image encoding device.

2. Circuits and, The circuit comprises a memory connected to the aforementioned circuit, The circuit, in its operation, performs processing along the boundary between a first partition having a non-rectangular shape in the image block and a second partition, and the processing is as follows: The first partition and the second partition are defined based on partition parameters relating to the shapes of the first partition and the second partition, Selecting the first motion vector of the first partition from the first motion vector candidate set, Using the first motion vector, a first prediction is made of multiple first values ​​of the pixel set of the first partition, Selecting the second motion vector of the second partition from the first motion vector candidate set, Using the second motion vector, a second prediction is made of multiple second values ​​of the pixel set of the first partition, This includes weighting the plurality of first values ​​and the plurality of second values, If the ratio of the width of the image block to the height of the image block is greater than 4, or if the ratio of the height to the width is greater than 4, the circuit disables the process. Image decoding device.

3. Circuits and, A bitstream generation device comprising a memory connected to the aforementioned circuit, In operation, the aforementioned circuit Information regarding the size of the image block and partition parameters relating to the shapes of the first and second partitions are generated to cause the decoder to perform partition processing along the boundary between the first and second partitions having a non-rectangular shape within the image block. The bitstream includes the aforementioned information and the partition parameters. The aforementioned partitioning process is, The first motion vector of the first partition in the image block is selected from the first motion vector candidate set, Multiple first values ​​of the pixel set of the first partition are calculated using the first motion vector, The second motion vector of the second partition is selected from the first motion vector candidate set, Multiple second values ​​of the aforementioned pixel set are calculated using the second motion vector, The plurality of third values ​​of the pixel set are calculated by weighting the plurality of first values ​​and the plurality of second values, If the ratio of the width of the image block to the height of the image block is greater than 4, or if the ratio of the height to the width is greater than 4, the bitstream generator disables the partitioning process. Bitstream generator.