Image encoding device, image decoding device, and bit stream generating device

By dividing image blocks into non-rectangular partitions and applying motion vectors, the solution enhances video coding efficiency and speed, addressing the challenges of handling large video data volumes in video coding technologies.

JP2026012309AActive Publication Date: 2026-01-23PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025182295
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-07-16
Filing Date
2025-10-29
Publication Date
2026-01-23
Estimated Expiration
2038-08-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently handling increasing amounts of digital video data, particularly in optimizing inter- or intra-prediction processes that divide image blocks into non-rectangular shapes, which can complicate encoding and decoding processes.

Method used

The proposed solution involves dividing image blocks into multiple partitions, including a first partition with a non-rectangular shape, such as a triangle, and selecting motion vectors for each partition, followed by a boundary smoothing process to enhance encoding efficiency and decoding accuracy.

Benefits of technology

This approach improves encoding efficiency, simplifies the encoding and decoding processes, and accelerates their speed by effectively utilizing appropriate filters and motion vectors for non-rectangular partitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012309000001_ABST
    Figure 2026012309000001_ABST
Patent Text Reader

Abstract

To provide an image encoding device for performing efficient encoding.SOLUTION: The image coding device 100 includes a circuit and a memory, and the circuit is configured to (a) select a first motion vector for a first partition from a first motion vector candidate set and obtain a first value for the first partition by using the first motion vector, and (b) select a second motion vector for a second partition from the first motion vector candidate set and obtain a second value for the second partition by using the second motion vector; (c) weight the first value and the second value for a plurality of pixels where the first partition and the second partition overlap, (d) perform a boundary smoothing process that encodes the image block using the weighted first value and the weighted second value, and disable the boundary smoothing process when the intra prediction is applied to the image block and the ratio of the width to the height is greater than 4 or the ratio of the height to the width is greater than 4.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to video coding, and to systems, components, and methods for encoding and decoding moving images, for example, by performing inter prediction to construct a current block based on a reference frame, or by performing intra prediction to construct a current block based on an encoded / decoded reference block within the current frame. [Background technology]

[0002] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there is a constant need to provide improvements and optimizations in video coding technology to handle ever-increasing amounts of digital video data in various applications. This disclosure further relates to advances, improvements, and optimizations in video coding, and in particular to inter- or intra-prediction that divides an image block into multiple partitions, including at least a first partition having a non-rectangular shape (e.g., a triangle) and a second partition. Summary of the Invention

[0003] According to one aspect, an image coding device is provided, comprising: a circuit and a memory coupled to the circuit, wherein the circuit is operable to determine whether intra prediction is applied to an image block; and if the intra prediction is not applied to the image block, to: (a) select a first motion vector for a boundary between a first partition having a non-rectangular shape in the image block and a second partition having a non-rectangular shape in the image block from a first motion vector candidate set and use the first motion vector to determine a first value for the first partition; and (b) select a second motion vector for the second partition from the first motion vector candidate set based on a first parameter. (c) weighting the first value and the second value for a plurality of pixels where the first partition and the second partition overlap; and (d) performing a boundary smoothing process to encode the image block using the weighted first value and the weighted second value, and disabling the boundary smoothing process for the image block when the intra prediction is applied to the image block and a ratio of the width of the image block to the height of the image block is greater than 4 or when a ratio of the height to the width is greater than 4.

[0004] Some implementations of the present disclosure may improve encoding efficiency, simplify encoding / decoding processes, accelerate encoding / decoding process speed, and efficiently select appropriate components / operations to be used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.

[0005] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. Benefits and / or advantages may be obtained by the various embodiments and features of the specification and drawings individually, and it is not necessary for all of the various embodiments and features of the specification and drawings to be present in order to obtain one or more such benefits and / or advantages.

[0006] It should be noted that the generic or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any combination thereof. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of an encoding device according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an example of block division. [Figure 3] FIG. 3 is a table showing the transformation basis functions corresponding to each transformation type. [Figure 4A] FIG. 4A is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter). [Figure 4B] FIG. 4B is a diagram showing another example of the shape of the filter used in ALF. [Figure 4C] FIG. 4C is a diagram showing another example of the shape of the filter used in ALF. [Figure 5A] FIG. 5A is a diagram showing 67 intra prediction modes in intra prediction. [Figure 5B] FIG. 5B is a flowchart for explaining an outline of the predictive image correction process using OBMC (overlapped block motion compensation) processing. [Figure 5C] FIG. 5C is a conceptual diagram for explaining an outline of the predicted image correction process using the OBMC process. [Figure 5D] FIG. 5D is a diagram showing an example of FRUC (frame rate up-conversion). [Figure 6] FIG. 6 is a diagram for explaining pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 7] FIG. 7 is a diagram for explaining pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 8] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. [Figure 9A] FIG. 9A is a diagram for explaining derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 9B] FIG. 9B is a diagram for explaining an outline of the motion vector derivation process in the merge mode. [Figure 9C] FIG. 9C is a conceptual diagram for explaining an outline of DMVR (dynamic motion vector refreshing) processing. [Figure 9D] FIG. 9D is a diagram for explaining an outline of a predicted image generation method using luminance correction processing by LIC (local illumination compensation) processing. [Figure 10] FIG. 10 is a block diagram illustrating a functional configuration of a decoding device according to an embodiment. [Figure 11] FIG. 11 is a flowchart illustrating an overall processing flow for dividing an image block into multiple partitions, including at least a first partition having a non-rectangular shape (e.g., a triangle), and a second partition for further processing, according to one embodiment. [Figure 12] FIG. 12 illustrates two exemplary methods for dividing an image block into a first partition having a non-rectangular shape (e.g., a triangle) and a second partition (also having a non-rectangular shape in the examples shown). [Figure 13] FIG. 13 illustrates an example of a boundary smoothing process that involves weighting multiple first values ​​of multiple boundary pixels predicted based on a first partition and multiple second values ​​of multiple boundary pixels predicted based on a second partition. [Figure 14] FIG. 14 illustrates three further examples of boundary smoothing processes that involve weighting first values ​​of boundary pixels predicted based on a first partition and second values ​​of boundary pixels predicted based on a second partition. [Figure 15]FIG. 15 is a table showing example parameters ("first index values") and a plurality of information sets respectively encoded by the parameters. [Figure 16] FIG. 16 is a table showing binarization of a plurality of parameters (a plurality of index values). [Figure 17] FIG. 17 is a flowchart illustrating a process for dividing an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape. [Figure 18] FIG. 18 is a diagram illustrating examples of dividing an image block into partitions, including a first partition and a second partition having a non-rectangular shape, which in the examples shown is a triangle. [Figure 19] FIG. 19 illustrates several further examples of dividing an image block into several partitions, including a first partition having a non-rectangular shape that is a polygon with at least five sides and corners in the examples shown, and a second partition. [Figure 20] FIG. 20 is a flowchart illustrating a boundary smoothing process that involves weighting a plurality of first values ​​of a plurality of boundary images predicted based on a first partition and a plurality of second values ​​of a plurality of boundary images predicted based on a second partition. [Figure 21A] FIG. 21A is a diagram illustrating an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Figure 21B] FIG. 21B is a diagram illustrating an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Figure 21C]FIG. 21C is a diagram illustrating an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Figure 21D] FIG. 21D is a diagram illustrating an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Figure 22] Figure 22 is a flowchart showing a method performed on the encoding device side to divide an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape, based on partition parameters indicating the division, and to write one or more parameters including the partition parameters into a bitstream during entropy encoding. [Figure 23] Figure 23 is a flowchart showing a method performed on a decoding device side that reads one or more parameters from a bitstream, including a partition parameter indicating division of an image block into multiple partitions including a first partition having a non-rectangular shape and a second partition, divides the image block into multiple partitions based on the partition parameter, and decodes the first partition and the second partition. [Figure 24] FIG. 24 is a table showing example partition parameters ("first index values") each indicating division of an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape, and multiple sets of information that may each be encoded together by the partition parameters. [Figure 25] FIG. 25 is a table of example combinations of first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is to be divided into multiple partitions, including a first partition and a second partition having a non-rectangular shape. [Figure 26] FIG. 26 is a block diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 27] FIG. 27 is a conceptual diagram showing an example of a coding structure in scalable coding. [Figure 28] FIG. 28 is a conceptual diagram showing an example of a coding structure for scalable coding. [Figure 29] FIG. 29 is a conceptual diagram showing an example of a display screen of a web page. [Figure 30] FIG. 30 is a conceptual diagram showing an example of a display screen of a web page. [Figure 31] FIG. 31 is a block diagram showing an example of a smartphone. [Figure 32] FIG. 32 is a block diagram showing an example of the configuration of a smartphone. DETAILED DESCRIPTION OF THE INVENTION

[0008] According to one aspect, there is provided an image coding device comprising: a circuit and a memory coupled to the circuit, the circuit being operative to: divide an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predict a first motion vector for the first partition, predict a second motion vector for the second partition, encode the first partition using the first motion vector, and encode the second partition using the second motion vector.

[0009] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and corners.

[0010] According to another aspect, the predicting includes selecting the first motion vector from a first motion vector candidate set and selecting the second motion vector from a second motion vector candidate set. For example, the first motion vector candidate set may include motion vectors of partitions adjacent to the first partition, and the second motion vector candidate set may include motion vectors of partitions adjacent to the second partition. The partitions adjacent to the first partition and the partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The adjacent partitions may be one or both of spatially adjacent partitions and temporally adjacent partitions. The first motion vector candidate set may be the same as or different from the second motion vector candidate set.

[0011] According to another aspect, the predicting includes selecting a first motion vector candidate from a first motion vector candidate set and deriving the first motion vector by adding a first differential motion vector to the first motion vector candidate, and selecting a second motion vector candidate from a second motion vector candidate set and deriving the second motion vector by adding a second differential motion vector to the second motion vector candidate.

[0012] According to another aspect, there is provided an image encoding device comprising: a division unit that, in operation, receives an original image and divides it into a plurality of blocks; an addition unit that, in operation, receives the plurality of blocks from the division unit and a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a transformation unit that, in operation, performs a transform on the plurality of residuals output from the addition unit and outputs a plurality of transform coefficients; a quantization unit that, in operation, quantizes the plurality of transform coefficients to generate a plurality of quantized transform coefficients; an entropy encoding unit that, in operation, encodes the plurality of quantized transform coefficients to generate a bitstream; an inter prediction unit that, in operation, generates a prediction of a current block based on a reference block in an encoded reference picture; an intra prediction unit that, in operation, generates a prediction of the current block based on an encoded reference block in the current picture; and the prediction control unit connected to a memory. In operation, the prediction control unit divides the multiple blocks into multiple partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, encodes the first partition using the first motion vector, and encodes the second partition using the second motion vector.

[0013] According to another aspect, there is provided an image encoding method including dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition and a second motion vector for the second partition, encoding the first partition using the first motion vector and encoding the second partition using the second motion vector.

[0014] According to one aspect, there is provided an image decoding device comprising: a circuit and a memory coupled to the circuit, the circuit being operative to: divide an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; predict a first motion vector for the first partition, predict a second motion vector for the second partition, decode the first partition using the first motion vector, and decode the second partition using the second motion vector.

[0015] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and corners.

[0016] According to another aspect, there is provided an image decoding device comprising: an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an adder that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit to a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; an inter prediction unit that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and the prediction control unit connected to a memory. In operation, the prediction control unit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, decodes the first partition using the first motion vector, and decodes the second partition using the second motion vector.

[0017] According to another aspect, there is provided an image decoding method including dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition and a second motion vector for the second partition, decoding the first partition using the first motion vector and decoding the second partition using the second motion vector.

[0018] According to one aspect, an image coding apparatus is provided, comprising: a circuit and a memory coupled to the circuit, the circuit being operable to perform a boundary smoothing operation along a boundary between a first partition divided from an image block and having a non-rectangular shape, and a second partition, the boundary smoothing operation including: first predicting a plurality of first values ​​of a set of pixels of the first partition along the boundary using information about the first partition; second predicting a plurality of second values ​​of the set of pixels of the first partition along the boundary using information about the second partition; weighting the plurality of first values ​​and the plurality of second values; and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0019] According to a further aspect, the non-rectangular shape is a triangle. According to another aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and corners. According to yet another aspect, the second partition has a non-rectangular shape.

[0020] According to another aspect, at least one of the first prediction and the second prediction is an inter prediction process that predicts the first values ​​and the second values ​​based on a reference partition in a coded reference picture, and the inter prediction process may predict the first values ​​of pixels in the first partition that includes the set of pixels, or may predict the second values ​​only for the set of pixels in the first partition.

[0021] According to another aspect, at least one of the first prediction and the second prediction is an intra prediction process that predicts the plurality of first values ​​and the plurality of second values ​​based on an already coded reference partition in the current picture.

[0022] According to another aspect, the prediction method used for the first prediction is different from the prediction method used for the second prediction.

[0023] According to a further aspect, the number of pixel sets in each row or column for predicting the first values ​​and the second values ​​is an integer. For example, if the number of pixel sets in each row or column is four, weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the first values ​​of the four pixels in the pixel set, respectively, and weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the second values ​​of the four pixels in the pixel set, respectively. As another example, if the number of pixel sets in each row or column is two, weights of 1 / 3 and 2 / 3 may be applied to the first values ​​of the two pixels in the pixel set, respectively, and weights of 2 / 3 and 1 / 3 may be applied to the second values ​​of the two pixels in the pixel set, respectively.

[0024] According to another aspect, the weights may be integer or fractional values.

[0025] According to another aspect, there is provided an image coding device including: a partitioning unit that receives an original image and divides it into a plurality of blocks; an adder that receives the blocks from the partitioning unit and a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a transformer that performs a transform on the residuals output from the adder and outputs a plurality of transform coefficients; a quantization unit that quantizes the transform coefficients to generate a plurality of quantized transform coefficients; an entropy coding unit that encodes the quantized transform coefficients to generate a bitstream; an inter prediction unit that generates a prediction of a current block based on a reference block in a coded reference picture; an intra prediction unit that generates a prediction of the current block based on a coded reference block in the current picture; and the prediction control unit connected to a memory. The prediction control unit performs a boundary smoothing operation along a boundary between a first partition and a second partition, each having a non-rectangular shape and divided from the image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of a set of pixels in the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the set of pixels in the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0026] According to another aspect, there is provided an image coding method for performing a boundary smoothing operation along a boundary between a first partition divided from an image block and having a non-rectangular shape and a second partition, the method generally including four steps: first predicting a plurality of first values ​​of a set of pixels of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the set of pixels of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0027] According to a further aspect, an image decoding device is provided, comprising: a circuit and a memory coupled to the circuit. The circuit is operative to perform a boundary smoothing operation along a boundary between a first partition divided from an image block and having a non-rectangular shape and a second partition. The boundary smoothing operation includes: first predicting a plurality of first values ​​of a set of pixels of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the set of pixels of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0028] According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and corners. According to another aspect, the second partition has a non-rectangular shape.

[0029] According to another aspect, at least one of the first prediction and the second prediction is an inter prediction process that predicts the first values ​​and the second values ​​based on a reference partition in a coded reference picture, and the inter prediction process may predict the first values ​​of pixels in the first partition that includes the set of pixels, or may predict the second values ​​only for the set of pixels in the first partition.

[0030] According to another aspect, at least one of the first prediction and the second prediction is an intra prediction process that predicts the plurality of first values ​​and the plurality of second values ​​based on an already coded reference partition in the current picture.

[0031] According to another aspect, there is provided an image decoding device including: an entropy decoding unit configured to receive and decode an encoded bitstream to obtain a plurality of quantized transform coefficients; an inverse quantization unit and an inverse transform unit configured to inversely quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and to inversely transform the plurality of transform coefficients to obtain a plurality of residuals; an adder configured to reconstruct a plurality of blocks by adding the plurality of residuals output from the inverse quantization unit and the inverse transform unit to a plurality of predictions output from a prediction control unit; an inter prediction unit configured to generate a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit configured to generate a prediction of the current block based on the decoded reference block in the current picture; and the prediction control unit connected to a memory. In the prediction control unit, the prediction control unit is configured to perform a boundary smoothing operation along a boundary between a first partition and a second partition, the first partition having a non-rectangular shape and divided from the image block. The boundary smoothing operation includes: first predicting a plurality of first values ​​of a set of pixels in the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the set of pixels in the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0032] According to another aspect, there is provided an image decoding method for performing a boundary smoothing operation along a boundary between a first partition divided from an image block and having a non-rectangular shape and a second partition, the method generally including four steps of: first predicting a plurality of first values ​​of a set of pixels of the first partition along the boundary using information of the first partition; second predicting a plurality of second values ​​of the set of pixels of the first partition along the boundary using information of the second partition; weighting the plurality of first values ​​and the plurality of second values; and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values.

[0033] According to one aspect, there is provided an image coding apparatus comprising: a circuit and a memory coupled to the circuit, the circuit being operable to perform a partition syntax operation, the partition syntax operation including: dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on a partition parameter indicating the division; encoding the first partition and the second partition; and writing one or more parameters including the partition parameter to a bitstream.

[0034] According to a further aspect, the partition parameters indicate that the first partition has a triangular shape.

[0035] According to another aspect, the partition parameters indicate that the second partition has a non-rectangular shape.

[0036] According to another aspect, the partition parameters indicate that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and corners.

[0037] According to another aspect, the partition parameters jointly encode a division direction applied to divide the image block into the plurality of partitions. For example, the division direction may include a direction from the upper left corner of the image block to the lower right corner thereof, and a direction from the upper right corner of the image block to the lower left corner thereof. The partition parameters jointly encode at least a first motion vector of the first partition.

[0038] According to another aspect, one or more parameters other than the partition parameter encode a division direction applied to divide the image block into the plurality of partitions, and the parameter encoding the division direction may encode at least a first motion vector of the first partition together.

[0039] According to another aspect, the partition parameters may include encoding at least a first motion vector for the first partition as a whole, and a second motion vector for the second partition as a whole.

[0040] According to another aspect, the one or more parameters other than the partition parameters may encode at least a first motion vector for the first partition.

[0041] According to another aspect, the one or more parameters are binarized according to a binarization scheme selected depending on the value of at least one of the one or more parameters.

[0042] According to a further aspect, there is provided an image encoding device including: a partitioning unit that receives an original image and divides it into a plurality of blocks; an adder that receives the blocks from the partitioning unit and a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual; a transformer that performs a transform on the plurality of residuals output from the adder and outputs a plurality of transform coefficients; a quantizer that quantizes the plurality of transform coefficients to generate a plurality of quantized transform coefficients; an entropy encoding unit that encodes the plurality of quantized transform coefficients to generate a bitstream; an inter prediction unit that generates a prediction of a current block based on a reference block in a previously encoded reference picture; an intra prediction unit that generates a prediction of the current block based on a previously encoded reference block in the current picture; and the prediction control unit connected to a memory. The prediction control unit divides the image block into a plurality of partitions, including a first partition and a second partition, each having a non-rectangular shape, based on a partition parameter indicating the division, and encodes the first partition and the second partition. The entropy coding unit operates to write one or more parameters, including the partition parameters, into a bitstream.

[0043] According to another aspect, there is provided an image coding method including a partition syntax operation, which generally includes three steps: dividing an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on a partition parameter indicating the division; encoding the first partition and the second partition; and writing one or more parameters including the partition parameter into a bitstream.

[0044] According to another aspect, there is provided an image decoding device comprising: a circuit and a memory coupled to the circuit, wherein the circuit is operable to perform a partition syntax operation including: interpreting one or more parameters from a bitstream, the parameters including partition parameters indicating partitioning of an image block into a plurality of partitions, the partition parameters including a first partition having a non-rectangular shape and a second partition; partitioning the image block into the plurality of partitions based on the partition parameters; and decoding the first partition and the second partition.

[0045] According to a further aspect, the partition parameters indicate that the first partition has a triangular shape.

[0046] According to another aspect, the partition parameters indicate that the second partition has a non-rectangular shape.

[0047] According to another aspect, the partition parameters indicate that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and corners.

[0048] According to another aspect, the partition parameters jointly encode a division direction applied to divide the image block into the plurality of partitions, for example, the division direction includes from the upper left corner of the image block to its lower right corner and from the upper right corner of the image block to its lower left corner, and the partition parameters jointly encode at least a first motion vector of the first partition.

[0049] According to another aspect, one or more parameters other than the partition parameter encode a division direction applied to divide the image block into the plurality of partitions, and the parameter encoding the division direction may encode at least a first motion vector of the first partition together.

[0050] According to another aspect, the partition parameters may include encoding at least a first motion vector for the first partition as a whole, and a second motion vector for the second partition as a whole.

[0051] According to another aspect, the one or more parameters other than the partition parameters may encode at least a first motion vector for the first partition.

[0052] According to another aspect, the one or more parameters are binarized according to a binarization scheme selected depending on the value of at least one of the one or more parameters.

[0053] According to a further aspect, there is provided an image decoding device comprising: an entropy decoding unit that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients; an inverse quantization unit and an inverse transform unit that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an adder that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit to a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks; an inter prediction unit that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and the prediction control unit connected to a memory. In operation, the entropy decoding unit decodes one or more parameters from a bitstream, including a partition parameter indicating division of an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, and divides the image block into the plurality of partitions based on the partition parameter, and decodes the first partition and the second partition.

[0054] According to another aspect, there is provided an image decoding method including a partition syntax operation, which generally includes three steps: decoding one or more parameters from a bitstream, including partition parameters indicating partitioning of an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition; partitioning the image block into the plurality of partitions based on the partition parameters; and decoding the first partition and the second partition.

[0055] In the drawings, the same reference numerals refer to the same elements, and the sizes and relative positions of elements in the drawings are not necessarily to scale.

[0056] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection forms, steps, step relationships and order, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Therefore, components disclosed in the following embodiments but not recited in the independent claims defining the broadest inventive concept may be understood as optional components.

[0057] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with regard to the processes and / or configurations applied to the embodiments, any of the following may be implemented.

[0058] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0059] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, or deleting a function or process. For example, any function or process may be replaced with or combined with another function or process described in any of the aspects of the present disclosure.

[0060] (3) In the method implemented by the encoding device or decoding device of the embodiment, some of the processes included in the method may be arbitrarily modified, such as by addition, replacement, deletion, etc. For example, any process in the method may be replaced with or combined with another process described in any of the aspects of the present disclosure.

[0061] (4) Some of the components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in any of the aspects of the present disclosure.

[0062] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0063] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the multiple processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.

[0064] (7) Some of the processes included in the method implemented by the encoding device or decoding device of the embodiments may be combined with the processes described in any of the aspects of the present disclosure.

[0065] (8) The implementation of the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for a purpose other than the video encoding or video decoding disclosed in the embodiments.

[0066] [Outline of the encoding device] First, an overview of a coding device according to an embodiment will be described. Fig. 1 is a block diagram showing the functional configuration of a coding device 100 according to an embodiment. The coding device 100 is a video coding device that codes a video on a block-by-block basis.

[0067] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0068] The encoding device 100 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. Alternatively, the encoding device 100 may be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0069] Each component included in the encoding device 100 will be described below.

[0070] [Divided part] The division unit 102 divides each picture included in the input video into multiple blocks and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides the picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks may be called coding tree units (CTUs). The division unit 102 then divides each of the fixed-size blocks into blocks of a variable size (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block division. These variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). In various processing examples, CUs, PUs, and TUs do not need to be distinguished from one another, and some or all of the blocks in a picture may serve as the processing units of CUs, PUs, and TUs.

[0071] Fig. 2 is a diagram showing an example of block division in an embodiment, in which solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0072] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0073] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0074] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14 and 15 (binary tree block division).

[0075] The lower-left 64x64 block is divided into four square 32x32 blocks (quadtree block decomposition). Of the four 32x32 blocks, the upper-left and lower-right blocks are further divided. The upper-left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block decomposition). The lower-right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block decomposition). As a result, the lower-left 64x64 block is divided into 16x32 block 16, two 16x16 blocks 17 and 18, two 32x32 blocks 19 and 20, and two 32x16 blocks 21 and 22.

[0076] The bottom right 64x64 block 23 is not split.

[0077] 2, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. This type of division is sometimes called QTBT (quad-tree plus binary tree) division.

[0078] In Fig. 2, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0079] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (a prediction sample input from a prediction control unit 128 described below) from the original signal (original sample) input from the division unit 102, for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also referred to as a residual) of a block to be coded (hereinafter referred to as a current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.

[0080] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing an image may also be referred to as a sample.

[0081] [Conversion section] The transform unit 106 transforms the spatial domain prediction errors into frequency domain transform coefficients and outputs the transform coefficients to the quantization unit 108. Specifically, the transform unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the spatial domain prediction errors.

[0082] The transform unit 106 may adaptively select a transform type from among a plurality of transform types and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform is sometimes called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0083] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Fig. 3 is a table showing transform basis functions corresponding to each transform type. In Fig. 3, N represents the number of input pixels. Selection of a transform type from among these multiple transform types may depend, for example, on the type of prediction (intra prediction or inter prediction) or the intra prediction mode.

[0084] Such information indicating whether EMT or AMT is applied (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transformation type are usually signaled at the CU level, but the signaling of this information does not need to be limited to the CU level and may be at other levels (e.g., the bit sequence level, picture level, slice level, tile level, or CTU level).

[0085] Furthermore, the transform unit 106 may retransform the transform coefficients (transform results). Such retransformation may be referred to as an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transform unit 106 performs retransformation for each sub-block (e.g., 4x4 sub-block) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether to apply NSST and information regarding the transform matrix used for NSST are typically signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may also be at other levels (e.g., the sequence level, picture level, slice level, tile level, or CTU level).

[0086] Separable transformation and non-separable transformation may be applied to the transformation unit 106. Separable transformation is a method of separating the input into directions for the number of dimensions and performing transformation multiple times, and non-separable transformation is a method of treating two or more dimensions of a multi-dimensional input as one dimension and performing transformation all at once.

[0087] For example, an example of a non-separable transformation is when the input is a 4x4 block, it is treated as a single array with 16 elements, and a 16x16 transformation matrix is ​​used to perform transformation processing on that array.

[0088] Another example of a non-separable transformation is to treat a 4x4 input block as a single array with 16 elements, and then perform a transformation (e.g., a Hypercube Givens Transform) on the array by performing multiple Givens rotations.

[0089] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order and quantizes the transform coefficients based on quantization parameters (QP) corresponding to the scanned transform coefficients. The quantization unit 108 then outputs the quantized transform coefficients of the current block (hereinafter referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112.

[0090] The predetermined scanning order is an order for quantizing / dequantizing transform coefficients, for example, the predetermined scanning order is defined as an ascending order (low frequency to high frequency) or a descending order (high frequency to low frequency).

[0091] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. In other words, as the value of the quantization parameter increases, the quantization error also increases.

[0092] [Entropy coding section] The entropy coding unit 110 generates a coded signal (coded bitstream) based on the quantized coefficients input from the quantization unit 108. Specifically, for example, the entropy coding unit 110 binarizes the quantized coefficients, arithmetically codes the binary signal, and outputs a compressed bitstream or sequence.

[0093] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. The inverse quantization unit 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0094] [Inverse conversion section] The inverse transform unit 114 restores prediction errors (residuals) by inverse transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores prediction errors of the current block by performing an inverse transform on the transform coefficients corresponding to the transform performed by the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction errors to the adder unit 116.

[0095] Note that the restored prediction error usually loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. In other words, the restored prediction error usually contains a quantization error.

[0096] [Adder] The adder 116 reconstructs a current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0097] [Block Memory] The block memory 118 is a storage unit for storing blocks that are referenced in intra prediction and are in the picture to be coded (referred to as the "current picture"). Specifically, the block memory 118 stores the reconstructed blocks output from the adder 116.

[0098] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0099] ALF applies a least squares error filter to remove coding artifacts, for example, for each 2x2 sub-block in the current block, one filter selected from multiple filters based on local gradient direction and activity.

[0100] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The sub-blocks are classified based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, the sub-blocks are classified into a plurality of classes based on the classification value C.

[0101] The gradient direction value D is derived by, for example, comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived by, for example, adding gradients in multiple directions and quantizing the sum.

[0102] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0103] The filter shape used in ALF is, for example, a circularly symmetric shape. FIGS. 4A to 4C are diagrams showing several examples of filter shapes used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the filter shape is usually signaled at the picture level. Note that signaling of the information indicating the filter shape does not need to be limited to the picture level, and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0104] Whether to turn on or off ALF may be determined at the picture level or the CU level. For example, whether to apply ALF for luma may be determined at the CU level, and whether to apply ALF for chroma may be determined at the picture level. Information indicating whether ALF is on or off is usually signaled at the picture level or the CU level. Note that signaling of information indicating whether ALF is on or off does not need to be limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0105] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level, although the signaling of coefficient sets need not be limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0106] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may also be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed blocks filtered by the loop filter unit 120.

[0107] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also referred to as intra-picture prediction) of the current block with reference to blocks in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates the intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0108] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes typically includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0109] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.

[0110] The plurality of directional prediction modes includes, for example, 33 prediction modes defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may also include 32 prediction modes in addition to the 33 directions (65 directional prediction modes in total).

[0111] 5A is a conceptual diagram showing all 67 intra-prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra-prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the two "non-directional" prediction modes are not shown in FIG. 5A).

[0112] In various processing examples, a luminance block may be referenced in intra prediction of a chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. This intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of the chrominance block that references such a luminance block (e.g., called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0113] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on gradients of reference pixels in the horizontal / vertical directions. Intra prediction involving such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is usually signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0114] [Inter prediction section] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also referred to as inter prediction) on the current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or current sub-block to find a reference block or sub-block within the reference picture that most closely matches the current block or sub-block. The inter prediction unit 126 then obtains motion information (e.g., a motion vector) that compensates for (or predicts) the movement or change from the reference block or sub-block to the current block or sub-block. The inter prediction unit 126 then performs motion compensation (or motion prediction) based on the motion information to generate an inter prediction signal for the current block or sub-block. The inter prediction unit 126 then outputs the generated inter prediction signal to the prediction control unit 128.

[0115] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms, such as a motion vector, or as a difference between a motion vector and a predicted motion vector.

[0116] Note that an inter-prediction signal may be generated using not only the motion information of the current block obtained by motion estimation but also the motion information of adjacent blocks. Specifically, an inter-prediction signal may be generated for each sub-block in the current block by weighting and adding a prediction signal based on the motion information obtained by motion estimation (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture). Such inter-prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0117] In the OBMC mode, information indicating the size of a sub-block for OBMC (e.g., referred to as an OBMC block size) may be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as an OBMC flag) may be signaled at the CU level. Note that the signaling level of this information is not limited to the sequence level and the CU level, and may be at another level (e.g., the picture level, slice level, tile level, CTU level, or sub-block level).

[0118] The OBMC mode will now be described in more detail. Figures 5B and 5C are a flowchart and a conceptual diagram illustrating the predicted image correction process using the OBMC process.

[0119] Referring to Figure 5C, first, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to the current block to be coded. In Figure 5C, the arrow "MV" points to a reference picture, indicating what the current block in the current picture references to obtain the predicted image.

[0120] Next, the motion vector (MV_L) already derived for the coded left neighboring block is applied (reused) to the current block to obtain a predicted image (Pred_L). The motion vector (MV_L) is indicated by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by overlapping the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between the neighboring blocks.

[0121] Similarly, a motion vector (MV_U) already derived for the coded upper adjacent block is applied (reused) to the current block to be coded to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. Then, a second correction of the predicted image is performed by superimposing the predicted image Pred_U on the predicted images (i.e., Pred and Pred_L) that have been corrected the first time. This has the effect of blending the boundaries between adjacent blocks, in one aspect. The predicted image obtained by the second correction is the final predicted image of the current block, in which the boundaries with the adjacent blocks have been blended (smoothed).

[0122] Although a two-stage correction method using the left adjacent block and the above adjacent block has been described here, it is also possible to configure a method in which correction is performed more than two times using the right adjacent block or the below adjacent block.

[0123] The area to be superimposed does not have to be the pixel area of ​​the entire block, but may be only a part of the area near the block boundary.

[0124] Here, the OBMC predicted image correction process has been described, in which additional predicted images Pred_L and Pred_U are superimposed based on one reference picture to obtain one predicted image Pred. However, when a predicted image is corrected based on multiple reference pictures, the same process may be applied to each of the multiple reference pictures. In such a case, the OBMC image correction based on multiple reference pictures is performed to obtain a corrected predicted image from each reference picture, and then the obtained multiple corrected predicted images are further superimposed to obtain a final predicted image.

[0125] In OBMC, the unit of the current block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.

[0126] As a method for determining whether to apply OBMC processing, for example, there is a method using obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, an encoding device determines whether a block to be encoded belongs to an area with complex motion, and if it belongs to an area with complex motion, sets the value "1" as obmc_flag and performs encoding by applying OBMC processing, and if it does not belong to an area with complex motion, sets the value "0" as obmc_flag and performs encoding without applying OBMC processing. On the other hand, a decoding device decodes obmc_flag described in a stream (i.e., a compressed sequence), and switches whether to apply OBMC processing depending on the value, and performs decoding.

[0127] The motion information may be derived on the decoding device side without being signaled from the encoding device side. For example, a merge mode defined in the H.265 / HEVC standard may be used. Alternatively, the motion information may be derived by performing motion estimation on the decoding device side. In this case, the motion estimation may be performed on the decoding device side without using pixel values ​​of the current block.

[0128] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.

[0129] An example of the FRUC process is shown in Figure 5D. First, a list of multiple candidates (which may be the same as the merge list) each having a predicted motion vector (MV) is generated by referring to the motion vectors of coded blocks spatially or temporally adjacent to the current block. Next, a best candidate MV is selected from the multiple candidate MVs registered in the candidate list. For example, an evaluation value of each candidate MV included in the candidate list is calculated, and one candidate MV is selected based on the evaluation value.

[0130] Then, a motion vector for the current block is derived based on the motion vector of the selected candidate. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as the motion vector for the current block. Also, for example, the motion vector for the current block may be derived by performing pattern matching in a peripheral area of ​​a position in the reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed using pattern matching in the reference picture and an evaluation value for the peripheral area of ​​the best candidate MV, and if an MV with a better evaluation value is found, the best candidate MV may be updated to that MV, and this may be used as the final MV for the current block. It is also possible to configure the system without performing the process of updating to an MV with a better evaluation value.

[0131] The same processing may be performed when processing is performed in sub-block units.

[0132] The evaluation value may be calculated by various methods. For example, a reconstructed image of an area in a reference picture corresponding to the motion vector may be compared with a reconstructed image of a predetermined area (for example, as described later, this may be an area of ​​another reference picture or an area of ​​an adjacent block in the current picture), and the difference in pixel values ​​between the two reconstructed images may be calculated and used as the evaluation value of the motion vector. The evaluation value may also be calculated using other information in addition to the difference value.

[0133] Next, an example of pattern matching will be described in detail. First, one candidate MV included in a candidate MV list (e.g., a merge list) is selected as a starting point for search by pattern matching. For example, first pattern matching or second pattern matching can be used as pattern matching. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0134] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture that is along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the candidate for an area in a reference picture.

[0135] 6 is a diagram illustrating an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of blocks that best match among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of a current block (Cur block). Specifically, for the current block, a difference is derived between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by a candidate MV and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV by the display time interval, and an evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, which can lead to good results.

[0136] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (CurPic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures temporally and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives two mirror-symmetric bidirectional motion vectors.

[0137] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0138] 7 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 7, in the second pattern matching, a motion vector of a current block is derived by searching a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in the current picture (Cur Pic). Specifically, a difference is derived between a reconstructed image of both or either of the coded areas adjacent to the left and / or above the current block and a reconstructed image at the same position in the coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and the candidate MV with the best evaluation value among the multiple candidate MVs can be selected as the best candidate MV.

[0139] Information indicating whether such a FRUC mode is applied (e.g., referred to as an FRUC flag) may be signaled at the CU level. Furthermore, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating a pattern matching method (e.g., first pattern matching or second pattern matching) (e.g., referred to as an FRUC mode flag) may be signaled at the CU level. Note that signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0140] Next, we will explain how to derive a motion vector. First, we will explain a mode in which a motion vector is derived based on a model that assumes uniform linear motion. This mode is sometimes called BIO (bi-directional optical flow) mode.

[0141] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x ,v y ) denotes a velocity vector, and τ0 and τ1 denote the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0 and Ref1), respectively. (MVx0,MVy0) denotes a motion vector corresponding to reference picture Ref0, and (MVx1,MVy1) denotes a motion vector corresponding to reference picture Ref1.

[0142] At this time, the velocity vector (v x ,v y ), (MVx0,MVy0) and (MVx1,MVy1) are respectively (v x τ0,v y τ0) and (-v x τ1,-v y τ1), and the following optical flow equation (1) holds:

[0143]

number

[0144] Here, I (k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like may be corrected pixel by pixel.

[0145] Note that the decoding device may derive motion vectors using a method other than that based on a model assuming constant-velocity linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0146] Next, a mode in which a motion vector is derived for each sub-block based on the motion vectors of multiple neighboring blocks will be described. This mode is sometimes called an affine motion compensation prediction mode.

[0147] FIG. 9A is a diagram for explaining the derivation of a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v x ,v y ) is derived.

[0148]

number

[0149] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting coefficient.

[0150] The affine motion compensation prediction mode may include several modes in which the methods of deriving the motion vectors of the upper-left and upper-right corner control points are different. Information indicating the affine motion compensation prediction mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating the affine motion compensation prediction mode does not need to be limited to the CU level, and may be at other levels (e.g., the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0151] [Predictive control unit] The prediction control unit 128 selects either an intra-prediction signal (a signal output from the intra-prediction unit 124) or an inter-prediction signal (a signal output from the inter-prediction unit 126), and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0152] As shown in FIG. 1 , in various processing examples, the prediction control unit 128 may output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bitstream and perform the same prediction process as that performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or any index, flag, or value based on or indicating the prediction process performed in the intra predictor 124, the inter predictor 126, and the prediction control unit 128.

[0153] FIG. 9B shows an example of a process for deriving a motion vector for the current picture in merge mode.

[0154] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include spatially adjacent prediction MVs, which are MVs held by multiple coded blocks spatially located around the target block, temporally adjacent prediction MVs, which are MVs held by blocks in the vicinity of the target block projected onto the coded reference picture, joint prediction MVs, which are MVs generated by combining the MV values ​​of the spatially adjacent prediction MVs and the temporally adjacent prediction MVs, and zero prediction MVs, which are MVs with a value of zero.

[0155] Next, one predicted MV is selected from the plurality of predicted MVs registered in the predicted MV list, and is determined as the MV for the target block.

[0156] Furthermore, the variable length coding unit encodes merge_idx, which is a signal indicating which predicted MV has been selected, into the stream.

[0157] Note that the predicted MVs registered in the predicted MV list described in Figure 9B are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include predicted MVs other than the types of predicted MVs shown in the figure.

[0158] The final MV may be determined by performing a DMVR (decoder motion vector refinement) process, which will be described later, using the MV of the current block derived in the merge mode.

[0159] FIG. 9C is a conceptual diagram illustrating an example of DMVR processing for determining an MV.

[0160] First, the optimal MVP set for the current block (for example, in merge mode) is set as the candidate MV. Then, reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 direction, according to the candidate MV (L0). Similarly, reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 direction, according to the candidate MV (L1). A template is generated by averaging these reference pixels.

[0161] Next, the template is used to search the surrounding areas of the candidate MVs in the first reference picture (L0) and the second reference picture (L1), and the MV with the smallest cost is determined as the final MV. Note that the cost value may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search area, the candidate MV value, etc.

[0162] Typically, the encoding device and the decoding device described below basically have the same configuration and operation for the processing described here.

[0163] Any processing may be used, not limited to the processing example described here, as long as it is a processing that can search around the candidate MVs and derive the final MV.

[0164] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.

[0165] FIG. 9D is a conceptual diagram for explaining an example of a predicted image generation method using luminance correction processing by LIC processing.

[0166] First, the MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.

[0167] Next, information indicating how the luminance values ​​of the current block have changed between the reference picture and the current picture is extracted. This extraction is performed based on the luminance pixel values ​​of the coded left-adjacent reference area (peripheral reference area) and the coded upper-adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the derived MV. Then, the information indicating how the luminance values ​​have changed is used to calculate luminance correction parameters.

[0168] A predicted image for the current block is generated by performing luminance correction processing that applies the luminance correction parameters to a reference image in a reference picture specified by the MV.

[0169] The shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.

[0170] Although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures, and the reference images obtained from each reference picture may be subjected to brightness correction processing in the same manner as described above before generating a predicted image.

[0171] One method for determining whether to apply LIC processing is to use lic_flag, which is a signal indicating whether to apply LIC processing. As a specific example, an encoding device determines whether the current block belongs to an area where a luminance change occurs, and if the current block belongs to an area where a luminance change occurs, sets the value of lic_flag to "1" and performs encoding by applying LIC processing, and if the current block does not belong to an area where a luminance change occurs, sets the value of lic_flag to "0" and performs encoding without applying LIC processing. On the other hand, a decoding device may decode lic_flag described in the stream, and switch whether to apply LIC processing depending on the value and perform decoding.

[0172] Another method for determining whether to apply LIC processing is to determine whether LIC processing has been applied to neighboring blocks.As a specific example, when the current block is in merge mode, it is determined whether the neighboring coded blocks selected when deriving MV in merge mode processing have been coded using LIC processing.Depending on the result, whether to apply LIC processing is switched and coding is performed.In this example, the same processing is also applied to the processing on the decoding device side.

[0173] [Overview of the decoding device] Next, an overview will be given of a decoding device capable of decoding the coded signal (coded bitstream) output from the above coding device 100. Fig. 10 is a block diagram showing the functional configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a video decoding device that decodes video on a block-by-block basis.

[0174] As shown in FIG. 10, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0175] The decoding device 200 is realized by, for example, a general-purpose processor and memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Alternatively, the decoding device 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0176] Each component included in the decoding device 200 will be described below.

[0177] [Entropy Decoding] The entropy decoding unit 202 entropy-decodes the coded bitstream. Specifically, for example, the entropy decoding unit 202 arithmetically decodes the coded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. The entropy decoding unit 202 outputs quantized coefficients to the inverse quantization unit 204 on a block-by-block basis. The entropy decoding unit 202 may output prediction parameters included in the coded bitstream (see FIG. 1 ) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the coding device side.

[0178] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of the block to be decoded (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. The inverse quantization unit 204 then outputs the inverse quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0179] [Inverse conversion section] The inverse transform unit 206 restores prediction errors (residuals) by inverse transforming the transform coefficients input from the inverse quantization unit 204.

[0180] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0181] Also, for example, if the information decoded from the coded bitstream indicates that NSST is to be applied, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0182] [Adder] The adder 208 reconstructs the current block by adding the prediction error input from the inverse transformer 206 and the prediction sample input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0183] [Block Memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter referred to as a current picture). Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208.

[0184] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0185] If the information indicating ALF on / off read from the encoded bitstream indicates that ALF is on, one filter is selected from multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0186] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0187] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction based on the intra prediction mode interpreted from the encoded bitstream, by referring to blocks in the current picture stored in the block memory 210. Specifically, the intra prediction unit 216 generates the intra prediction signal by performing intra prediction by referring to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0188] Note that when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0189] Furthermore, when information interpreted from the coded bitstream (for example, prediction parameters output from the entropy decoding unit 202) indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical directions.

[0190] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the coded bitstream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.

[0191] If the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion estimation, but also the motion information of neighboring blocks.

[0192] Furthermore, if the information interpreted from the coded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the coded bitstream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0193] Furthermore, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming constant-velocity linear motion. Furthermore, when information interpreted from the coded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks.

[0194] [Predictive control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as a prediction signal to the adder 208. Overall, the configurations, functions, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side may correspond to the configurations, functions, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.

[0195] [Non-rectangular division] In the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126 of the encoding device (see FIG. 1 ), as well as in the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218 of the decoding device (see FIG. 10 ), the partitions (or variable-size blocks or sub-blocks) obtained from the division of each block and from which motion information (e.g., motion vectors) is obtained are conventionally always rectangular, as shown in FIG. 2 . The inventors have discovered that generating partitions having non-rectangular shapes, such as triangular shapes, can lead to improved image quality and coding efficiency in various implementations, depending on the image content of a picture. Hereinafter, various embodiments will be described in which at least one partition divided from an image block for prediction purposes has a non-rectangular shape. Note that these embodiments are equally applicable to the encoding device side (prediction control unit 128 connected to intra prediction unit 124 and inter prediction unit 126) and to the decoding device side (prediction control unit 220 connected to intra prediction unit 216 and inter prediction unit 218), and may be implemented in the encoding device of Figure 1 or the decoding device of Figure 10.

[0196] FIG. 11 is a flowchart illustrating an example of a process for dividing an image block into multiple partitions including at least a first partition and a second partition having a non-rectangular shape (e.g., a triangle), and then encoding (or decoding) the image block as a reconstructed combination of the first and second partitions.

[0197] In step S1001, an image block is divided into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, which may or may not have a non-rectangular shape. For example, as shown in FIG. 12, the image block may be divided from the upper left corner of the image block to the lower right corner of the image block to create a first partition and a second partition, both of which have a non-rectangular shape (e.g., a triangle). Alternatively, the image block may be divided from the upper right corner of the image block to the lower left corner of the image block to create a first partition and a second partition, both of which have a non-rectangular shape (e.g., a triangle). Various examples of non-rectangular divisions are described below with reference to FIG. 12 and FIGS. 17-19.

[0198] In step S1002, the process predicts a first motion vector for the first partition and a second motion vector for the second partition. For example, predicting the first and second motion vectors may include selecting a first motion vector from a first set of motion vector candidates and selecting a second motion vector from a second set of motion vector candidates.

[0199] In step S1003, a motion compensation process is performed to obtain a first partition using the first motion vector derived in step S1002 above, and to obtain a second partition using the second motion vector derived in step S1002 above.

[0200] In step S1004, a prediction process is performed on the image block as a (reconstructed) combination of the first partition and the second partition. The prediction process includes a boundary smoothing process to smooth the boundary between the first partition and the second partition. For example, the boundary smoothing process involves weighting a plurality of first values ​​of a plurality of boundary pixels predicted based on the first partition and a plurality of second values ​​of a plurality of boundary pixels predicted based on the second partition. Various implementations of the boundary smoothing process are described later with reference to Figures 13, 14, 20, and 21A-21D.

[0201] In step S1005, the process encodes or decodes the image block using one or more parameters including a partition parameter indicating division of the image block into a first partition having a non-rectangular shape and a second partition. As summarized in the table of Figure 15, for example, the partition parameter ("first index value") may encode, for example, a division direction to be applied to the division (e.g., from upper left to lower right or from upper right to lower left as shown in Figure 12) together with the first and second motion vectors derived in step S1002 described above. Details of such partition syntax operations involving one or more parameters including the partition parameter will be described in detail later with reference to Figures 15, 16, and 22-25.

[0202] FIG. 17 is a flowchart showing a process 2000 for dividing an image block. In step S2001, the process divides an image into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. As shown in FIG. 12, the image block is divided into a first partition having a triangular shape and a second partition also having a triangular shape. There are many other examples in which an image block is divided into a plurality of partitions, including a first partition and a second partition, where at least the first partition has a non-rectangular shape. The non-rectangular shape may be a triangle, a trapezoid, or a polygon having at least five sides and corners.

[0203] For example, as shown in Figure 18, an image block may be divided into two triangular partitions. An image block may also be divided into more than two triangular partitions (e.g., three triangular partitions). An image block may be divided into a combination of one or more triangular partitions and one or more rectangular partitions. Alternatively, an image block may be divided into a combination of one or more triangular partitions and one or more polygonal partitions.

[0204] 19, an image block may be divided into L-shaped (polygonal) partitions and rectangular-shaped partitions. An image block may be divided into pentagonal (polygonal) partitions and triangular-shaped partitions. An image block may be divided into hexagonal (polygonal) partitions and pentagonal (polygonal) partitions. Alternatively, an image block may be divided into multiple polygonal-shaped partitions.

[0205] Referring back to FIG. 17 , in step S2002, the process predicts a first motion vector for a first partition, e.g., by selecting a first partition from a first motion vector candidate set, and predicts a second motion vector for a second partition, e.g., by selecting a second partition from a second motion vector candidate set. For example, the first motion vector candidate set may include motion vectors of partitions adjacent to the first partition, and the second motion vector candidate set may include motion vectors of partitions adjacent to the second partition. The adjacent partitions may be one or both of spatially adjacent partitions and temporally adjacent partitions. Some examples of spatially adjacent partitions include partitions located to the left, lower-left, lower, lower-right, right, upper-right, upper, or upper-left of the partition being processed. Examples of temporally adjacent partitions include co-located partitions in reference pictures of the image block.

[0206] In various implementations, the partitions adjacent to the first partition and the partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The first motion vector candidate set may be the same as or different from the second motion vector candidate set. Furthermore, at least one of the first motion vector candidate set and the second motion vector candidate set may be the same as another third motion vector candidate set prepared for the image block.

[0207] In some implementations, in response to determining in step S2002 that the second partition, like the first partition, has a non-rectangular shape (e.g., a triangle), process 2000 creates (for the non-rectangular-shaped second partition) a second motion vector candidate set that includes multiple motion vectors of multiple partitions adjacent to the second partition, excluding the first partition (i.e., excluding the motion vector of the first partition).On the other hand, in response to determining that the second partition, not like the first partition, has a rectangular shape, process 2000 creates (for the rectangular-shaped second partition) a second motion vector candidate set that includes multiple motion vectors of multiple partitions adjacent to the second partition, including the first partition.

[0208] In step S2003, the process encodes or decodes the first partition using the first motion vector derived in step S2002 described above, and encodes or decodes the second partition using the second motion vector derived in step S2002 described above.

[0209] An image block division process such as process 2000 in Figure 17 may be performed by an image encoding device including, for example, a circuit such as that shown in Figure 1 and a memory connected to the circuit. The circuit operates to divide an image block into a plurality of partitions, including a first partition and a second partition, each having a non-rectangular shape (step S2001), predict a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encode the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0210] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a partitioning unit 102 that receives and divides an original image into a plurality of blocks in operation, an adder 104 that receives and divides the plurality of blocks from the partitioning unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a transform unit 106 that performs a transform on the plurality of residuals output from the adder 104 in operation, and outputs a plurality of transform coefficients, a quantization unit 108 that quantizes the plurality of transform coefficients in operation, to generate a plurality of quantized transform coefficients, an entropy coding unit 110 that encodes the plurality of quantized transform coefficients in operation, and generates a bitstream, an inter prediction unit 126 that generates and divides a prediction of a current block based on a reference block in a previously coded reference picture in operation, an intra prediction unit 124 that generates and divides a prediction of the current block based on a previously coded reference block in the current picture in operation, and a prediction control unit 128 connected to the memories 118 and 122. In operation, the prediction control unit 128 divides multiple blocks into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0211] According to another embodiment, there is provided an image decoding device including a circuit and a memory connected to the circuit, for example as shown in Figure 10. The circuit operates to divide an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape (Figure 17, step S2001), predict a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decode the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0212] Further according to an embodiment, the image decoding device shown in FIG. 10 is provided including: an entropy decoding unit 202 that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients; an inverse quantization unit 204 and an inverse transform unit 206 that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an adder 208 that, in operation, adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks; an inter prediction unit 218 that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit 216 that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and a prediction control unit 220 connected to the memories 210, 214. In operation, the prediction control unit 220 divides an image block into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0213] [Boundary smoothing] As described above in FIG. 11 , step S1004, according to various embodiments, performing prediction processing on an image block as a (reconstructed) combination of a first partition having a non-rectangular shape and a second partition may involve applying a boundary smoothing process along the boundary between the first partition and the second partition.

[0214] For example, Figure 21B shows an example of a boundary smoothing process that involves weighting a plurality of first values ​​of a plurality of boundary pixels that are first predicted based on a first partition and a plurality of second values ​​of a plurality of boundary pixels that are second predicted based on a second partition.

[0215] 20 is a flowchart illustrating an overall boundary smoothing process 3000 that involves weighting a plurality of first values ​​of a plurality of boundary pixels that are first predicted based on a first partition and a plurality of second values ​​of a plurality of boundary pixels that are second predicted based on a second partition, according to one embodiment. In step S3001, an image block is divided along its boundary into a first partition and a second partition, where at least the first partition has a non-rectangular shape, as shown in FIG. 21A or the above-described FIGS. 12, 18, and 19.

[0216] In step S3002, a plurality of first values ​​(e.g., color, luminance, transparency, etc.) of a set of pixels in a first partition along the boundary ("boundary pixels" in FIG. 21A) are first predicted using information about the first partition. In step S3003, a plurality of second values ​​of a set of pixels in the (same) first partition along the boundary are second predicted using information about the second partition. In some embodiments, at least one of the first prediction and the second prediction is an inter-prediction process that predicts the plurality of first values ​​and the plurality of second values ​​based on a reference partition in a coded reference picture. With reference to FIG. 21D, in some implementations, the prediction process predicts a plurality of first values ​​of all pixels in a first partition ("first sample set") that includes a set of pixels where the first and second partitions overlap, and predicts second values ​​only for a set of pixels where the first and second partitions overlap ("second sample set"). In other implementations, at least one of the first prediction and the second prediction is an intra-prediction process that predicts the first values ​​and the second values ​​based on a coded reference partition in the current picture. In some implementations, the prediction method used for the first prediction is different from the prediction method used for the second prediction. For example, the first prediction may include an inter-prediction process, and the second prediction may include an intra-prediction process. Information used for the first prediction of the first values ​​or the second prediction of the second values ​​may include motion vectors of the first partition or the second partition, multiple intra-prediction directions, etc.

[0217] In step S3004, the first values ​​predicted using the first partition and the second values ​​predicted using the second partition are weighted, and in step S3005, the first partition is encoded or decoded using the weighted first values ​​and second values.

[0218] 21B shows an example of a boundary smoothing operation in which the first partition and the second partition overlap by (a maximum of) five pixels in each row. That is, the number of pixel sets in each row or column for which the plurality of first values ​​are predicted based on the first partition and the plurality of second values ​​are predicted based on the second partition is at most five. FIG. 21C shows another example of a boundary smoothing operation in which the first partition and the second partition overlap by (a maximum of) three pixels in each row or column. That is, the number of pixel sets in each row or column for which the plurality of first values ​​are predicted based on the first partition and the plurality of second values ​​are predicted based on the second partition is at most three.

[0219] 13 shows another example of a boundary smoothing operation in which the first partition and the second partition overlap at (maximum) four pixels in each row or column. That is, the number of pixel sets in each row or column for which first values ​​are predicted based on the first partition and second values ​​are predicted based on the second partition is at most four. In the example shown, weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the first values ​​of the four pixels in the set, respectively, and weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the second values ​​of the four pixels in the set, respectively.

[0220] 14 further illustrates several examples of boundary smoothing operations in which the first and second partitions overlap by 0 pixels in each row or column (i.e., they do not overlap), overlap by (at most) 1 pixel in each row or column, and overlap by (at most) 2 pixels in each row or column. In examples in which the first and second partitions do not overlap, zero weights are applied. In examples in which the first and second partitions overlap by 1 pixel in each row or column, a weight of 1 / 2 may be applied to the first values ​​of the pixels in the set predicted based on the first partition, and a weight of 1 / 2 may be applied to the second values ​​of the pixels in the set predicted based on the second partition. In an example where the first and second partitions overlap by two pixels in each row or column, weights of 1 / 3 and 2 / 3 may be applied to the first values ​​of two pixels in the set predicted based on the first partition, respectively, and weights of 2 / 3 and 1 / 3 may be applied to the second values ​​of two pixels in the set predicted based on the second partition, respectively.

[0221] According to some embodiments described above, the number of pixels in the set where the first and second partitions overlap is an integer. In other implementations, the number of overlapping pixels in the set may be, for example, a non-integer or fractional. The weights applied to the first and second values ​​of the pixel set may also be fractional or integer, depending on the application.

[0222] A boundary smoothing process such as process 3000 in Figure 20 may be performed by an image encoding device including, for example, a circuit such as that shown in Figure 1 and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition and a second partition, each partition having a non-rectangular shape, which are divided from an image block (Figure 20, step S3001). The boundary smoothing operation includes predicting a plurality of first values ​​of a set of pixels in the first partition along the boundary using information about the first partition (step S3002), predicting a plurality of second values ​​of the set of pixels in the first partition along the boundary using information about the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0223] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a partitioning unit 102 that receives and divides an original image into a plurality of blocks in operation, an adder 104 that receives and divides the plurality of blocks from the partitioning unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a transform unit 106 that performs a transform on the plurality of residuals output from the adder 104 in operation, and outputs a plurality of transform coefficients, a quantization unit 108 that quantizes the plurality of transform coefficients in operation, to generate a plurality of quantized transform coefficients, an entropy coding unit 110 that encodes the plurality of quantized transform coefficients in operation, and generates a bitstream, an inter prediction unit 126 that generates and divides a prediction of a current block based on a reference block in a previously coded reference picture in operation, an intra prediction unit 124 that generates and divides a prediction of the current block based on a previously coded reference block in the current picture in operation, and a prediction control unit 128 connected to the memories 118 and 122. In operation, the prediction control unit 128 performs a boundary smoothing operation along the boundary between a first partition and a second partition, the first partition having a non-rectangular shape and divided from the image block (FIG. 20, step S3001). The boundary smoothing operation includes: first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information about the first partition (step S3002); second predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information about the second partition (step S3003); weighting the plurality of first values ​​and the plurality of second values ​​(step S3004); and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0224] According to another embodiment, there is provided an image decoding device including a circuit and a memory connected to the circuit, for example as shown in Fig. 10. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape divided from an image block and a second partition (Fig. 20, step S3001). The boundary smoothing operation includes: first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information of the first partition (step S3002); second predicting a plurality of second values ​​of the pixel set of the first partition along the boundary using information of the second partition (step S3003); weighting the plurality of first values ​​and the plurality of second values ​​(step S3004); and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0225] According to another embodiment, an image decoding apparatus shown in FIG. 10 is provided, including: an entropy decoding unit 202 that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients in operation; an inverse quantization unit 204 and an inverse transform unit 206 that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an adder 208 that, in operation, adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks; an inter prediction unit 218 that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit 216 that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and a prediction control unit 220 connected to the memories 210, 214. In operation, prediction control unit 220 performs a boundary smoothing operation along the boundary between a first partition and a second partition, each partition having a non-rectangular shape and divided from an image block (FIG. 20, step S3001). The boundary smoothing operation includes: first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information about the first partition (step S3002); second predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information about the second partition (step S3003); weighting the plurality of first values ​​and the plurality of second values ​​(step S3004); and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0226] Entropy coding and decoding using partition parameter syntax As shown in Figure 11, step S1005, according to various embodiments, an image block divided into a first partition and a second partition having a non-rectangular shape may be encoded or decoded using one or more parameters, including a partition parameter indicating the non-rectangular division of the image block. In various embodiments, such partition parameter may jointly encode, for example, a division direction applied to the division (e.g., from upper left to lower right or from upper right to lower left, see Figure 12) and the first and second motion vectors predicted in step S1002, as described more fully below.

[0227] 15 is a table showing sample partition parameters ("first index values") and information sets that are jointly encoded by the partition parameters. The partition parameters ("first index values") range from 0 to 6 and jointly encode the direction of dividing an image block into a first partition and a second partition, both of which are triangular (see FIG. 12), the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002). In particular, partition parameter 0 encodes that the division direction is from the upper left corner to the lower right corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0228] Partition parameter 1 encodes that the division direction is from the upper right corner to the lower left corner, that the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the division direction is from the upper right corner to the lower left corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the division direction is from the upper left corner to the lower right corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the partition direction is from the upper right corner to the lower left corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the partition direction is from the upper left corner to the lower right corner, that the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 6 encodes that the partition direction is from the upper left corner to the lower right corner, that the first motion vector is the "fourth" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0229] FIG. 22 is a flowchart showing a method 4000 performed on the encoding device side. In step S4001, the process divides an image block into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape, based on a partition parameter indicating the division. For example, as shown in FIG. 15 above, the partition parameter may indicate a direction in which the image block is divided (e.g., from the upper right corner to the lower left corner or from the upper left corner to the lower right corner). In step S4002, the process encodes the first partition and the second partition. In step S4003, the process writes one or more parameters including the partition parameter into a bitstream that can be received and decoded by the decoding device side to perform the same prediction process (as performed on the encoding device side) on the first and second partitions on the decoding device side. The one or more parameters including the partition parameter jointly or separately encode various information, such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used to divide the image block to obtain the first and second partitions, the first motion vector of the first partition, the second motion vector of the second partition, etc.

[0230] FIG. 23 is a flowchart showing a method 5000 performed on the decoding device side. In step S5001, the process reads one or more parameters from a bitstream, the one or more parameters including partition parameters indicating division of an image block into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape. The one or more parameters including the partition parameters read from the bitstream may encode various information required for the decoding device side to perform the same prediction process as that performed on the encoding device side, such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used to divide the image block to obtain the first and second partitions, a first motion vector for the first partition, a second motion vector for the second partition, etc. In step S5002, the process 5000 divides the image block into multiple partitions based on the partition parameters read from the bitstream. In step S5003, the process decodes the first and second partitions divided from the image block.

[0231] 24 is a table of sample partition parameters ("first index values") and a plurality of information sets respectively coded as a whole by a plurality of partition parameters, similar in characteristics to the sample table described above in FIG. 15. In FIG. 24, the partition parameters ("first index values") range from 0 to 6 and jointly code the shapes of the first and second partitions divided from the image block, the direction in which the image block is divided into the first and second partitions, the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002). In particular, a partition parameter of 0 codes that neither the first nor the second partition has a triangular shape, and therefore the partition direction information is "N / A," the first motion vector information is "N / A," and the second motion vector information is "N / A."

[0232] Partition parameter 1 encodes that the first partition and the second partition are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the first partition and the second partition are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the first and second partitions are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the first and second partitions are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the first and second partitions are triangular, the division direction is from the top right corner to the bottom left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition.Partition parameter 6 encodes that the first and second partitions are triangular, the division direction is from the top left corner to the bottom right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0233] According to some implementations, the partition parameters (index values) may be binarized according to a binarization scheme selected depending on the value of at least one or more parameters. Figure 16 shows an example of a binarization scheme for binarizing the index values ​​(partition parameter values).

[0234] 25 is a table showing example combinations of first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is to be divided into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape. In this example, the partition parameter may be used to indicate that the image block is to be divided without integrally encoding other information encoded by one or more of the other parameters.

[0235] 25, the first parameter is used to indicate the image block size, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used to indicate, for example, 1) that there is no triangular-shaped partition when the image block size is larger than 64x64, or 2) that there is no triangular-shaped partition when the ratio of the width to height of the image block is larger than 4 (e.g., 64x4).

[0236] 25, the first parameter is used to indicate a prediction mode, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, to indicate that there is no triangular partition when the image block is coded in intra mode.

[0237] In the third example of Figure 25, the first parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape, and the second parameter is used to indicate the prediction mode. Such a combination of the first parameter and the second parameter may be used, for example, to indicate that the image block should be inter-coded if at least one of the partitions divided from the image block has a triangular shape.

[0238] 25, the first parameter indicates the motion vector of the neighboring block, and the second parameter is used as a partition parameter indicating the direction of dividing the image block into two triangles. Such a combination of the first and second parameters may be used to indicate, for example, 1) if the motion vector of the neighboring block is diagonal, that the direction of dividing the image block into two triangles is from the upper left corner to the lower right corner.

[0239] In the fifth example of Figure 25, the first parameter indicates the intra-prediction direction of the neighboring block, and the second parameter is used as a partition parameter indicating the direction of dividing the image block into two triangles. Such a combination of the first parameter and the second parameter may be used to indicate, for example, 1) if the intra-prediction direction of the neighboring block is the reverse diagonal direction, the direction of dividing the image block into two triangles is from the upper right corner to the lower left corner.

[0240] The tables of one or more parameters including the partition parameters, and which information is coded together or separately as shown in Figures 15, 24, and 25, are presented as examples only, and it should be understood that numerous other ways of coding various information together or separately as part of the partition syntax operations described above are within the scope of this disclosure. For example, the partition parameters may indicate that the first partition is a triangle, a trapezoid, or a polygon having at least five sides and corners. The partition parameters may indicate that the second partition has a non-rectangular shape, such as a triangle, a trapezoid, or a polygon having at least five sides and corners. The partition parameters may indicate one or more pieces of information about the partition, such as the non-rectangular shape of the first partition, the shape of the second partition (which may be non-rectangular or rectangular), and the division direction (e.g., from the upper left corner of the image block to its lower right corner, and from the upper right corner of the image block to its lower left corner) applied to divide the image block into multiple partitions. The partition parameters jointly encode further information, such as a first motion vector of the first partition, a second motion vector of the second partition, an image block size, a prediction mode, a motion vector of a neighboring block, an intra-prediction direction of the neighboring block, etc. Alternatively, any of the further information may be separately encoded by one or more parameters other than the partition parameters.

[0241] A partition syntax operation such as process 4000 in Figure 22 may be performed by an image coding device including a circuit and a memory connected to the circuit, such as that shown in Figure 1. In operation, the circuit performs the partition syntax operation including dividing an image block into multiple partitions, including a first partition and a second partition, each having a non-rectangular shape, based on partition parameters indicating the division (Figure 22, step S4001), encoding the first partition and the second partition (S4002), and writing one or more parameters including the partition parameters to a bitstream (S4003).

[0242] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a partitioning unit 102 that receives and divides an original image into a plurality of blocks in operation, an adder 104 that receives and divides the plurality of blocks from the partitioning unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a transform unit 106 that performs a transform on the plurality of residuals output from the adder 104 in operation, and outputs a plurality of transform coefficients, a quantization unit 108 that quantizes the plurality of transform coefficients in operation, to generate a plurality of quantized transform coefficients, an entropy coding unit 110 that encodes the plurality of quantized transform coefficients in operation, and generates a bitstream, an inter prediction unit 126 that generates and divides a prediction of a current block based on a reference block in a previously coded reference picture in operation, an intra prediction unit 124 that generates and divides a prediction of the current block based on a previously coded reference block in the current picture in operation, and a prediction control unit 128 connected to the memories 118 and 122. In operation, the prediction control unit 128 divides an image block into a plurality of partitions, including a first partition and a second partition, each having a non-rectangular shape, based on a partition parameter indicating the division (FIG. 22, step S4001), and encodes the first partition and the second partition (step S4002). In operation, the entropy coding unit 110 writes one or more parameters, including the partition parameter, into a bitstream (step S4003).

[0243] According to another embodiment, there is provided an image decoding device including a circuit and a memory coupled to the circuit, for example as shown in Figure 10. In operation, the circuit performs a partition syntax operation including: reading one or more parameters from a bitstream, including a partition parameter indicating partitioning of an image block into a plurality of partitions, including a first partition and a second partition, each having a non-rectangular shape (Figure 23, step S5001), partitioning the image block into the plurality of partitions based on the partition parameter (S5002), and decoding the first partition and the second partition (S5003).

[0244] Further according to an embodiment, the image decoding device shown in FIG. 10 is provided including: an entropy decoding unit 202 that, in operation, receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients; an inverse quantization unit 204 and an inverse transform unit 206 that, in operation, inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an adder 208 that, in operation, adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks; an inter prediction unit 218 that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit 216 that, in operation, generates a prediction of the current block based on a decoded reference block in the current picture; and a prediction control unit 220 connected to the memories 210, 214. In operation, the entropy decoding unit 202, in some implementations, in cooperation with the prediction control unit 220, reads one or more parameters from the bitstream, including partition parameters indicating division of the image block into multiple partitions, including a first partition having a non-rectangular shape and a second partition (Figure 23, step S5001), divides the image block into multiple partitions based on the partition parameters (S5002), and decodes the first partition and the second partition (S5003).

[0245] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can usually be realized by an MPU (micro processing unit), memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memory. It is also possible to realize each functional block by hardware (dedicated circuitry). Various combinations of hardware and software may be employed.

[0246] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Furthermore, the processor that executes the program may be a single processor or multiple processors. In other words, centralized processing or distributed processing may be performed.

[0247] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.

[0248] Furthermore, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems implementing the application examples will be described. Such systems may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device including both. Other configurations of such systems can be appropriately changed depending on the situation.

[0249] [Usage example] 26 is a diagram showing the overall configuration of an appropriate content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0250] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may connect any combination of the above devices. In various implementations, the devices may be connected to each other directly or indirectly via a telephone network or short-range wireless communication, without the intervention of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 may be connected to a terminal in a hotspot on an airplane ex117, etc., via a satellite ex116.

[0251] Note that wireless access points, hotspots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0252] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handy-phone System) compatible with mobile communication systems such as 2G, 3G, 3.9G, 4G, and the upcoming 5G.

[0253] The home appliance ex114 is a refrigerator or an appliance included in a home fuel cell cogeneration system.

[0254] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content captured by a user using the terminal, may multiplex the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and may transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0255] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client may be a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117, which is capable of decoding the encoded data. Each device that receives the distributed data may decode and play the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0256] [Distributed processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented as a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server can be dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing latency. Furthermore, when certain types of errors occur or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the failed portion of the network, thereby achieving high-speed and stable distribution.

[0257] In addition to the distributed processing of the distribution itself, the encoding of captured data can be performed on each device, on the server side, or shared among devices. For example, encoding generally involves two processing loops. The first loop detects the image complexity or code size for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a device can perform the first encoding process, and the server that receives the content can perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each device. In this case, if there is a request for near-real-time reception and decoding, the data encoded by a device can be received and played back on another device, enabling more flexible real-time distribution.

[0258] As another example, the camera ex113 or the like extracts features (quantities of features or characteristics) from an image, compresses the data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning of the image (or the importance of the content), for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server recompresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a heavy processing load such as CABAC (context-adaptive binary arithmetic coding).

[0259] As another example, in a stadium, shopping mall, factory, etc., there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.

[0260] Since multiple video data are of nearly the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referential. The server may also receive encoded data from each terminal and change the reference relationships between multiple data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.

[0261] Furthermore, the server may perform transcoding to change the encoding format of the video data before distributing it. For example, the server may convert an MPEG-based encoding format into a VP-based encoding format (e.g., VP9), or convert H.264 to H.265.

[0262] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0263] [3D, multi-angle] Images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other are increasingly being integrated for use. The videos taken by each device can be integrated based on the relative positions of the devices acquired separately or on areas where feature points in the videos match.

[0264] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a 3D shape of the scene based on not only the 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point clouds, etc., or may select or reconstruct images to be transmitted to the receiving terminal from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0265] In this way, a user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting device, or can enjoy content in which a video from a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0266] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, and may perform encoding that allows reference between the viewpoint images using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0267] In the case of AR images, the server may superimpose virtual object information in the virtual space onto camera information in the real space based on the 3D position or the movement of the user's viewpoint. The decoding device may acquire or store virtual object information and 3D data, generate a 2D image according to the movement of the user's viewpoint, and smoothly connect the images to create superimposed data. Alternatively, the decoding device may send the user's viewpoint movement to the server in addition to a request for virtual object information. The server may create superimposed data based on the viewpoint movement received from the 3D data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data typically has an α value indicating transparency in addition to RGB. The server may set the α value of parts other than the object created from the 3D data to 0, for example, to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object. The predetermined RGB value may be predetermined.

[0268] Similarly, the decoding of distributed data may be performed by the client (e.g., a terminal), by the server, or by both parties. As an example, a terminal may first send a reception request to a server, and then another terminal may receive and decode content according to the request, and then transmit the decoded signal to a device with a display. By distributing the processing and selecting appropriate content regardless of the performance of the communication-capable terminals themselves, it is possible to reproduce data with high image quality. As another example, large-sized image data may be received on a TV or the like, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of ​​responsibility or an area they wish to view in more detail.

[0269] In situations where multiple short-, medium-, or long-range wireless communications are available, both indoors and outdoors, it may be possible to seamlessly receive content using distribution system standards such as MPEG-DASH. Users may freely select and switch between decoding and display devices, such as their own devices and indoor / outdoor displays, in real time. Decoding can also be performed by switching between decoding and display devices using location information. This allows information to be mapped and displayed on a part of the wall or ground of a neighboring building with an embedded display device while the user is moving toward their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as if the encoded data is cached on a server that can be quickly accessed from the receiving device or copied to an edge server in a content delivery service.

[0270] [Scalable Coding] Regarding content switching, we will explain it using a scalable stream, as shown in Figure 27, compressed and encoded using the video encoding method described in each of the above embodiments. The server may have multiple streams with the same content but different qualities as individual streams, or it may be configured to switch content by taking advantage of the characteristics of a temporally / spatially scalable stream, which is achieved by encoding the content separately into layers as shown. In other words, the decoding side determines which layer to decode based on internal factors such as performance and external factors such as communication bandwidth, allowing the decoding side to freely switch between low-resolution and high-resolution content. For example, if a user wants to continue watching a video they were watching on their smartphone ex115 while on the go on a device such as an Internet TV after returning home, the device can simply decode the same stream up to a different layer, thereby reducing the burden on the server.

[0271] Furthermore, as described above, pictures are coded for each layer, and in addition to the configuration in which scalability is achieved by an enhancement layer above the base layer, the enhancement layer may include meta-information based on image statistics, etc. The decoding side may generate high-quality content by super-resolving pictures in the base layer based on the meta-information. Super-resolution may improve the signal-to-noise ratio while maintaining and / or expanding the resolution. The meta-information may include information for specifying linear or non-linear filter coefficients used in super-resolution processing, or information for specifying parameter values ​​in filter processing, machine learning, or least-squares calculations used in super-resolution processing.

[0272] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of objects in the image. The decoding side selects tiles to decode, thereby decoding only a portion of the area. Furthermore, by storing the object's attributes (such as a person, a car, or a ball) and its position in the video (such as a coordinate position in the same image) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 28, the meta information may be stored using a data storage structure different from pixel data, such as an SEI (supplemental enhancement information) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0273] Meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in the video, and by combining the picture-by-picture information with the time information, it can identify the picture in which the object exists and determine the position of the object within the picture.

[0274] [Webpage optimization] FIG. 29 is a diagram showing an example of a web page display screen on a computer ex111 or the like. FIG. 30 is a diagram showing an example of a web page display screen on a smartphone ex115 or the like. As shown in FIGS. 29 and 30, a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device on which the page is viewed. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture included in each content as a link image until the user explicitly selects the link image, or until the link image approaches the center of the screen or until the entire link image is within the screen, or may display a video such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the video.

[0275] When a link image is selected by a user, the display device performs decoding, for example, giving top priority to the base layer. Note that if the HTML constituting the web page contains information indicating that the content is scalable, the display device may also decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device decodes and displays only forward-referenced pictures (I pictures, P pictures, and forward-reference-only B pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Furthermore, the display device may intentionally ignore the picture reference relationships and roughly decode all B and P pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.

[0276] [Autonomous driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0277] In this case, since a vehicle, drone, or airplane including a receiving terminal is moving, the receiving terminal can transmit location information of the receiving terminal, thereby realizing seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0278] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0279] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distribution companies, but also low-quality, short-duration content from individuals via unicast or multicast. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, using the following configuration.

[0280] During shooting, either in real time or after accumulating and shooting, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image data or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.

[0281] Personal content may contain content that, if left as is, violates copyright, moral rights, or portrait rights, and may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. Furthermore, the server may recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons. The server may replace the specified area with another image or blur the focus. If the image contains a person, the server may track the person in the video and replace the image of the person's face.

[0282] Because viewing personal content with small data volumes requires real-time performance, the decoding device may first receive the base layer as a top priority and decode and play it back, depending on the bandwidth. The decoding device may also receive the enhancement layer during this time, and if the content is played back more than once, such as when playback is looped, play back high-quality video including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smoother and the image quality improves. In addition to scalable encoding, a similar experience can also be provided by configuring a single stream consisting of a rough stream played the first time and a second stream that is encoded with reference to the first video.

[0283] [Other application examples] Furthermore, these encoding or decoding processes are generally performed by an LSI ex500 possessed by each terminal. The LSI (large scale integration circuitry) ex500 (see FIG. 26) may be a single chip or may be configured with multiple chips. It is also possible to incorporate video encoding or decoding software into some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by a computer ex111 or the like, and perform the encoding or decoding process using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data may be data encoded by the LSI ex500 possessed by the smartphone ex115.

[0284] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal may download a codec or application software and then acquire and play the content.

[0285] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, the content supply system ex100 is more suited to multicast than the content supply system ex100, which is more suited to unicast, but similar applications are possible with regard to encoding and decoding processes.

[0286] [Hardware configuration] FIG. 31 is a diagram illustrating further details of the smartphone ex115 illustrated in FIG. 26. FIG. 32 is a diagram illustrating an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and other data received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing encoded or decoded data such as captured video or still images, recorded voice, received video or still images, and email, and a slot unit ex464 that serves as an interface with a SIM ex468 for identifying users and authenticating access to various data, including networks. In addition, an external memory may be used instead of the memory unit ex467.

[0287] A main control unit ex460 that can comprehensively control the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0288] When the power key is turned on by a user, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state and supplies power to each unit from the battery pack.

[0289] The smartphone ex115 processes calls, data communications, and other communications under the control of a main control unit ex460, which includes a CPU, ROM, RAM, and other components. During a call, the audio signal collected by the audio input unit ex456 is converted to a digital audio signal by the audio signal processing unit ex454, which then performs spectrum spread processing on the modulation / demodulation unit ex452. The resulting signal is then transmitted via the antenna ex450. The received data is then amplified, subjected to frequency conversion and analog-to-digital conversion, subjected to spectrum despreading processing on the modulation / demodulation unit ex452, and converted to an analog audio signal by the audio signal processing unit ex454, which then outputs the resulting signal from the audio output unit ex457. During data communications, text, still images, or video data can be transmitted via the operation input control unit ex462 under the control of the main control unit ex460, based on the operation of the main unit's operation unit ex466, etc. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the video encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the video or still image is being captured by the camera unit ex465, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the multiplexed video data and audio data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, before transmitting the multiplexed video data and audio data via the antenna ex450. The predetermined method may be determined in advance.

[0290] In the case of receiving video attached to an e-mail or chat, or video linked to a web page, for example, the multiplexed data received via the antenna ex450 is decoded by the multiplexing / demultiplexing unit ex453, which separates the multiplexed data into a video data bitstream and an audio data bitstream. The multiplexing / demultiplexing unit ex453 then demultiplexes the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and displays the video or still images contained in the linked video file on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. As real-time streaming becomes increasingly common, audio playback may be socially inappropriate depending on the user's circumstances. Therefore, it is preferable that the initial setting be a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0291] Although the smartphone ex115 has been used as an example, other implementations of the terminal are possible, such as a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed onto video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Furthermore, the video data itself may be received or transmitted instead of the multiplexed data.

[0292] While the main control unit ex460, which includes a CPU, has been described as controlling the encoding and decoding processes, various terminals often include a GPU. Therefore, a configuration in which a memory shared by the CPU and GPU, or a memory with addresses managed for common use, is also possible, leveraging the GPU's performance to process a large area at once. This shortens encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation and quantization processes at a picture level or other unit in the GPU rather than the CPU.

Claims

1. The circuit and a memory connected to the circuit; The circuit, in operation, determining whether intra prediction is applied to the image block; When the intra prediction is not applied to the image block, for a boundary between a first partition having a non-rectangular shape in the image block and a second partition having a non-rectangular shape in the image block based on a first parameter, (a) selecting a first motion vector for the first partition from a first motion vector candidate set and determining a first value for the first partition using the first motion vector, (b) selecting a second motion vector for the second partition from the first motion vector candidate set and determining a second value for the second partition using the second motion vector, (c) weighting the first value and the second value for a plurality of pixels where the first partition and the second partition overlap, and (d) performing a boundary smoothing process to encode the image block using the weighted first value and the weighted second value; Disabling the boundary smoothing process for the image block when the intra prediction is applied to the image block and a ratio of a width of the image block to a height of the image block is greater than 4, or when a ratio of the height to the width is greater than 4. Image encoding device.

2. The circuit and a memory connected to the circuit; The circuit, in operation, determining whether intra prediction is applied to the image block; When the intra prediction is not applied to the image block, for a boundary between a first partition having a non-rectangular shape in the image block and a second partition having a non-rectangular shape in the image block based on a first parameter, (a) selecting a first motion vector for the first partition from a first motion vector candidate set and determining a first value for the first partition using the first motion vector, (b) selecting a second motion vector for the second partition from the first motion vector candidate set and determining a second value for the second partition using the second motion vector, (c) weighting the first value and the second value for a plurality of pixels where the first partition and the second partition overlap, and (d) performing a boundary smoothing process to decode the image block using the weighted first value and the weighted second value; Disabling the boundary smoothing process for the image block when the intra prediction is applied to the image block and a ratio of a width of the image block to a height of the image block is greater than 4, or when a ratio of the height to the width is greater than 4. Image decoding device.

3. The circuit and a memory connected to the circuit; The circuit, in operation, generating first parameters for causing a decoding device to perform a boundary smoothing process on a boundary between a first partition having a non-rectangular shape in an image block and a second partition having a non-rectangular shape in the image block; including the first parameter in a bitstream; If intra prediction is not applied to the image block, the boundary smoothing process includes: a first motion vector for the first partition is selected from a first set of motion vector candidates; a first value of the first partition is calculated using the first motion vector; a second motion vector for the second partition is selected from the first set of motion vector candidates; a second value of the second partition is calculated using the second motion vector; the first and second values ​​of a plurality of pixels between the first and second partitions are weighted; the image block is decoded using the weighted first value and the weighted second value; When the intra prediction is applied to the image block and a ratio of a width of the image block to a height of the image block is greater than 4, or when a ratio of the height to the width is greater than 4, the decoding device disables the boundary smoothing process. Bitstream generator.

Citation Information

Patent Citations

  • Image processing device and image processing method

    JP2012023597A

  • Method for decomposing a video sequence frame

    US20080101707A1

  • Smoothing overlapped regions resulting from geometric motion partitioning

    US20110200110A1