Image Encoding Device, Image Decoding Device, and Bitstream Generation Device
By using non-rectangular image block partitions for motion vector selection and filter applications, the video coding technology enhances encoding efficiency and decoding speed, addressing the limitations of rectangular divisions in existing video coding systems.
Patent Information
- Application Number
- JP2024188064
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-16
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2038-08-10
AI Technical Summary
Existing video coding technologies face challenges in efficiently handling the increasing amount of digital video data, particularly in inter and intra predictions, where rectangular block divisions are limiting the optimization of motion vectors and filter applications.
The implementation of non-rectangular shapes, such as triangles and polygons, for image block partitions allows for the selection of motion vectors and filter applications that enhance encoding and decoding efficiency, including boundary smoothing operations and partition syntax for improved video coding.
This approach improves encoding efficiency, simplifies processing, and accelerates encoding/decoding speeds by allowing for more precise motion vector selection and filter application, optimizing video coding for complex image structures.
Smart Images

Figure 0007710081000003 
Figure 0007710081000004 
Figure 0007710081000005
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding, for example, constructing a current block based on a reference frame by performing inter prediction, or constructing a current block based on an encoded / decoded reference block within a current frame by performing intra prediction, and relates to systems, components, and methods in video encoding and decoding of moving images and the like.
Background Art
[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there has always been a need to provide improvements and optimizations to video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure further relates to advancements, improvements, and optimizations in video coding, particularly to inter prediction or intra prediction that divides an image block into a plurality of partitions including at least a first partition having a non-rectangular shape (e.g., a triangle) and a second partition.
Summary of the Invention
[0003] According to one aspect, an image encoding device is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit selects, for a first partition having a non-rectangular shape in an image block, a first motion vector of the first partition from a set of motion vector candidates, and for a second partition having a non-rectangular shape in the image block, a second motion vector of the second partition from the set of motion vector candidates. The first partition is encoded with the selected first motion vector, and the second partition is encoded with the selected second motion vector. From the set of motion vector candidates, only a single-prediction motion vector is selected without a condition related to the size of the image block.
[0004] In some examples of embodiments of the present disclosure, it is possible to improve encoding efficiency, simplify encoding / decoding processing, accelerate encoding / decoding processing speed, and efficiently select appropriate components / operations used in encoding and decoding, such as an appropriate filter, block size, motion vector, reference picture, reference block, etc.
[0005] Further benefits and advantages of the disclosed embodiments will become apparent from the specification and drawings. The benefits and / or advantages may be obtained individually by various embodiments and features of the specification and drawings, and it is not necessary to provide all the features of various embodiments and the specification and drawings in order to obtain one or more of such benefits and / or advantages.
[0006] Note that the comprehensive or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any combination thereof.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21A
Figure 21B
Figure 21C
Figure 21D
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
DETAILED DESCRIPTION OF THE INVENTION
[0008] According to one aspect, there is provided an image encoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, encodes the first partition using the first motion vector, and encodes the second partition using the second motion vector.
[0009] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles.
[0010] According to another aspect, the predicting includes selecting the first motion vector from a first set of motion vector candidates and selecting the second motion vector from a second set of motion vector candidates. For example, the first set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the first partition, and the second set of motion vector candidates may include motion vectors of a plurality of partitions adjacent to the second partition. The plurality of partitions adjacent to the first partition and the plurality of partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The plurality of adjacent partitions may be one or both of a plurality of spatially adjacent partitions and a plurality of temporally adjacent partitions. The first set of motion vector candidates may be the same as or different from the second set of motion vector candidates.
[0011] According to another aspect, the predicting includes selecting a first motion vector candidate from a first set of motion vector candidates, deriving the first motion vector by adding a first differential motion vector to the first motion vector candidate, selecting a second motion vector candidate from a second set of motion vector candidates, and deriving the second motion vector by adding a second differential motion vector to the second motion vector candidate.
[0012] According to another aspect, in operation, a splitting unit that receives an original image and splits it into a plurality of blocks, in operation, an addition unit that receives the plurality of blocks from the splitting unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual, in operation, a conversion unit that performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients, in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients, in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bit stream, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture, in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture, and an image encoding apparatus including the prediction control unit connected to a memory are provided. The prediction control unit, in operation, splits the plurality of blocks into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, encodes the first partition using the first motion vector, and encodes the second partition using the second motion vector.
[0013] According to another aspect, there is provided an image encoding method including splitting an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition, predicting a second motion vector for the second partition, encoding the first partition using the first motion vector, and encoding the second partition using the second motion vector.
[0014] According to one aspect, there is provided an image decoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, decodes the first partition using the first motion vector, and decodes the second partition using the second motion vector.
[0015] According to a further aspect, the second partition has a non-rectangular shape. According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles.
[0016] According to another aspect, in operation, there is provided an image decoding apparatus including an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, an inverse quantization unit and an inverse transform unit that, in operation, inverse-quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse-transform the plurality of transform coefficients to obtain a plurality of residuals, an addition unit that, in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, an inter prediction unit that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit that, in operation, generates a prediction of a current block based on a decoded reference block in a current picture, and the prediction control unit connected to the memory. In operation, the prediction control unit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicts a first motion vector for the first partition, predicts a second motion vector for the second partition, decodes the first partition using the first motion vector, and decodes the second partition using the second motion vector.
[0017] According to another aspect, there is provided an image decoding method including dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, predicting a first motion vector for the first partition, predicting a second motion vector for the second partition, decoding the first partition using the first motion vector, and decoding the second partition using the second motion vector.
[0018] According to one aspect, there is provided an image encoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0019] According to a further aspect, the non-rectangular shape is a triangle. According to another aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles. According to yet another aspect, the second partition has a non-rectangular shape.
[0020] According to another aspect, at least one of the first prediction and the second prediction is an inter-prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in an encoded reference picture. The inter-prediction process may predict the plurality of first values of the plurality of pixels in the first partition including the pixel set, or may predict the plurality of second values only of the pixel set in the first partition.
[0021] According to another aspect, at least one of the first prediction and the second prediction is an intra-prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in a current picture.
[0022] According to another aspect, the prediction method used for the first prediction is different from the prediction method used for the second prediction.
[0023] According to a further aspect, the number of the pixel sets in each row or each column for predicting the plurality of first values and the plurality of second values is an integer. For example, when the number of the pixel sets in each row or each column is 4, a plurality of weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be respectively applied to the plurality of first values of the 4 pixels in the pixel set, and a plurality of weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be respectively applied to the plurality of second values of the 4 pixels in the pixel set. As another example, when the number of the pixel sets in each row or each column is 2, a plurality of weights of 1 / 3 and 2 / 3 may be respectively applied to the plurality of first values of the 2 pixels in the pixel set, and a plurality of weights of 2 / 3 and 1 / 3 may be respectively applied to the plurality of second values of the 2 pixels in the pixel set.
[0024] According to another aspect, the plurality of weights may be integer values or fractional values.
[0025] According to another aspect, in operation, a splitting unit that receives an original image and splits it into a plurality of blocks, in operation, a adding unit that receives the plurality of blocks from the splitting unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual, in operation, a conversion unit that performs a conversion on the plurality of residuals output from the adding unit and outputs a plurality of conversion coefficients, in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients, in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bitstream, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture, in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture, and an image encoding apparatus including the prediction control unit connected to a memory are provided. The prediction control unit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are split from an image block, in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0026] According to another aspect, there is provided an image encoding method for performing a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The method generally includes four steps of: first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values and the plurality of second values; and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0027] According to a further aspect, there is provided an image decoding apparatus including a circuit and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The boundary smoothing operation includes: first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition; second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition; weighting the plurality of first values and the plurality of second values; and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0028] According to another aspect, the non-rectangular shape is a triangle. According to a further aspect, the non-rectangular shape is selected from the group consisting of a triangle, a trapezoid, and a polygon having at least five sides and angles. According to another aspect, the second partition has a non-rectangular shape.
[0029] According to another aspect, at least one of the first prediction and the second prediction is an inter prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in an encoded reference picture. The inter prediction process may predict the plurality of first values of the plurality of pixels of the first partition including the pixel set, or may predict the plurality of second values only of the pixel set of the first partition.
[0030] According to another aspect, at least one of the first prediction and the second prediction is an intra prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in a current picture.
[0031] According to another aspect, in operation, an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, and in operation, inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, an inverse quantization unit and an inverse transform unit, and in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, an addition unit, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit that generates a prediction of the current block based on a decoded reference block in a current picture, and an image decoding apparatus including the prediction control unit connected to a memory are provided. The prediction control unit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block, in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition, second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition, weighting the plurality of first values and the plurality of second values, and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0032] According to another aspect, an image decoding method is provided that performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block. The method generally includes four steps: predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition; predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition; weighting the plurality of first values and the plurality of second values; and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values.
[0033] According to one aspect, an image encoding apparatus is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit performs a partition syntax operation, which includes dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the division, encoding the first partition and the second partition, and writing one or more parameters including the partition parameter into a bitstream.
[0034] According to a further aspect, the partition parameter indicates that the first partition has a triangular shape.
[0035] According to another aspect, the partition parameter indicates that the second partition has a non-rectangular shape.
[0036] According to another aspect, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.
[0037] According to another aspect, the partition parameter encodes the division direction applied to divide the image block into the plurality of partitions as a whole. For example, the division direction may include from the upper left corner to the lower right corner of the image block, and from the upper right corner to the lower left corner of the image block. The partition parameter may encode at least the first motion vector of the first partition as a whole.
[0038] According to another aspect, one or more parameters other than the partition parameter encode the division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode at least the first motion vector of the first partition as a whole.
[0039] According to another aspect, the partition parameter may encode at least the first motion vector of the first partition as a whole. The partition parameter may encode the second motion vector of the second partition as a whole.
[0040] According to another aspect, one or more parameters other than the partition parameter may encode at least the first motion vector of the first partition.
[0041] According to another aspect, the one or more parameters are binarized according to a binarization method selected according to at least one value of the one or more parameters.
[0042] According to a further aspect, in operation, a splitting unit that receives an original image and splits it into a plurality of blocks, in operation, an addition unit that receives the plurality of blocks from the splitting unit, receives a plurality of predictions from a prediction control unit, subtracts each prediction from its corresponding block, and outputs a residual, in operation, a conversion unit that performs a conversion on the plurality of residuals output from the addition unit and outputs a plurality of conversion coefficients, in operation, a quantization unit that quantizes the plurality of conversion coefficients to generate a plurality of quantized conversion coefficients, in operation, an entropy encoding unit that encodes the plurality of quantized conversion coefficients to generate a bitstream, in operation, an inter prediction unit that generates a prediction of a current block based on a reference block in an encoded reference picture, in operation, an intra prediction unit that generates a prediction of a current block based on an encoded reference block in a current picture, and a prediction control unit connected to a memory are provided. The prediction control unit, in operation, splits an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating splitting, and encodes the first partition and the second partition. The entropy encoding unit, in operation, writes one or more parameters including the partition parameter into the bitstream.
[0043] According to another aspect, an image encoding method including a partition syntax operation is provided. The method generally includes three steps: splitting an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating splitting, encoding the first partition and the second partition, and writing one or more parameters including the partition parameter into the bitstream.
[0044] According to another aspect, an image decoding apparatus is provided that includes a circuit and a memory connected to the circuit. In operation, the circuit performs a partition syntax operation, and the partition syntax operation includes decoding, from a bitstream, one or more parameters including partition parameters indicating dividing an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition, dividing the image block into the plurality of partitions based on the partition parameters, and decoding the first partition and the second partition.
[0045] According to a further aspect, the partition parameter indicates that the first partition has a triangular shape.
[0046] According to another aspect, the partition parameter indicates that the second partition has a non-rectangular shape.
[0047] According to another aspect, the partition parameter indicates that the non-rectangular shape is one of a triangle, a trapezoid, and a polygon having at least five sides and angles.
[0048] According to another aspect, the partition parameter encodes, integrally, a division direction applied to divide the image block into the plurality of partitions. For example, the division direction includes from the upper left corner to the lower right corner of the image block and from the upper right corner to the lower left corner of the image block. The partition parameter may encode, integrally, at least a first motion vector of the first partition.
[0049] According to another aspect, one or more parameters other than the partition parameter encode a division direction applied to divide the image block into the plurality of partitions. The parameter encoding the division direction may encode, integrally, at least a first motion vector of the first partition.
[0050] According to another aspect, the partition parameter may at least integrally encode the first motion vector of the first partition. The partition parameter may integrally encode the second motion vector of the second partition.
[0051] According to another aspect, one or more parameters other than the partition parameter may at least encode the first motion vector of the first partition.
[0052] According to another aspect, the one or more parameters are binarized according to a binarization method selected according to at least one value of the one or more parameters.
[0053] According to a further aspect, in operation, an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, and in operation, inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inverse transforms the plurality of transform coefficients to obtain a plurality of residuals, an inverse quantization unit and an inverse transform unit, and in operation, adds the plurality of residuals output from the inverse quantization unit and the inverse transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks. An addition unit, an inter prediction unit that generates a prediction of a current block based on a reference block in a decoded reference picture in operation, an intra prediction unit that generates a prediction of a current block based on a decoded reference block in a current picture in operation, and an image decoding apparatus including the prediction control unit connected to a memory is provided. The entropy decoding unit reads, from the bitstream, one or more parameters including a partition parameter indicating that an image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition in operation, divides the image block into the plurality of partitions based on the partition parameter, and decodes the first partition and the second partition.
[0054] According to another aspect, an image decoding method including a partition syntax operation is provided. The method generally includes three steps of decoding one or more parameters including partition parameters indicating to divide an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition from a bit stream, dividing the image block into the plurality of partitions based on the partition parameters, and decoding the first partition and the second partition.
[0055] In the drawings, the same reference numerals denote the same components. The sizes and relative positions of the components in the drawings do not necessarily follow the scale ratio.
[0056] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that the embodiments described below are all examples showing comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Therefore, although not disclosed in the following embodiments, components not described in the independent claims defining the broadest inventive concept may be understood as arbitrary components.
[0057] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device to which the processing and / or configuration described in each aspect of the present disclosure is applicable. The processing and / or configuration can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with respect to the processing and / or configuration applied to the embodiments, for example, any of the following may be implemented.
[0058] (1) Any one of the plurality of components of the encoding device or decoding device according to the embodiments described in each aspect of the present disclosure may be replaced or combined with any other component described in any one of the aspects of the present disclosure.
[0059] (2) In the encoding device or decoding device according to the embodiments, any change such as addition, replacement, deletion, etc. may be made to the functions or processes performed by some of the plurality of components of the encoding device or decoding device. For example, any function or process may be replaced or combined with any other function or process described in any one of the aspects of the present disclosure.
[0060] (3) In the method performed by the encoding device or decoding device according to the embodiments, any change such as addition, replacement, and deletion may be made to some of the plurality of processes included in the method. For example, any process in the method may be replaced or combined with any other process described in any one of the aspects of the present disclosure.
[0061] (4) Some of the plurality of components constituting the encoding device or decoding device according to the embodiments may be combined with the components described in any one of the aspects of the present disclosure, or may be combined with the components having a part of the functions described in any one of the aspects of the present disclosure, or may be combined with the components that perform a part of the processes performed by the components described in each aspect of the present disclosure.
[0062] (5) The components having a part of the functions of the encoding device or decoding device according to the embodiments, or the components that perform a part of the processes of the encoding device or decoding device according to the embodiments may be combined or replaced with the components described in any one of the aspects of the present disclosure, the components having a part of the functions described in any one of the aspects of the present disclosure, or the components that perform a part of the processes described in any one of the aspects of the present disclosure.
[0063] (6) In the method implemented by the encoding device or decoding device of the embodiment, any one of a plurality of processes included in the method may be replaced or combined with the processes described in any of the aspects of the present disclosure or any similar processes.
[0064] (7) Some of the plurality of processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.
[0065] (8) The implementation manners of the processes and / or configurations described in each aspect of the present disclosure are not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or configurations may be implemented in a device used for a purpose different from the moving image encoding or moving image decoding disclosed in the embodiment.
[0066] [Overview of Encoding Device] First, the overview of the encoding device according to the embodiment will be described. FIG. 1 is a block diagram showing the functional configuration of the encoding device 100 according to the embodiment. The encoding device 100 is a moving image encoding device that encodes a moving image in block units.
[0067] As shown in FIG. 1, the encoding device 100 is a device that encodes an image in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.
[0068] The symbolization device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as a splitting unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a loop filter unit 120, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. Further, the symbolization device 100 may be realized as one or more dedicated electronic circuits corresponding to the splitting unit 102, the subtraction unit 104, the conversion unit 106, the quantization unit 108, the entropy encoding unit 110, the inverse quantization unit 112, the inverse conversion unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0069] Each component included in the symbolization device 100 will be described below.
[0070] [Splitting Unit] The splitting unit 102 splits each picture included in the input moving image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128). These fixed-size blocks are sometimes called coding tree units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into variable-size blocks (e.g., 64x64 or less) based on recursive quadtree and / or binary tree block splitting. These variable-size blocks are sometimes called coding units (CUs), prediction units (PUs), or transform units (TUs). In various processing examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be processing units for CUs, PUs, and TUs.
[0071] FIG. 2 is a diagram showing an example of block splitting in the embodiment. In FIG. 2, solid lines represent block boundaries by quadtree block splitting, and broken lines represent block boundaries by binary tree block splitting.
[0072] Here, block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quad-tree block division).
[0073] The upper-left 64x64 block is further vertically divided into two rectangular 32x64 blocks, and the left 32x64 block is further vertically divided into two rectangular 16x64 blocks (binary-tree block division). As a result, the upper-left 64x64 block is divided into two 16x64 blocks 11, 12 and a 32x64 block 13.
[0074] The upper-right 64x64 block is horizontally divided into two rectangular 64x32 blocks 14, 15 (binary-tree block division).
[0075] The lower-left 64x64 block is divided into four square 32x32 blocks (quad-tree block division). Among the four 32x32 blocks, the upper-left block and the lower-right block are further divided. The upper-left 32x32 block is vertically divided into two rectangular 16x32 blocks, and the right 16x32 block is further horizontally divided into two 16x16 blocks (binary-tree block division). The lower-right 32x32 block is horizontally divided into two 32x16 blocks (binary-tree block division). As a result, the lower-left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.
[0076] The lower-right 64x64 block 23 is not divided.
[0077] As described above, in FIG. 2, block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.
[0078] In FIG. 2, one block was divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to these. For example, one block may be divided into three blocks (ternary-tree block division). Such a division including ternary-tree block division may be called MBT (multi type tree) division.
[0079] [Subtraction unit] The subtraction unit 104 receives an input from the division unit 102 and subtracts a prediction signal (prediction samples input from the prediction control unit 128 shown below) from the original signal (original samples) in units of blocks divided by the division unit 102. That is, the subtraction unit 104 calculates the prediction error (also referred to as the residual) of the block to be encoded (hereinafter referred to as the current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.
[0080] The original signal is the input signal of the encoding device 100 and is a signal representing the image of each picture constituting the moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.
[0081] [Conversion unit] The conversion unit 106 converts the prediction error in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain.
[0082] Note that the conversion unit 106 may adaptively select a conversion type from a plurality of conversion types and convert the prediction error into conversion coefficients using a transform basis function corresponding to the selected conversion type. Such a conversion may be called EMT (explicit multiple core transform) or AMT (adaptive multiple transform).
[0083] The plurality of transform types includes, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. FIG. 3 is a table showing the transform basis functions corresponding to each transform type. In FIG. 3, N indicates the number of input pixels. The selection of a transform type from among these plurality of transform types may depend on, for example, the type of prediction (intra prediction and inter prediction), or may depend on the intra prediction mode.
[0084] Information indicating whether to apply such EMT or AMT (for example, called an EMT flag or an AMT flag) and information indicating the selected transform type are usually signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, the bit sequence level, the picture level, the slice level, the tile level, or the CTU level).
[0085] Also, the conversion unit 106 may re-convert the conversion coefficients (conversion results). Such re-conversion may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (for example, a 4x4 sub-block) included in the block of conversion coefficients corresponding to the intra prediction error. Information indicating whether to apply NSST and information regarding the conversion matrix used for NSST are usually signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level).
[0086] The conversion unit 106 may be applied with separable conversion and non-separable conversion. Separable conversion is a method of performing conversion multiple times by separating for each direction by the number of dimensions of the input, and non-separable conversion is a method of treating two or more dimensions as one dimension when the input is multi-dimensional and performing conversion collectively.
[0087] For example, as an example of non-separable conversion, when the input is a 4×4 block, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.
[0088] In a further example of non-separable conversion, after regarding a 4×4 input block as an array having 16 elements, a conversion (for example, Hypercube Givens Transform) in which Givens rotation is performed multiple times on the array may be performed.
[0089] [Quantization Unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans the conversion coefficients of the current block in a predetermined scanning order, and quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0090] The predetermined scanning order is an order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).
[0091] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. That is, as the value of the quantization parameter increases, the quantization error increases.
[0092] [Entropy Encoding Unit] The entropy encoding unit 110 generates an encoded signal (encoded bit stream) based on the quantized coefficients input from the quantization unit 108. Specifically, for example, the entropy encoding unit 110 binarizes the quantized coefficients, arithmetic-encodes the binary signal, and outputs a compressed bit stream or sequence.
[0093] [Inverse Quantization Unit] The inverse quantization unit 112 inverse-quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse-quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114.
[0094] [Inverse Transform Unit] The inverse transform unit 114 restores the prediction error (residual) by inverse-transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients to restore the prediction error of the current block. Then, the inverse transform unit 114 outputs the restored prediction error to the addition unit 116.
[0095] Note that since information is usually lost due to quantization, the restored prediction error does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error usually contains a quantization error.
[0096] [Addition Unit] The adder 116 reconstructs the current block by adding the prediction error input from the inverse conversion unit 114 and the prediction sample input from the prediction control unit 128. Then, the adder 116 outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block may also be called a local decoding block.
[0097] [Block Memory] The block memory 118 is a storage unit for storing blocks within the coded target picture (referred to as the "current picture") that are referenced in intra prediction. Specifically, the block memory 118 stores the reconstructed block output from the adder 116.
[0098] [Loop Filter Unit] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder 116 and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used within the coding loop, and includes, for example, a deblocking filter (DF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).
[0099] In the case of the ALF, a least squares error filter for removing coding distortion is applied, and for example, for each 2x2 sub-block within the current block, one filter selected from a plurality of filters is applied based on the local gradient direction and activity.
[0100] Specifically, first, sub-blocks (for example, 2x2 sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed based on the gradient direction and activity. For example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes.
[0101] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantizing the addition result.
[0102] Based on the results of such classification, a filter for a sub-block is determined from among a plurality of filters.
[0103] As the shape of the filter used in ALF, for example, a circularly symmetric shape is utilized. FIGS. 4A to 4C are diagrams showing a plurality of examples of the shape of the filter used in ALF. FIG. 4A shows a 5x5 diamond-shaped filter, FIG. 4B shows a 7x7 diamond-shaped filter, and FIG. 4C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Note that the signaling of the information indicating the shape of the filter need not be limited to the picture level and may be at other levels (for example, sequence level, slice level, tile level, CTU level, or CU level).
[0104] The on / off of ALF may be determined at the picture level or the CU level. For example, for luminance, it may be determined whether to apply ALF at the CU level, and for chrominance difference, it may be determined whether to apply ALF at the picture level. Information indicating the on / off of ALF is usually signaled at the picture level or the CU level. Note that the signaling of the information indicating the on / off of ALF need not be limited to the picture level or the CU level and may be at other levels (for example, sequence level, slice level, tile level, or CTU level).
[0105] A set of coefficients for a plurality of selectable filters (e.g., filters up to 15 or 25) is usually signaled at the picture level. Note that the signaling of the coefficient set does not have to be limited to the picture level and may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0106] [Frame memory] Frame memory 122 is a storage unit for storing reference pictures used for inter prediction and may also be called a frame buffer, for example. Specifically, frame memory 122 stores the reconstructed blocks filtered by loop filter unit 120.
[0107] [Intra prediction unit] Intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called in-picture prediction) of the current block with reference to a block within the current picture stored in block memory 118. Specifically, intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block and outputs the intra prediction signal to prediction control unit 128.
[0108] For example, intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes usually include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0109] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC standard.
[0110] The plurality of directional prediction modes include, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may include, in addition to the 33 directions, a further 32-direction prediction mode (a total of 65 directional prediction modes).
[0111] FIG. 5A is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the 2 "non-directional" prediction modes are not shown in FIG. 5A).
[0112] In various processing examples, in the intra prediction of a chrominance block, a luminance block may be referred to. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. This intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that refers to such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.
[0113] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradient of the reference pixels in the horizontal / vertical direction. Such intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called a PDPC flag) is usually signaled at the CU level. Note that the signaling of this information is not necessarily limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, tile level or CTU level).
[0114] [Inter Prediction Unit] The inter prediction unit 126 generates a prediction signal (inter prediction signal) by performing inter prediction (also called inter - picture prediction) of a current block with reference to a reference picture stored in the frame memory 122 that is different from the current picture. The inter prediction is performed in units of a current block or a current sub - block (e.g., 4x4 block) within the current block. For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or current sub - block, and finds a reference block or sub - block in the reference picture that most matches the current block or sub - block. Then, the inter prediction unit 126 obtains motion information (e.g., a motion vector) that compensates for (or predicts) the motion or change from the reference block or sub - block to the current block or sub - block. Then, the inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, and generates an inter prediction signal for the current block or sub - block. Then, the inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.
[0115] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms. For example, a motion vector may be signaled. As another example, the difference between a motion vector and a predicted motion vector may be signaled.
[0116] Note that an inter prediction signal may be generated using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks. Specifically, an inter prediction signal may be generated in units of sub - blocks within the current block by weighted - adding a prediction signal based on the motion information obtained by motion search (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture). Such inter prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation).
[0117] In the OBMC mode, information indicating the size of sub-blocks for OBMC (e.g., referred to as the OBMC block size) may be signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (e.g., referred to as the OBMC flag) may be signaled at the CU level. Note that the signaling level of these pieces of information does not necessarily have to be limited to the sequence level and the CU level, and may be at other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0118] The OBMC mode will be described more specifically. FIGS. 5B and 5C are a flowchart and a conceptual diagram for explaining the predictive image correction process by OBMC processing.
[0119] Referring to FIG. 5C, first, a predictive image (Pred) by normal motion compensation is obtained using the motion vector (MV) assigned to the block to be coded (current). In FIG. 5C, the arrow "MV" points to the reference picture, indicating what the current block in the current picture refers to in order to obtain the predictive image.
[0120] Next, the motion vector (MV_L) already derived for the encoded left adjacent block is applied (reused) to the block to be coded (current) to obtain a predictive image (Pred_L). The motion vector (MV_L) is indicated by the arrow "MV_L" pointing from the current block to the reference picture. Then, the first correction of the predictive image is performed by superimposing the two predictive images Pred and Pred_L. This has the effect of mixing the boundaries between adjacent blocks.
[0121] Similarly, the motion vector (MV_U) already derived for the encoded upper adjacent block is applied (reused) to the block to be encoded (current) block to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. Then, the predicted image Pred_U is superposed on the predicted image after the first correction (i.e., Pred and Pred_L) to perform the second correction of the predicted image. This has the effect of blending the boundaries between adjacent blocks in one aspect. The predicted image obtained by the second correction is the final predicted image of the current block in which the boundary with the adjacent block is blended (smoothed).
[0122] Here, a two-stage correction method using the left adjacent block and the upper adjacent block has been described. However, it is also possible to perform correction more times than two stages using the right adjacent block or the lower adjacent block.
[0123] Note that the region for superposition may be only a partial region near the block boundary, rather than the entire pixel region of the block.
[0124] Here, the prediction image correction process of OBMC for obtaining one predicted image Pred by superposing additional predicted images Pred_L and Pred_U based on one reference picture has been described. However, when the predicted image is corrected based on a plurality of reference pictures, the same process may be applied to each of the plurality of reference pictures. In such a case, by performing OBMC image correction based on a plurality of reference pictures, after obtaining a corrected predicted image from each reference picture, the obtained plurality of corrected predicted images are further superposed to obtain a final predicted image.
[0125] Note that in OBMC, the unit of the target block may be a prediction block unit or a sub-block unit obtained by further dividing the prediction block.
[0126] As a method for determining whether to apply OBMC processing, for example, there is a method of using an obmc_flag, which is a signal indicating whether to apply OBMC processing. As a specific example, in an encoding device, it is determined whether an encoding target block belongs to a region with complex motion. If it belongs to a region with complex motion, the value "1" is set as the obmc_flag, and OBMC processing is applied for encoding. If it does not belong to a region with complex motion, the value "0" is set as the obmc_flag, and encoding is performed without applying OBMC processing. On the other hand, in a decoding device, by decoding the obmc_flag described in a stream (i.e., a compressed sequence), decoding is performed by switching whether to apply OBMC processing according to the value.
[0127] Note that the motion information may be derived on the decoding device side without being signaled from the encoding device side. For example, the merge mode defined in the H.265 / HEVC standard may be used. Also, for example, motion information may be derived by performing motion search on the decoding device side. In this case, on the decoding device side, motion search may be performed without using the pixel values of the current block.
[0128] Here, a mode of performing motion search on the decoding device side will be described. This mode of performing motion search on the decoding device side may be called the PMMVD (pattern matched motion vector derivation) mode or the FRUC (frame rate up-conversion) mode.
[0129] An example of FRUC processing is shown in FIG. 5D. First, by referring to the motion vectors of encoded blocks that are spatially or temporally adjacent to the current block, a list of a plurality of candidates (which may be common to the merge list) each having a predicted motion vector (MV) is generated. Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate list. For example, an evaluation value of each candidate MV included in the candidate list is calculated, and one candidate MV is selected based on the evaluation value.
[0130] Then, based on the motion vectors of the selected candidates, a motion vector for the current block is derived. Specifically, for example, the motion vector of the selected candidate (best candidate MV) is directly derived as the motion vector for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the motion vector of the selected candidate, by performing pattern matching, a motion vector for the current block may be derived. That is, search is performed on the region around the best candidate MV using pattern matching and evaluation values in the reference picture. Further, if there is an MV with a better evaluation value, the best candidate MV may be updated to the MV and used as the final MV of the current block. It is also possible to adopt a configuration in which the process of updating to an MV with a better evaluation value is not performed.
[0131] When processing is performed in units of sub-blocks, the same processing may be adopted.
[0132] Note that the evaluation value may be calculated in various ways. For example, the reconstructed image of the region in the reference picture corresponding to the motion vector may be compared with the reconstructed image of a predetermined region (for example, as described later, it may be a region of another reference picture or a region of an adjacent block of the current picture), and the difference between the pixel values of the two reconstructed images may be calculated and used as the evaluation value of the motion vector. Note that, in addition to the difference value, other information may be used to calculate the evaluation value.
[0133] Next, an example of pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (for example, the merge list) is selected as the start point of the search by pattern matching. For example, as the pattern matching, first pattern matching or second pattern matching may be used. The first pattern matching and the second pattern matching may be called bilateral matching and template matching, respectively.
[0134] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures, which are two blocks along the motion trajectory of the current block. Therefore, in the first pattern matching, for the area in the reference picture, as the predetermined area for calculating the evaluation value of the above-mentioned candidate, the area in another reference picture along the motion trajectory of the current block is used.
[0135] FIG. 6 is a diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As shown in FIG. 6, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1), which are two blocks along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among a plurality of candidate MVs as the final MV, which can bring good results.
[0136] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, two mirror-symmetric bidirectional motion vectors are derived.
[0137] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (for example, the upper and / or left adjacent blocks)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined region for calculating the evaluation value of the candidate described above.
[0138] FIG. 7 is a diagram for explaining an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in FIG. 7, in the second pattern matching, the motion vector of the current block is derived by searching for the block that most matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic) in the reference picture (Ref0). Specifically, for the current block, the difference between the reconstructed images of both or either of the left adjacent and upper adjacent coded regions and the reconstructed image at the equivalent position in the coded reference picture (Ref0) specified by the candidate MV is derived, and an evaluation value is calculated using the obtained difference value, and it is possible to select the candidate MV with the best evaluation value among the plurality of candidate MVs as the best candidate MV.
[0139] Information indicating whether or not to apply such a FRUC mode (for example, called a FRUC flag) may be signaled at the CU level. Also, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the pattern matching method (for example, the first pattern matching or the second pattern matching) (for example, called a FRUC mode flag) may be signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0140] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode may be called the BIO (bi - directional optical flow) mode.
[0141] FIG. 8 is a diagram for explaining a model assuming uniform linear motion. In FIG. 8, (v x , v y ) represents a velocity vector, and τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the motion vector corresponding to the reference picture Ref1.
[0142] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented by (v x τ0, v y τ0) and (-v x τ1, -v y τ1), and the following optical flow equation (1) holds.
[0143]
Equation
[0144] Here, I (k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the merge list or the like may be corrected in pixel units.
[0145] Note that the motion vector may be derived on the decoder side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0146] Next, a mode of deriving a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks will be described. This mode may be called an affine motion compensation prediction mode.
[0147] FIG. 9A is a diagram for explaining the derivation of a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks. In FIG. 9A, the current block includes 16 4x4 sub-blocks. Here, the motion vector v0 of the upper left control point of the current block is derived based on the motion vectors of the adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, using the two motion vectors v0 and v1, the motion vector (v x , v y ) of each sub-block within the current block is derived by the following equation (2).
[0148]
Equation
[0149] Here, x and y respectively indicate the horizontal position and the vertical position of the sub-block, and w indicates a predetermined weight coefficient.
[0150] This affine motion compensation prediction mode may include several modes in which the methods for deriving the motion vectors of the upper left and upper right corner control points are different. Information indicating this affine motion compensation prediction mode (for example, called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating this affine motion compensation prediction mode need not be limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0151] [Prediction control unit] The prediction control unit 128 selects either an intra prediction signal (a signal output from the intra prediction unit 124) or an inter prediction signal (a signal output from the inter prediction unit 126), and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.
[0152] As shown in FIG. 1, in various processing examples, the prediction control unit 128 may output prediction parameters input to the entropy encoding unit 110. The entropy encoding unit 110 may generate an encoded bit stream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantized coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bit stream, and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (for example, a motion vector, a prediction type, or a prediction mode used in the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or indicating the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0153] FIG. 9B shows an example of a process for deriving the motion vector of a current picture in the merge mode.
[0154] First, a predicted MV list in which candidates for the predicted MV are registered is generated. As candidates for the predicted MV, there are a spatial adjacent predicted MV which is the MV of a plurality of encoded blocks spatially adjacent to the target block, a temporal adjacent predicted MV which is the MV of a neighboring block obtained by projecting the position of the target block in the encoded reference picture, a combined predicted MV which is an MV generated by combining the MV values of the spatial adjacent predicted MV and the temporal adjacent predicted MV, and a zero predicted MV which is an MV with a value of zero, and the like.
[0155] Next, one predicted MV is selected from among the plurality of predicted MVs registered in the predicted MV list, and thus determined as the MV of the target block.
[0156] Furthermore, in the variable length encoding unit, merge_idx, which is a signal indicating which predicted MV is selected, is described in the stream and encoded.
[0157] Note that the predicted MVs registered in the predicted MV list described in FIG. 9B are merely examples, and the number may be different from that in the figure, the configuration may not include some types of the predicted MVs in the figure, or the configuration may include predicted MVs other than the types of the predicted MVs in the figure.
[0158] The final MV may be determined by performing DMVR (decoder motion vector refinement) processing, which will be described later, using the MV of the target block derived in the merge mode.
[0159] FIG. 9C is a conceptual diagram for explaining an example of the DMVR processing for determining the MV.
[0160] First, take the optimal MVP set in the current block (e.g., in the merge mode) as the candidate MV. Then, according to the candidate MV (L0), identify reference pixels from the first reference picture (L0), which is the encoded picture in the L0 direction. Similarly, according to the candidate MV (L1), identify reference pixels from the second reference picture (L1), which is the encoded picture in the L1 direction. Generate a template by taking the average of these reference pixels.
[0161] Next, use the template to search the peripheral regions of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) respectively, and determine the MV with the minimum cost as the final MV. Note that the cost value may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search region, and the candidate MV value, etc.
[0162] Typically, in the encoder and the decoder described later, the configurations and operations of the processes described here are basically common.
[0163] Even if it is not the processing example described here, any processing may be used as long as it can search the periphery of the candidate MV to derive the final MV.
[0164] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) processing will be described.
[0165] FIG. 9D is a conceptual diagram for explaining an example of a predicted image generation method using luminance correction processing by LIC processing.
[0166] First, derive an MV from the encoded reference picture and obtain a reference image corresponding to the current block.
[0167] Next, for the current block, information indicating how the luminance values change between the reference picture and the current picture is extracted. This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV. Then, using the information indicating how the luminance values change, luminance correction parameters are calculated.
[0168] By performing a luminance correction process of applying the luminance correction parameters to the reference image in the reference picture specified by the MV, a predicted image for the current block is generated.
[0169] Note that the shape of the peripheral reference region in FIG. 9D is an example, and other shapes may be used.
[0170] Also, although the process of generating a predicted image from one reference picture has been described here, the same applies when generating a predicted image from multiple reference pictures. Luminance correction processing may be performed on the reference images obtained from each reference picture in the same manner as described above, and then a predicted image may be generated.
[0171] As a method for determining whether to apply the LIC process, for example, there is a method using lic_flag, which is a signal indicating whether to apply the LIC process. As a specific example, in the encoding device, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value "1" is set as lic_flag and encoding is performed by applying the LIC process. If it does not belong to a region where a luminance change has occurred, the value "0" is set as lic_flag and encoding is performed without applying the LIC process. On the other hand, in the decoding device, by decoding the lic_flag described in the stream, decoding may be performed by switching whether to apply the LIC process according to the value.
[0172] As another method for determining whether to apply the LIC process, for example, there is also a method of determining according to whether the LIC process is applied to peripheral blocks. As a specific example, when the current block is in the merge mode, it is determined whether the encoded peripheral blocks selected during the derivation of the MV in the merge mode process are encoded by applying the LIC process. Encoding is performed by switching whether to apply the LIC process according to the result. Note that even in this example, the same process is applied to the process on the decoder side.
[0173] [Outline of Decoder] Next, an outline of a decoder capable of decoding the encoded signal (encoded bit stream) output from the above-described encoder 100 will be described. FIG. 10 is a block diagram showing a functional configuration of a decoder 200 according to an embodiment. The decoder 200 is a moving image decoder that decodes a moving image in units of blocks.
[0174] As shown in FIG. 10, the decoder 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.
[0175] The decoder 200 is realized by, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. Further, the decoder 200 may be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.
[0176] Each component included in the decoder 200 will be described below.
[0177] [Entropy Decoding Unit] The entropy decoding unit 202 entropy-decodes the encoded bit stream. Specifically, the entropy decoding unit 202, for example, arithmetically decodes the encoded bit stream into a binary signal. Then, the entropy decoding unit 202 de-binarizes the binary signal. The entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may output prediction parameters included in the encoded bit stream (see FIG. 1) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can execute the same prediction processes as those performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoder side.
[0178] [Inverse Quantization Unit] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the decoding target block (hereinafter referred to as the current block) input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0179] [Inverse Transform Unit] The inverse transform unit 206 restores the prediction error (residual) by inverse-transforming the transform coefficients input from the inverse quantization unit 204.
[0180] For example, when the information read from the encoded bit stream indicates that EMT or AMT is applied (for example, the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the read transform type.
[0181] For example, when the information decoded from the encoded bitstream indicates that the NSST is to be applied, the inverse conversion unit 206 applies an inverse reconversion to the conversion coefficients.
[0182] [Addition unit] The addition unit 208 reconstructs the current block by adding the prediction error input from the inverse conversion unit 206 and the prediction sample input from the prediction control unit 220. Then, the addition unit 208 outputs the reconstructed block to the block memory 210 and the loop filter unit 212.
[0183] [Block memory] The block memory 210 is a storage unit for storing blocks within the decoded target picture (hereinafter referred to as the current picture) that are referenced in intra prediction. Specifically, the block memory 210 stores the reconstructed block output from the addition unit 208.
[0184] [Loop filter unit] The loop filter unit 212 applies a loop filter to the block reconstructed by the addition unit 208 and outputs the filtered reconstructed block to the frame memory 214 and a display device or the like.
[0185] When the information indicating the on / off of the ALF decoded from the encoded bitstream indicates that the ALF is on, one filter is selected from a plurality of filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.
[0186] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 214 stores the reconstructed block filtered by the loop filter unit 212.
[0187] [Intra prediction unit] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance difference values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.
[0188] Note that when an intra prediction mode that refers to a luminance block in the intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0189] Also, when the information decoded from the encoded bit stream (e.g., the prediction parameter output from the entropy decoding unit 202) indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0190] [Inter prediction unit] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) within the current block. For example, the inter prediction unit 218 generates an inter prediction signal for the current block or sub-block by performing motion compensation using the motion information (e.g., motion vector) decoded from the encoded bit stream (e.g., the prediction parameter output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.
[0191] When the information decoded from the encoded bit stream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.
[0192] Further, when the information decoded from the encoded bitstream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the encoded stream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.
[0193] Also, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. Also, when the information decoded from the encoded bitstream indicates that the affine motion compensation prediction mode is to be applied, the inter prediction unit 218 derives a motion vector in sub-block units based on the motion vectors of a plurality of adjacent blocks.
[0194] [Prediction control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal to the addition unit 208 as the prediction signal. Overall, the configurations, functions, and processes of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoder side may correspond to the configurations, functions, and processes of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoder side.
[0195] [Non-rectangular partitioning] Also in the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126 of the symbolization device (see FIG. 1), and also in the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218 of the decoding device (see FIG. 10), conventionally, a plurality of partitions (or a plurality of variable-size blocks, or a plurality of sub-blocks) obtained from the division of each block and from which motion information (for example, a plurality of motion vectors) is obtained are always rectangular as shown in FIG. 2. The inventors have discovered that generating a plurality of partitions having a non-rectangular shape such as a triangular shape can lead to improvements in image quality and coding efficiency according to the content of the image in the picture in various implementations. Hereinafter, various embodiments will be described in which at least one partition divided from an image block for the purpose of prediction has a non-rectangular shape. Note that these embodiments are equally applicable to the encoding device side (the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126) and to the decoding device side (the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218), and may be implemented in the encoding device of FIG. 1 or the decoding device of FIG. 10 or the like.
[0196] FIG. 11 is a flowchart showing an example of a process of dividing an image block into a plurality of partitions including at least a first partition having a non-rectangular shape (for example, a triangle) and a second partition, and further performing a process of encoding (or decoding) the image block as a reconfigured combination of the first partition and the second partition.
[0197] In step S1001, the image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. For example, as shown in FIG. 12, the image block may be divided from the upper left corner to the lower right corner of the image block to create a first partition and a second partition, both having a non-rectangular shape (e.g., a triangle). Alternatively, the image block may be divided from the upper right corner to the lower left corner of the image block to create a first partition and a second partition, both having a non-rectangular shape (e.g., a triangle). With reference to FIGS. 12 and 17 to 19, various examples of non-rectangular division will be described later.
[0198] In step S1002, the process predicts a first motion vector for the first partition and a second motion vector for the second partition. For example, predicting the first motion vector and the second motion vector may include selecting the first motion vector from a set of first motion vector candidates and selecting the second motion vector from a set of second motion vector candidates.
[0199] In step S1003, motion compensation processing is performed to obtain the first partition using the first motion vector derived in step S1002 above and to obtain the second partition using the second motion vector derived in step S1002 above.
[0200] In step S1004, prediction processing is performed on the image block as a (reconfigured) combination of the first partition and the second partition. The prediction processing includes boundary smoothing processing for smoothing the boundary between the first partition and the second partition. For example, the boundary smoothing processing involves weighting a plurality of first values of a plurality of boundary pixels predicted based on the first partition and a plurality of second values of a plurality of boundary pixels predicted based on the second partition. Various implementations of the boundary smoothing processing will be described later with reference to FIGS. 13, 14, 20, and 21A to 21D.
[0201] In step S1005, the process encodes or decodes the image block using one or more parameters including a first partition having a non-rectangular shape and partition parameters indicating dividing the image block into the second partition. As summarized in the table of FIG. 15, for example, the partition parameter (“first index value”) may encode, for example, the division direction applied to the division (e.g., from top left to bottom right or from top right to bottom left as shown in FIG. 12) and the first motion vector and the second motion vector derived in step S1002 described above, together. Details of such partition syntax operations with one or more parameters including the partition parameter will be described in detail later with reference to FIGS. 15, 16, and 22 to 25.
[0202] FIG. 17 is a flowchart showing a process 2000 of dividing an image block. In step S2001, the process divides an image into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. As shown in FIG. 12, the image block is divided into a first partition having a triangular shape and a second partition also having a triangular shape. There are many other examples where an image block is divided into a plurality of partitions including a first partition and a second partition, where at least the first partition has a non-rectangular shape. The non-rectangular shape may be a triangle, a trapezoid, or a polygon having at least five sides and angles.
[0203] For example, as shown in FIG. 18, an image block may be divided into two triangular-shaped partitions. The image block may be divided into more than two triangular-shaped partitions (e.g., three triangular-shaped partitions). The image block may be divided into a combination of one or more triangular-shaped partitions and one or more rectangular-shaped partitions. Alternatively, the image block may be divided into a combination of one or more triangular-shaped partitions and one or more polygon-shaped partitions.
[0204] Furthermore, as shown in FIG. 19, an image block may be divided into an L-shaped (polygon-shaped) partition and a rectangular-shaped partition. The image block may be divided into a pentagon (polygon) shaped partition and a triangular-shaped partition. The image block may be divided into a hexagon (polygon) shaped partition and a pentagon (polygon) shaped partition. Alternatively, the image block may be divided into a plurality of polygon-shaped partitions.
[0205] Referring again to FIG. 17, in step S2002, the process predicts a first motion vector for the first partition by, for example, selecting the first partition from a first set of motion vector candidates, and predicts a second motion vector for the second partition by, for example, selecting the second partition from a second set of motion vector candidates. For example, the first set of motion vector candidates may include a plurality of motion vectors of a plurality of partitions adjacent to the first partition, and the second set of motion vector candidates may include a plurality of motion vectors of a plurality of partitions adjacent to the second partition. The plurality of adjacent partitions may be one or both of a plurality of spatially adjacent partitions and a plurality of temporally adjacent partitions. Some examples of the plurality of spatially adjacent partitions include partitions located to the left, lower left, lower, lower right, right, upper right, upper, or upper left of the partition being processed. Some examples of the plurality of temporally adjacent partitions include a plurality of co-located partitions in a plurality of reference pictures of an image block.
[0206] In various implementations, the plurality of partitions adjacent to the first partition and the plurality of partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The first set of motion vector candidates may be the same as or different from the second set of motion vector candidates. Further, at least one of the first set of motion vector candidates and the second set of motion vector candidates may be the same as another third set of motion vector candidates prepared for the image block.
[0207] In some implementations, in step S2002, in response to a determination that the second partition, like the first partition, has a non-rectangular shape (e.g., triangular), process 2000 creates (for the non-rectangular shaped second partition) a second set of motion vector candidates that includes the motion vectors of a plurality of partitions adjacent to the second partition, excluding the first partition (i.e., excluding the motion vector of the first partition). On the other hand, in response to a determination that the second partition has a rectangular shape and is not the same as the first partition, process 2000 creates (for the rectangular shaped second partition) a second set of motion vector candidates that includes the motion vectors of a plurality of partitions adjacent to the second partition, including the first partition.
[0208] In step S2003, the process encodes or decodes the first partition using the first motion vector derived in step S2002 described above, and encodes or decodes the second partition using the second motion vector derived in step S2002 described above.
[0209] An image block splitting process such as process 2000 in FIG. 17 may be performed by an image encoding apparatus including, for example, a circuit and a memory connected to the circuit as shown in FIG. 1. The circuit, in operation, splits an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).
[0210] According to another embodiment, as shown in FIG. 1, an image encoding apparatus includes a division unit 102 that receives an original image in operation and divides it into a plurality of blocks, an addition unit 104 that receives a plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the addition unit 104 in operation and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bit stream, an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in a current picture in operation, and a prediction control unit 128 connected to memories 118 and 122, and is provided. The prediction control unit 128 divides a plurality of blocks into a plurality of partitions including a first partition and a second partition having a non-rectangular shape in operation (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).
[0211] According to another embodiment, an image decoding apparatus including a circuit and a memory connected to the circuit as shown in FIG. 10, for example, is provided. The circuit performs, in operation, dividing an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape (FIG. 17, step S2001), predicting a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decoding the first partition using the first motion vector and the second partition using the second motion vector (step S2003).
[0212] According to a further embodiment, as shown in FIG. 10, the image decoding apparatus receives a coded bitstream in operation, decodes it, and acquires a plurality of quantized transform coefficients. An entropy decoding unit 202, in operation, inverse quantizes the plurality of quantized transform coefficients to acquire a plurality of transform coefficients, and inverse transforms the plurality of transform coefficients to acquire a plurality of residuals. An inverse quantization unit 204 and an inverse transform unit 206, in operation, add the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to the plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks. An addition unit 208, and, in operation, an inter prediction unit 218 that generates a prediction of a current block based on a reference block in a decoded reference picture, and an intra prediction unit 216 that generates a prediction of the current block based on a decoded reference block in the current picture, and memories 210, 214 are provided, connected to a prediction control unit 220. The prediction control unit 220, in operation, divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (FIG. 17, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).
[0213] [Boundary Smoothing] As described above in FIG. 11, step S1004, according to various embodiments, performing prediction processing on an image block as a (reconstructed) combination of a first partition having a non-rectangular shape and a second partition may involve applying a boundary smoothing process along the boundary between the first partition and the second partition.
[0214] For example, FIG. 21B shows an example of a boundary smoothing process that involves weighting a plurality of first values of a plurality of first predicted boundary pixels based on a first partition and a plurality of second values of a plurality of second predicted boundary pixels based on a second partition.
[0215] FIG. 20 is a flowchart showing an overall boundary smoothing process 3000 involving weighting a plurality of first values of a plurality of first predicted boundary pixels based on a first partition and a plurality of second values of a plurality of second predicted boundary pixels based on a second partition. In step S3001, as shown in FIG. 21A or as shown in FIGS. 12, 18, and 19 described above, an image block is divided along a boundary into a first partition and a second partition, where at least the first partition has a non-rectangular shape.
[0216] In step S3002, a plurality of first values (e.g., color, luminance, transparency, etc.) of a pixel set of the first partition (the "plurality of boundary pixels" in FIG. 21A) along the boundary are first predicted using the information of the first partition. In step S3003, a plurality of second values of the (same) pixel set of the first partition along the boundary are second predicted using the information of the second partition. In some embodiments, at least one of the first prediction and the second prediction is an inter-prediction process that predicts the plurality of first values and the plurality of second values based on a reference partition in the encoded reference picture. Referring to FIG. 21D, in some implementations, the prediction process predicts a plurality of first values of all pixels of a first partition (the "first sample set") including a pixel set where the first partition and the second partition overlap, and predicts second values only for a pixel set (the "second sample set") where the first partition and the second partition overlap. In other implementations, at least one of the first prediction and the second prediction is an intra-prediction process that predicts the plurality of first values and the plurality of second values based on an encoded reference partition in the current picture. In some implementations, the prediction method used for the first prediction is different from the prediction method used for the second prediction. For example, the first prediction may include an inter-prediction process, or the second prediction may include an intra-prediction process. The information used for the first prediction of the plurality of first values or the second prediction of the plurality of second values may be a plurality of motion vectors, a plurality of intra-prediction directions, etc. of the first partition or the second partition.
[0217] In step S3004, the plurality of first values predicted using the first partition and the plurality of second values predicted using the second partition are weighted. In step S3005, the first partition is encoded or decoded using the weighted plurality of first values and the plurality of second values.
[0218] FIG. 21B shows an example of a boundary smoothing operation in which the first partition and the second partition overlap by (a maximum of) 5 pixels per row or per row. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 5. FIG. 21C shows another example of a boundary smoothing operation in which the first partition and the second partition overlap by (a maximum of) 3 pixels per row or per column. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 3.
[0219] FIG. 13 shows another example of a boundary smoothing operation in which the first partition and the second partition overlap by (a maximum of) 4 pixels per row or per column. That is, the number of pixel sets per row or per column for which a plurality of first values are predicted based on the first partition and a plurality of second values are predicted based on the second partition is at most 4. In the example shown, a plurality of weights of 1 / 8, 1 / 4, 3 / 4, and 7 / 8 may be applied to the plurality of first values of the 4 pixels in the set, respectively, and a plurality of weights of 7 / 8, 3 / 4, 1 / 4, and 1 / 8 may be applied to the plurality of second values of the 4 pixels in the set, respectively.
[0220] FIG. 14 further shows a plurality of examples of boundary smoothing operations where the first partition and the second partition overlap at 0 pixels of each row or column (i.e., they do not overlap), overlap at (maximum) 1 pixel of each row or column, and overlap at (maximum) 2 pixels of each row or column. In an example where the first partition and the second partition do not overlap, a plurality of zero weights are applied. In an example where the first partition and the second partition overlap at 1 pixel of each row or column, a weight of 1 / 2 may be applied to a plurality of first values of a plurality of pixels in a set predicted based on the first partition, and a weight of 1 / 2 may be applied to a plurality of second values of a plurality of pixels in a set predicted based on the second partition. In an example where the first partition and the second partition overlap at 2 pixels of each row or column, weights of 1 / 3 and 2 / 3 may be applied, respectively, to a plurality of first values of two pixels in a set predicted based on the first partition, and weights of 2 / 3 and 1 / 3 may be applied, respectively, to a plurality of second values of two pixels in a set predicted based on the second partition.
[0221] According to the plurality of embodiments described above, the number of a plurality of pixels in a set where the first partition and the second partition overlap is an integer. In other implementations, the number of overlapping pixels in a set may be, for example, a non-integer or a fraction. The plurality of weights applied to the plurality of first values and the plurality of second values of the pixel set may also be a fraction or an integer, depending on each application.
[0222] Boundary smoothing processing such as processing 3000 in FIG. 20 may be performed by an image encoding apparatus including a circuit and a memory connected to the circuit as shown in FIG. 1, for example. In operation, the circuit performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition (FIG. 20, step S3001). The boundary smoothing operation includes predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition (step S3002), second-predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).
[0223] According to another embodiment, as shown in FIG. 1, an image encoding apparatus includes a splitting unit 102 that receives an original image in operation and splits it into a plurality of blocks, an adding unit 104 that receives a plurality of blocks from the splitting unit and a plurality of predictions from a prediction control unit 128 in operation, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the adding unit 104 in operation and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bit stream, an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in the current picture in operation, and a prediction control unit 128 connected to memories 118 and 122, and is provided. The prediction control unit 128 performs a boundary smoothing operation (FIG. 20, step S3001) along a boundary between a first partition having a non-rectangular shape and a second partition, which are split from an image block, in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and encoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).
[0224] According to another embodiment, for example, an image decoding apparatus including a circuit and a memory connected to the circuit, as shown in FIG. 10, is provided. The circuit performs a boundary smoothing operation (FIG. 20, step S3001) along a boundary between a first partition having a non-rectangular shape, which is divided from an image block, and a second partition in operation. The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using the information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using the information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).
[0225] According to another embodiment, the image decoding apparatus shown in FIG. 10 includes an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantized transform coefficients, an inverse quantization unit 204 that inverse quantizes the plurality of quantized transform coefficients in operation to obtain a plurality of transform coefficients, and inverse-transforms the plurality of transform coefficients to obtain a plurality of residuals, an inverse transform unit 206, an addition unit 208 that adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and the plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter prediction unit 218 that generates a prediction of a current block based on a reference block in a decoded reference picture in operation, an intra prediction unit 216 that generates a prediction of the current block based on a decoded reference block in the current picture in operation, and a prediction control unit 220 connected to memories 210 and 214. The prediction control unit 220 performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape and a second partition, which are divided from an image block, in operation (FIG. 20, step S3001). The boundary smoothing operation includes first predicting a plurality of first values of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values of the pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values and the plurality of second values (step S3004), and decoding the first partition using the weighted plurality of first values and the weighted plurality of second values (step S3005).
[0226] [Entropy Encoding and Decoding Using Partition Parameter Syntax] As shown in FIG. 11 and step S1005, according to various embodiments, an image block divided into a first partition and a second partition having a non-rectangular shape may be encoded or decoded using one or more parameters including partition parameters indicating the non-rectangular division of the image block. In various embodiments, such partition parameters may encode, for example, the division direction applied to the division (e.g., from the upper left to the lower right, or from the upper right to the lower left, see FIG. 12) and the first motion vector and the second motion vector predicted in step S1002 as a whole, as will be described in more detail later.
[0227] FIG. 15 is a table of a plurality of sample partition parameters ("first index values") and a plurality of information sets each encoded as a whole by a plurality of partition parameters. The plurality of partition parameters ("first index values") range from 0 to 6 and encode, as a whole, the direction of dividing the image block into a first partition and a second partition, both of which are triangular (see FIG. 12), the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002). In particular, partition parameter 0 encodes that the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.
[0228] Partition parameter 1 encodes that the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 6 encodes that the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "fourth" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.
[0229] FIG. 22 is a flowchart showing a method 4000 performed on the encoder side. In step S4001, the process divides the image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on a partition parameter indicating the division. For example, as shown in FIG. 15 described above, the partition parameter may indicate the direction of dividing the image block (e.g., from the upper right corner to the lower left corner, or from the upper left corner to the lower right corner). In step S4002, the process encodes the first partition and the second partition. In step S4003, the process writes one or more parameters including the partition parameter into a bitstream that can be received and decoded so that the decoder side can obtain one or more parameters and perform the same prediction process on the first partition and the second partition on the decoder side (as performed on the encoder side). One or more parameters including the partition parameter encode various information such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used for dividing the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., either integrally or separately.
[0230] FIG. 23 is a flowchart showing a method 5000 performed on the decoder side. In step S5001, the process decodes one or more parameters from a bitstream including partition parameters indicating that the image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition. One or more parameters including the partition parameters decoded from the bitstream may include the non-rectangular shape of the first partition, the shape of the second partition, the division direction used for dividing the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., and various information required for the decoder side to perform the same prediction process as that performed on the encoder side may be encoded together or separately. In step S5002, the process 5000 divides the image block into a plurality of partitions based on the partition parameters decoded from the bitstream. In step S5003, the process decodes the first partition and the second partition, such as those divided from the image block.
[0231] FIG. 24 is a table of a sample table whose characteristics are similar to those described above in FIG. 15, and a plurality of information sets integrally encoded by a plurality of sample partition parameters ( "first index value") and a plurality of partition parameters respectively. In FIG. 24, the partition parameter ( "first index value") ranges from 0 to 6, and encodes the shapes of the first partition and the second partition divided from the image block, the direction of dividing the image block into the first partition and the second partition, the first motion vector predicted for the first partition (FIG. 11, step S1002), and the second motion vector predicted for the second partition (FIG. 11, step S1002) integrally. In particular, partition parameter 0 encodes that neither the first partition nor the second partition has a triangular shape, thus the division direction information is "N / A", the first motion vector information is "N / A", and the second motion vector information is "N / A".
[0232] Partition parameter 1 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper right corner to the lower left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition.The partition parameter 6 encodes that the first partition and the second partition are triangles, the splitting direction is from the upper left corner to the lower right corner, the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.
[0233] According to some implementations, a plurality of partition parameters (a plurality of index values) may be binarized according to a binarization method selected according to at least one or more parameter values. FIG. 16 shows an example of a binarization method for binarizing a plurality of index values (a plurality of partition parameter values).
[0234] FIG. 25 is a table of an example combination of a first parameter and a second parameter, where one of the first parameter and the second parameter is a partition parameter indicating dividing an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape. In this example, the partition parameter may be used to indicate dividing an image block without integrally encoding other information encoded by one or more of the other plurality of parameters.
[0235] In the first example in FIG. 25, the first parameter is used to indicate the image block size, and the second parameter is used as a partition parameter (flag) for indicating that at least one of the plurality of partitions divided from the image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, to indicate that 1) there is no triangular-shaped partition when the image block size is larger than 64×64, or 2) there is no triangular-shaped partition when the ratio of the width to the height of the image block is larger than 4 (e.g., 64×4).
[0236] In the second example of FIG. 25, the first parameter is used to indicate a prediction mode, and the second parameter is used as a partition parameter (flag) to indicate that at least one of a plurality of partitions divided from an image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that there is no triangular partition when the image block is encoded in the intra mode.
[0237] In the third example of FIG. 25, the first parameter is used as a partition parameter (flag) to indicate that at least one of a plurality of partitions divided from an image block has a triangular shape, and the second parameter is used to indicate a prediction mode. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the image block must be inter-encoded when at least one of a plurality of partitions divided from the image block has a triangular shape.
[0238] In the fourth example of FIG. 25, the first parameter indicates a motion vector of an adjacent block, and the second parameter is used as a partition parameter indicating a direction in which the image block is divided into two triangles. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the direction in which the image block is divided into two triangles is from the upper left corner to the lower right corner when the motion vector of the adjacent block is in an oblique direction.
[0239] In the fifth example of FIG. 25, the first parameter indicates an intra prediction direction of an adjacent block, and the second parameter is used as a partition parameter indicating a direction in which the image block is divided into two triangles. Such a combination of the first parameter and the second parameter may be used, for example, 1) to indicate that the direction in which the image block is divided into two triangles is from the upper right corner to the lower left corner when the intra prediction direction of the adjacent block is in an inverse oblique direction.
[0240] Multiple tables of one or more parameters including partition parameters, and as shown in FIGS. 15, 24, and 25, which information is encoded together or separately is presented only as multiple examples, and it should be understood that numerous other ways of encoding various information together or separately as part of the partition syntax operations described above are within the scope of the present disclosure. For example, the partition parameter may indicate that the first partition is a triangle, a trapezoid, or a polygon having at least five sides and angles. The partition parameter may indicate that the second partition has a non-rectangular shape, such as a triangle, a trapezoid, and a polygon having at least five sides and angles. The partition parameter may indicate one or more pieces of information about the partition, such as the non-rectangular shape of the first partition, the shape of the second partition (which may be non-rectangular or rectangular), the division direction applied to divide the image block into multiple partitions (e.g., from the upper left corner to the lower right corner of the image block, and from the upper right corner to the lower left corner of the image block). The partition parameter encodes further information together, such as the first motion vector of the first partition, the second motion vector of the second partition, the image block size, the prediction mode, the motion vectors of adjacent blocks, the intra prediction direction of adjacent blocks, etc. Alternatively, any of the further information may be encoded separately by one or more parameters other than the partition parameter.
[0241] A partition syntax operation such as the process 4000 in FIG. 22 may be performed by an image encoding apparatus including a circuit and a memory connected to the circuit, as shown in FIG. 1 for example. In operation, the circuit divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on partition parameters indicating division (FIG. 22, step S4001), encodes the first partition and the second partition (S4002), and performs a partition syntax operation including writing one or more parameters including the partition parameters to a bit stream (S4003).
[0242] According to another embodiment, an image encoding apparatus as shown in FIG. 1 includes, in operation, a division unit 102 that receives an original image and divides it into a plurality of blocks, an addition unit 104 that receives a plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128, subtracts each prediction from its corresponding block, and outputs a residual, a conversion unit 106 that performs a conversion on the plurality of residuals output from the addition unit 104 and outputs a plurality of conversion coefficients, a quantization unit 108 that quantizes the plurality of conversion coefficients in operation to generate a plurality of quantized conversion coefficients, an entropy encoding unit 110 that encodes the plurality of quantized conversion coefficients in operation to generate a bit stream, and an inter prediction unit 126 that generates a prediction of a current block based on a reference block in an encoded reference picture in operation, and an intra prediction unit 124 that generates a prediction of the current block based on an encoded reference block in the current picture in operation, and a prediction control unit 128 connected to memories 118 and 122. In operation, the prediction control unit 128 divides an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition based on partition parameters indicating division (FIG. 22, step S4001), and encodes the first partition and the second partition (step S4002). The entropy encoding unit 110 writes one or more parameters including the partition parameters to a bit stream in operation (step S4003).
[0243] According to another embodiment, for example, an image decoding apparatus including a circuit and a memory connected to the circuit as shown in FIG. 10 is provided. The circuit, in operation, reads one or more parameters including a partition parameter indicating to divide an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition from a bit stream (FIG. 23, step S5001), divides the image block into a plurality of partitions based on the partition parameter (S5002), and performs a partition syntax operation including decoding the first partition and the second partition (S5003).
[0244] According to a further embodiment, the image decoding apparatus shown in FIG. 10 includes, in operation, an entropy decoding unit 202 that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, an inverse quantization unit 204 that, in operation, inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inverse-transforms the plurality of transform coefficients to obtain a plurality of residuals, an inverse transform unit 206, an addition unit 208 that, in operation, adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 and the plurality of predictions output from the prediction control unit 220 to reconstruct a plurality of blocks, an inter prediction unit 218 that, in operation, generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit 216 that, in operation, generates a prediction of a current block based on a decoded reference block in a current picture, and a prediction control unit 220 connected to memories 210 and 214. In some implementations, the entropy decoding unit 202, in operation, cooperates with the prediction control unit 220 to read, from the bitstream, one or more parameters including a partition parameter indicating to partition an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (FIG. 23, step S5001), partitions the image block into the plurality of partitions based on the partition parameter (S5002), and decodes the first partition and the second partition (S5003).
[0245] According to a plurality of other examples, the inter prediction unit may perform the following processing.
[0246] All motion vector candidates included in the first motion vector candidate set may be a plurality of single prediction motion vectors. That is, the inter prediction unit may determine only a plurality of single prediction motion vectors as the plurality of motion vector candidates in the first motion vector candidate set.
[0247] The inter prediction unit may select only a plurality of single prediction motion vector candidates from the first motion vector candidate set.
[0248] Only a single prediction motion vector may be used for predicting small blocks. A dual prediction motion vector may be used for predicting large blocks. As an example, the prediction process may include determining the size of an image block. If it is determined that the size of the image block is greater than a threshold value, the prediction may include selecting a first motion vector from a first set of motion vector candidates, and the first set of motion vector candidates may include single prediction motion vectors and / or dual prediction motion vectors. If it is determined that the size of the image block is not greater than the threshold value, the prediction may include selecting a first motion vector from a first set of motion vector candidates, and the first set of motion vector candidates may include only a plurality of single prediction motion vectors.
[0249] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can generally be realized by an MPU (micro processing unit), a memory, and the like. Also, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (program) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memories. Note that it is also possible to realize each functional block by hardware (a dedicated circuit). Various combinations of hardware and software can be adopted.
[0250] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using a plurality of devices. Also, the processor that executes the above program may be singular or plural. That is, centralized processing or distributed processing may be performed.
[0251] Aspects of the present disclosure are not limited to the above embodiments, and various modifications are possible, and these are also included within the scope of the aspects of the present disclosure.
[0252] Furthermore, here, application examples of the moving image encoding method (image encoding method) or the moving image decoding method (image decoding method) shown in each of the above embodiments, and various systems for implementing such application examples will be described. Such a system may be characterized by having an image encoding device using the image encoding method, an image decoding device using the image decoding method, or an image encoding / decoding device having both. Other configurations of such a system can be appropriately changed as the case may be.
[0253] [Usage Example] FIG. 26 is a diagram showing the overall configuration of a suitable content supply system ex100 for realizing a content distribution service. The communication service providing area is divided into a desired size, and base stations ex106, ex107, ex108, ex109, ex110, which are fixed radio stations in the illustrated example, are installed in each cell.
[0254] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may be configured by connecting a combination of any of the above devices. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network or short-range wireless communication without passing through base stations ex106 to ex110. Further, the streaming server ex103 may be connected to devices such as a computer ex111, a game machine ex112, a camera ex113, home appliances ex114, and a smartphone ex115 via the Internet ex101 or the like. Also, the streaming server ex103 may be connected to terminals within a hotspot in an airplane ex117 via a satellite ex116.
[0255] Note that, instead of the base stations ex106 to ex110, a wireless access point, a hot spot, or the like may be used. Further, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.
[0256] The camera ex113 is a device capable of taking still images and video, such as a digital camera. Further, the smartphone ex115 is a smartphone device, a mobile phone, or a PHS (Personal Handy-phone System) or the like compatible with the mobile communication system standards called 2G, 3G, 3.9G, 4G, and in the future 5G.
[0257] The home appliance ex114 is a refrigerator or a device included in a household fuel cell cogeneration system or the like.
[0258] In the content supply system ex100, a terminal having a photographing function is connected to the streaming server ex103 through the base station ex106 or the like, enabling live distribution and the like. In live distribution, the terminal (the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, the smartphone ex115, and the terminal in the airplane ex117, etc.) may perform the encoding process described in each of the above embodiments on the still image or video content photographed by the user using the terminal, may multiplex the video data obtained by encoding with the audio data obtained by encoding the sound corresponding to the video, and may transmit the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to an aspect of the present disclosure.
[0259] On the one hand, the streaming server ex103 streams the transmitted content data to the requested client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117, etc., which can decode the encoded data. Each device that receives the distributed data may decode and play the received data. That is, each device may function as an image decoding device according to an aspect of the present disclosure.
[0260] [Distributed Processing] Also, the streaming server ex103 may be a plurality of servers or a plurality of computers that distribute, process, and record data. For example, the streaming server ex103 may be implemented by a CDN (Content Delivery Network), and content delivery may be realized by a network connecting a large number of edge servers distributed around the world. In a CDN, an edge server physically close to the client can be dynamically assigned according to the client. By caching and distributing the content to the edge server, the delay can be reduced. In addition, when some types of errors occur or the communication state changes due to an increase in traffic, etc., the processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or the part of the network with a failure can be bypassed to continue the distribution, so high-speed and stable distribution can be realized.
[0261] Moreover, not limited to the distributed processing of the distribution itself, the encoding process of the captured data may be performed on each terminal, on the server side, or shared between them. As an example, in general encoding processing, the processing loop is performed twice. In the first loop, the complexity of the image in units of frames or scenes, or the amount of code is detected. In the second loop, processing is performed to improve the encoding efficiency while maintaining the image quality. For example, if the terminal performs the first encoding process and the server that receives the content performs the second encoding process, it is possible to improve the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a requirement to receive and decode almost in real time, the other terminals can receive and play the already encoded data processed by the terminal, so that more flexible real-time distribution is also possible.
[0262] As another example, cameras such as ex113 extract feature amounts (amounts of features or characteristics) from images, compress the data related to the feature amounts as metadata, and transmit it to the server. The server performs compression according to the meaning of the image (or the importance of the content), for example, by judging the importance of the object from the feature amounts and switching the quantization accuracy. Feature amount data is particularly effective in improving the accuracy and efficiency of motion vector prediction during re-compression on the server. Also, simple encoding such as VLC (Variable Length Coding) may be performed on the terminal, and encoding with a large processing load such as CABAC (Context Adaptive Binary Arithmetic Coding) may be performed on the server.
[0263] As yet another example, in a stadium, a shopping mall, or a factory, etc., there may be a plurality of video data in which substantially the same scene is captured by a plurality of terminals. In this case, using the plurality of terminals that performed the shooting, and other terminals and servers that did not perform the shooting as necessary, encoding processing is respectively assigned and distributed processing is performed, for example, in units of GOP (Group of Picture), picture units, or tile units obtained by dividing the picture. Thereby, the delay can be reduced and more real-time performance can be realized.
[0264] Since the plurality of video data is of substantially the same scene, the server may manage and / or give instructions so that the video data captured by each terminal can be referenced to each other. Also, the server may receive the encoded data from each terminal and change the reference relationship among the plurality of data, or correct or replace the picture itself and re-encode it. Thereby, a stream with improved quality and efficiency of each piece of data can be generated.
[0265] Furthermore, the server may perform transcoding to change the encoding method of the video data and then distribute the video data. For example, the server may convert an MPEG-based encoding method to a VP-based (e.g., VP9) method, or convert H.264 to H.265, etc.
[0266] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, hereinafter, descriptions such as "server" or "terminal" are used as the subject performing the process, but part or all of the processes performed by the server may be performed by the terminal, or part or all of the processes performed by the terminal may be performed by the server. Also, regarding these, the same applies to the decoding process.
[0267] [3D, Multi-angle] There is an increasing trend to integrate and use different scenes captured by terminals such as a plurality of cameras ex113 and / or smartphones ex115 that are substantially synchronized with each other, or images or videos of the same scene captured from different angles. The videos captured by each terminal can be integrated based on the relative positional relationship between the terminals obtained separately, or the regions where the feature points included in the videos match.
[0268] The server may not only encode two-dimensional moving images, but also automatically encode still images based on scene analysis of the moving images or at a time specified by the user, and transmit them to the receiving terminal. When the server can further obtain the relative positional relationship between the shooting terminals, it can generate the three-dimensional shape of the scene based on not only two-dimensional moving images but also videos shot from different angles of the same scene. The server may separately encode the three-dimensional data generated by a point cloud or the like, or select or reconstruct the video to be transmitted to the receiving terminal from the videos shot by multiple terminals based on the results of recognizing or tracking a person or an object using the three-dimensional data.
[0269] In this way, the user can arbitrarily select each video corresponding to each shooting terminal to enjoy the scene, or enjoy the content obtained by cutting out the video of the selected viewpoint from the three-dimensional data reconstructed using multiple images or videos. Furthermore, the sound is also collected from a plurality of different angles together with the video, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.
[0270] In recent years, content that associates the real world with the virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server may create viewpoint images for the right eye and the left eye respectively, and perform encoding that allows reference between each viewpoint video by Multi-View Coding (MVC) or the like, or encode them as separate streams without referring to each other. When decoding the separate streams, it is advisable to synchronize and play them so that a virtual three-dimensional space is reproduced according to the user's viewpoint.
[0271] In the case of an AR image, the server may superimpose virtual object information in the virtual space on the camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and create superimposed data by smoothly connecting them. Alternatively, in addition to requesting virtual object information, the decoding device may transmit the movement of the user's viewpoint to the server. The server may create superimposed data in accordance with the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data typically has an α value indicating transparency in addition to RGB, and the server may set the α value of the portion other than the object created from the three-dimensional data to 0 or the like and encode it in a state where the portion is transparent. Alternatively, the server may set an RGB value of a predetermined value as a background like a chroma key and generate data in which the portion other than the object is the background color. The RGB value of the predetermined value may be determined in advance.
[0272] Similarly, the decoding process of the distributed data may be performed by the client (e.g., a terminal), on the server side, or shared between them. As an example, a certain terminal may once send a reception request to the server, receive the content corresponding to the request by another terminal, perform the decoding process, and the decoded signal may be transmitted to a device having a display. By dispersing the process regardless of the performance of the communicable terminal itself and selecting appropriate content, it is possible to reproduce data with good image quality. Also, as another example, while receiving large-size image data on a TV or the like, a part of the area such as a tile in which the picture is divided may be decoded and displayed on the personal terminal of the viewer. Thereby, while sharing the overall image, it is possible to check the area of one's own field of responsibility or the area to be confirmed in more detail at hand.
[0273] In a situation where multiple short-range, medium-range, or long-range wireless communications inside and outside the house are available, it may be possible to seamlessly receive content by using a delivery system standard such as MPEG-DASH. The user may freely select and switch in real time between the user's terminal, a decoding device or display device such as a display arranged inside and outside the house. Also, using the user's own location information and the like, decoding can be performed while switching between the terminal for decoding and the terminal for display. As a result, while the user is moving to a destination, it becomes possible to map and display information on a part of the wall surface or ground of an adjacent building in which a displayable device is embedded. Also, based on the ease of access to the encoded data on the network, such as the encoded data being cached in a server that can be accessed from the receiving terminal in a short time, or being copied to an edge server in a content delivery service, it is also possible to switch the bitrate of the received data.
[0274] [Scalable Encoding] Regarding the switching of content, it will be described using a scalable stream that is compression-encoded by applying the moving image encoding method shown in each of the above embodiments shown in FIG. 27. The server may have a plurality of streams with the same content but different qualities as individual streams, but by taking advantage of the characteristics of a temporally / spatially scalable stream realized by encoding in layers as shown in the figure, it may be configured to switch content. That is, by determining which layer to decode according to internal factors such as performance and external factors such as the state of the communication band on the decoding side, the decoding side can freely switch between low-resolution content and high-resolution content for decoding. For example, when the user wants to watch the continuation of a video that was being viewed on the smartphone ex115 while moving, on a device such as an Internet TV after returning home, for example, the device only needs to decode the same stream to a different layer, so the burden on the server side can be reduced.
[0275] Furthermore, as described above, pictures are encoded layer by layer. In addition to the configuration that realizes scalability in the enhancement layer above the base layer, the enhancement layer may include meta information based on statistical information of the image or the like. The decoding side may generate high-quality content by super-resolving the picture of the base layer based on the meta information. The super-resolution may improve the signal-to-noise ratio while maintaining and / or enlarging the resolution. The meta information includes information for specifying linear or non-linear filter coefficients for use in super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used in super-resolution processing, and the like.
[0276] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of an object in the image or the like. The decoding side decodes only a part of the area by selecting the tile to be decoded. Further, by storing the attributes of the object (such as a person, a car, a ball, etc.) and the position in the video (such as the coordinate position in the same image) as meta information, the decoding side can specify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 28, the meta information may be stored using a data storage structure different from the pixel data, such as the SEI (supplemental enhancement information) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.
[0277] The meta information may be stored in a unit composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. The decoding side can obtain the time when a specific person appears in the video, etc., and by combining the picture unit information and the time information, can specify the picture in which the object exists and determine the position of the object in the picture.
[0278] [Optimization of Web Page] FIG. 29 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 30 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 29 and 30, a web page may include a plurality of link images that are links to image content, and the appearance may differ depending on the device being viewed. When a plurality of link images are visible on the screen, until the user explicitly selects a link image, or until the link image approaches the vicinity of the center of the screen or the entire link image enters the screen, the display device (decoding device) may display a still image or an I picture that each content has as a link image, or may display a video such as a gif animation with a plurality of still images or I pictures, etc., or may receive only the base layer, decode and display the video.
[0279] When a link image is selected by the user, the display device performs decoding, for example, with the base layer having the highest priority. If there is information indicating that the HTML constituting the web page is scalable content, the display device may decode up to the enhancement layer. Further, in order to ensure real-time performance, before being selected or when the communication bandwidth is very strict, the display device can reduce the delay between the decoding time and the display time of the first picture (the delay from the start of content decoding to the start of display) by decoding and displaying only forward-reference pictures (I pictures, P pictures, B pictures with only forward reference). Furthermore, the display device may deliberately ignore the reference relationship of the pictures, roughly decode all B pictures and P pictures with forward reference, and perform normal decoding as the received pictures increase over time.
[0280] [Autonomous Driving] Also, when transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for autonomous driving or driving assistance of a vehicle, the receiving terminal may receive, in addition to the image data belonging to one or more layers, weather or construction information, etc. as meta information, and decode them in association with each other. Note that the meta information may belong to a layer or may simply be multiplexed with the image data.
[0281] In this case, since vehicles, drones, airplanes, etc. including the receiving terminal move, the receiving terminal can realize seamless reception and decoding by transmitting the position information of the receiving terminal while switching between the base stations ex106 to ex110. Further, the receiving terminal can dynamically switch how much meta information to receive or how much to update the map information according to the user's selection, the user's situation, and / or the state of the communication band.
[0282] In the content supply system ex100, the client can receive, decode, and reproduce the encoded information transmitted by the user in real time.
[0283] [Delivery of Personal Content] Also, in the content supply system ex100, not only high-quality and long-duration content by video delivery providers but also unicast or multicast delivery of low-quality and short-duration content by individuals is possible. Such personal content is expected to increase in the future. In order to make personal content into better content, the server may perform an editing process and then an encoding process. This can be realized, for example, using the following configuration.
[0284] During shooting in real time or accumulating and after shooting, the server performs recognition processing such as shooting error, scene search, semantic analysis, and object detection on the original picture data or encoded data. Then, based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes unimportant scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes the edges of objects, or changes the color tone for editing. The server encodes the edited data based on the editing results. Also, it is known that if the shooting time is too long, the viewing rate will decrease. The server may automatically clip not only unimportant scenes but also scenes with little movement within a specific time range according to the shooting time so that the content is within that time range, based on the image processing results. Or, the server may generate and encode a digest based on the result of the semantic analysis of the scene.
[0285] In some cases, personal content may contain elements that, as they are, would infringe on copyright, moral rights of the author, or the right of portrait, etc., and there may be inconvenient situations for individuals, such as the sharing scope exceeding the intended scope. Therefore, for example, the server may intentionally change the image of a person's face in the peripheral part of the screen or inside a house to an out-of-focus image and then encode it. Furthermore, the server may recognize whether a face of a person different from the pre-registered person appears in the image to be encoded, and if it does, perform processing such as applying a mosaic to the face part. Or, as pre-processing or post-processing of encoding, the user may specify a person or background area that the user wants to process the image from the perspective of copyright, etc. The server may perform processing such as replacing the specified area with another video or blurring the focus. For a person, in a moving image, the person can be tracked and the video of the face part of the person can be replaced.
[0286] Since viewing personal content with a small amount of data requires strong real-time performance, depending on the bandwidth, the decoding device may first receive the base layer with the highest priority and perform decoding and playback. During this time, the decoding device may receive the enhancement layer and, when the playback is looped or played two or more times, play back high-quality video including the enhancement layer. For a stream with scalable encoding like this, the video is rough when not selected or at the beginning of viewing, but it can provide an experience where the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream played for the first time and a second stream encoded with reference to the first video are configured as one stream.
[0287] [Other implementation and application examples] Also, these encoding or decoding processes are generally processed in the LSIex500 that each terminal has. The LSI (large scale integration circuitry) ex500 (see Fig. 26) may be a one-chip configuration or a configuration consisting of multiple chips. Note that software for video encoding or decoding may be incorporated into some recording medium (such as a CD-ROM, flexible disk, or hard disk) that can be read by a computer ex111 or the like, and the encoding or decoding process may be performed using the software. Further, when the smartphone ex115 has a camera, the video data acquired by the camera may be transmitted. The video data at this time may be data encoded by the LSIex500 that the smartphone ex115 has.
[0288] Note that the LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the content encoding method or has the ability to execute a specific service. If the terminal does not support the content encoding method or does not have the ability to execute a specific service, the terminal may download a codec or application software and then acquire and play the content.
[0289] In addition, not limited to the content supply system ex100 via the Internet ex101, at least one of the moving image encoding device (image encoding device) or the moving image decoding device (image decoding device) of the above-described embodiments can be incorporated into a digital broadcast system. Since multiplexed data in which video and audio are multiplexed is transmitted and received by riding on a broadcast radio wave using a satellite or the like, there is a difference in that it is more suitable for multicast than the unicast-oriented configuration of the content supply system ex100, but the same application is possible for the encoding process and the decoding process.
[0290] [Hardware Configuration] FIG. 31 is a diagram showing further details of the smartphone ex115 shown in FIG. 26. Further, FIG. 32 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying data obtained by decoding videos captured by the camera unit ex465 and videos received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing captured videos or still images, recorded audio, received videos or still images, encoded data such as emails, or decoded data, and a slot unit ex464 which is an interface unit with the SIM ex468 for identifying the user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.
[0291] A main control unit ex460 capable of comprehensively controlling the display unit ex458, the operation unit ex466, etc., a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / demultiplexing unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 are connected via a synchronization bus ex470.
[0292] When the power key is turned on by the user's operation, the power supply circuit unit ex461 activates the smartphone ex115 to an operable state and supplies power to each unit from the battery pack.
[0293] The smartphone ex115 performs processes such as calls and data communications based on the control of the main control unit ex460 having a CPU, ROM, RAM, etc. During a call, the voice signal picked up by the voice input unit ex456 is converted into a digital voice signal by the voice signal processing unit ex454, subjected to spread spectrum processing by the modulation / demodulation unit ex452, and subjected to digital-to-analog conversion processing and frequency conversion processing by the transmission / reception unit ex451, and the resulting signal is transmitted via the antenna ex450. Also, received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, subjected to inverse spread spectrum processing by the modulation / demodulation unit ex452, converted into an analog voice signal by the voice signal processing unit ex454, and then output from the voice output unit ex457. In the data communication mode, text, still images, or video data can be sent under the control of the main control unit ex460 via the operation input control unit ex462 based on operations such as those of the operation unit ex466 of the main body unit. Similar transmission and reception processes are performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image encoding method shown in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The voice signal processing unit ex454 encodes the voice signal picked up by the voice input unit ex456 while a video or still image is being captured by the camera unit ex465, and sends the encoded voice data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and the encoded voice data in a predetermined manner, performs modulation processing and conversion processing by the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits it via the antenna ex450. The predetermined manner may be determined in advance.
[0294] When receiving video attached to an email or chat, or video linked to a web page, etc., in order to decode the multiplexed data received via the antenna ex450, the multiplexing / demultiplexing unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the moving image encoding method shown in each of the above embodiments, and the video or still image included in the linked moving image file is displayed from the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. Since real-time streaming is becoming increasingly popular, depending on the user's situation, it may not be socially appropriate to play audio. Therefore, as an initial value, it is desirable to have a configuration that plays only video data without playing the audio signal, and the audio may be played synchronously only when the user performs an operation such as clicking on the video data.
[0295] Also, although the smartphone ex115 has been described as an example here, as the terminal, in addition to a transceiver type terminal having both an encoder and a decoder, other implementation forms such as a transmitting terminal having only an encoder and a receiving terminal having only a decoder are conceivable. In the digital broadcast system, it has been described as receiving or transmitting multiplexed data in which audio data is multiplexed with video data. However, in the multiplexed data, character data related to the video etc. may be multiplexed in addition to the audio data. Also, instead of the multiplexed data, the video data itself may be received or transmitted.
[0296] Although the main control unit ex460 including the CPU has been described as controlling the encoding or decoding process, many types of terminals are often equipped with a GPU. Therefore, a configuration may be adopted in which a wide area is processed in a batch by taking advantage of the performance of the GPU using a memory shared by the CPU and the GPU or a memory whose address is managed so as to be commonly used. As a result, the encoding time can be shortened, real-time performance can be ensured, and low latency can be achieved. In particular, it is efficient to perform the processes of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transform / quantization in units such as pictures using the GPU instead of the CPU in a batch.
Claims
1. A circuit, and a memory connected to the circuit, wherein the circuit, in operation, selects, from a set of motion vector candidates, a first motion vector of a first partition having a non-rectangular shape in an image block; selects, from the set of motion vector candidates, a second motion vector of a second partition having a non-rectangular shape in the image block; encodes the first partition with the selected first motion vector; encodes the second partition with the selected second motion vector; and only a single-prediction motion vector is selected from the set of motion vector candidates without any condition related to the size of the image block, An image encoding device.
2. A circuit, and a memory connected to the circuit, wherein the circuit, in operation, selects, from a set of motion vector candidates, a first motion vector of a first partition having a non-rectangular shape in an image block; selects, from the set of motion vector candidates, a second motion vector of a second partition having a non-rectangular shape in the image block; decodes the first partition with the selected first motion vector; decodes the second partition with the selected second motion vector; and only a single-prediction motion vector is selected from the set of motion vector candidates without any condition related to the size of the image block, An image decoding device.
3. A circuit, and a memory connected to the circuit, wherein the circuit, in operation, generates parameters for causing a decoding device to execute partition processing; includes the parameters in a bitstream; wherein the partition processing includes: selecting a first motion vector of a first partition having a non-rectangular shape in an image block from a set of motion vector candidates; selecting a second motion vector of a second partition having a non-rectangular shape in the image block from the set of motion vector candidates; decoding the first partition with the selected first motion vector; decoding the second partition with the selected second motion vector; and only a single-prediction motion vector is selected from the set of motion vector candidates without any condition related to the size of the image block, Bit stream generation device.
Citation Information
Patent Citations
Image processing device and image processing method
JP2012023597A
Coding device and coding method
JP2015216632A
Method for decomposing a video sequence frame
US20080101707A1
Method for Determining a Corner Video Part of a Partition of a Video Coding Block
US20160234503A1
Image decoding apparatus, image decoding method and image encoding apparatus
WO2013047805A1
Cited By
Image encoding device, image decoding device, and bit stream generating device
JP2025143441A