Image encoding device, image decoding device, and bitstream generating device

The video coding technology optimizes interprediction by predicting pixel sets for image block partitions using single motion vectors and weighting overlapping pixels, addressing inefficiencies in existing technologies and enhancing coding efficiency and speed.

JP7675143B2Active Publication Date: 2025-05-12PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023151749
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-07
Filing Date
2023-09-19
Publication Date
2025-05-12
Estimated Expiration
2039-09-03

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in optimizing interprediction functions for constructing predictions of the current frame based on reference frames, particularly in handling the increasing amount of digital video data.

Method used

The proposed solution involves an image encoding device and method that predicts pixel sets for partitions of an image block using single predicted motion vectors, weights these pixel sets for overlapping portions, and encodes these partitions using weighted pixels, thereby optimizing the interprediction process.

Benefits of technology

This approach improves coding efficiency, simplifies the encoding and decoding processes, and enhances the speed of these processes by efficiently utilizing motion vectors and weighted pixels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675143000004
    Figure 0007675143000004
  • Figure 0007675143000005
    Figure 0007675143000005
  • Figure 0007675143000006
    Figure 0007675143000006
Patent Text Reader

Abstract

To provide an image encoder which can suppress a memory amount.SOLUTION: An image encoder includes a circuit and a memory coupled to the circuit. The circuit, in operation, predicts a first set of pixels for a first partition of an image block with a first motion vector being a single prediction motion vector, and predicts a second set of pixels for a second partition of the image block with a second motion vector being a single prediction motion vector. The first set of pixels and the second set of pixels are weighted for a plurality of pixels of a first portion where the first partition and the second partition overlap each other. The first motion vector and the second motion vector are stored in the memory as both prediction motion vectors and as the motion vector of the first portion. The first partition is encoded using at least the weighted pixels of the first portion.SELECTED DRAWING: Figure 68
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to video coding, and more particularly to systems, components, and methods in video encoding and decoding for performing inter prediction functions that build a prediction of a current frame based on a reference frame. [Background technology]

[0002] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). With this progress, there is a constant need to provide improvements and optimizations in video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure relates to further advances, improvements and optimizations in video coding, especially in the inter prediction function that builds a prediction of a current frame based on a reference frame. Summary of the Invention

[0003] An image encoding device according to one embodiment of the present disclosure includes a circuit and a memory connected to the circuit, wherein the circuit, in operation, predicts a first set of pixels for a first partition of an image block using a first motion vector that is a uni-predictive motion vector, predicts a second set of pixels for a second partition of the image block using a second motion vector that is a uni-predictive motion vector, weights the first set of pixels and the second set of pixels for a first portion where the first partition and the second partition overlap, stores the first motion vector and the second motion vector in the memory as bi-predictive motion vectors and as motion vectors for the first portion, and encodes the first partition using at least the weighted pixels of the first portion.

[0004] An image decoding device according to one embodiment of the present disclosure includes a circuit and a memory connected to the circuit, wherein the circuit, in operation, predicts a first set of pixels for a first partition of an image block using a first motion vector that is a uni-predictive motion vector, predicts a second set of pixels for a second partition of the image block using a second motion vector that is a uni-predictive motion vector, weights the first set of pixels and the second set of pixels for a first portion where the first partition and the second partition overlap, stores the first motion vector and the second motion vector in the memory as bi-predictive motion vectors and as motion vectors for the first portion, and decodes the first partition using at least the weighted plurality of pixels of the first portion.

[0005] According to one aspect of the present disclosure, there is provided a bitstream generating apparatus comprising: a circuit; and a memory connected to the circuit, the circuit being operable to: ,single Using the first motion vector, which is the predicted motion vector, The first set of pixels for the first partition of the image block is prediction And, Using the second motion vector, which is the predicted motion vector, A second set of pixels for a second partition of the image block is prediction death For a plurality of pixels in a first portion where the first partition and the second partition overlap, the first pixel set and the second pixel set are of Weighting death , the first motion vector and the second motion vector of stored in the memory as a bi-predictive motion vector and as the motion vector of the first portion. And heavy using at least the first portion of the plurality of pixels located encoding the first partition and generating a bitstream including information used to encode the image block using the first partition; .

[0006] In one embodiment, an image encoding device includes a circuit and a memory coupled to the circuit. The circuit, in an operation, predicts a first set of pixels for a first partition of a current picture using a first motion vector and predicts a second set of pixels for a second partition of the current picture using a second motion vector. Weights the first set of pixels and the second set of pixels for a first portion where the first partition and the second partition overlap. Stores one of the first motion vector and the second motion vector as a motion vector for the first portion in the memory based on a first condition. Stores both the first motion vector and the second motion vector as a motion vector for the first portion in the memory based on a second condition different from the first condition. Encodes the first partition using at least the weighted pixels of the first portion.

[0007] In one aspect, an image encoding device includes: a division unit that, in an operation, receives an original picture and divides it into a plurality of blocks; a first adder that, in an operation, receives the plurality of blocks from the division unit and a plurality of predictions from a prediction control unit, and subtracts each prediction from a corresponding block to output a residual; a transformation unit that, in an operation, performs a transformation on the plurality of residuals output from the first adder unit to output a plurality of transformation coefficients; a quantization unit that, in an operation, quantizes the plurality of transformation coefficients to generate a plurality of quantized transformation coefficients; an entropy coding unit that, in an operation, encodes the plurality of quantized transformation coefficients to generate a bitstream; an inverse quantization transformation unit that, in an operation, dequantizes the plurality of quantized transformation coefficients to obtain the plurality of transformation coefficients and inversely transforms the plurality of transform coefficients to obtain the plurality of residuals; a second adder that, in an operation, adds the plurality of residuals output from the inverse quantization transformation unit and the plurality of predictions output from the prediction control unit to reconstruct the plurality of blocks; and the prediction control unit connected to an inter prediction unit, an intra prediction unit, and a memory. The inter prediction unit, in an operation, generates a prediction of a current block based on a reference block in a coded reference picture. The intra prediction unit, in an operation, generates a prediction of a current block based on a coded reference block in the current picture. The inter prediction unit, in an operation, predicts a first set of pixels for a first partition of the current picture using a first motion vector and predicts a second set of pixels for a second partition of the current picture using a second motion vector. The inter prediction unit, in an operation, weights the first set of pixels and the second set of pixels for a plurality of pixels in a first portion where the first partition and the second partition overlap. Based on a first condition, one of the first motion vector and the second motion vector is stored in the memory as a motion vector of the first portion. Based on a second condition different from the first condition, both the first motion vector and the second motion vector are stored in the memory as motion vectors of the first portion. The first partition is encoded using at least the weighted plurality of pixels in the first portion.

[0008] In one embodiment, an image coding method includes predicting a first set of pixels for a first partition of a current picture using a first motion vector and predicting a second set of pixels for a second partition of the current picture using a second motion vector. The first set of pixels and the second set of pixels are weighted for a first portion of pixels where the first partition and the second partition overlap. Based on a first condition, one of the first motion vector and the second motion vector are stored in a memory as a motion vector for the first portion. Based on a second condition different from the first condition, both the first motion vector and the second motion vector are stored in the memory as motion vectors for the first portion. The first partition is coded using at least the weighted pixels of the first portion.

[0009] In one embodiment, an image decoding device includes a circuit and a memory connected to the circuit. The circuit, in an operation, predicts a first set of pixels for a first partition of a current picture using a first motion vector and predicts a second set of pixels for a second partition of the current picture using a second motion vector. For a plurality of pixels in a first portion where the first partition and the second partition overlap, weights the first set of pixels and the second set of pixels. Based on a first condition, one of the first motion vector and the second motion vector is stored in the memory as a motion vector for the first portion. Based on a second condition different from the first condition, both the first motion vector and the second motion vector are stored in the memory as motion vectors for the first portion. The first partition is decoded using at least the weighted plurality of pixels in the first portion.

[0010] In one aspect, the present invention includes an entropy decoding unit that receives and decodes an encoded bitstream to obtain a plurality of quantized transform coefficients, an inverse quantization transform unit that inverse quantizes the plurality of quantized transform coefficients to obtain a plurality of transform coefficients, and inversely transforms the plurality of transform coefficients to obtain a plurality of residuals, an adder that adds the plurality of residuals output from the inverse quantization transform unit and a plurality of predictions output from a prediction control unit to reconstruct a plurality of blocks, and the prediction control unit is connected to an inter prediction unit, an intra prediction unit, and a memory. The inter prediction unit in operation generates a prediction of a current block based on a reference block in a decoded reference picture. The intra prediction unit in operation generates a prediction of a current block based on a decoded reference block in the current picture. The inter prediction unit in operation predicts a first set of pixels for a first partition of the current picture using a first motion vector and predicts a second set of pixels for a second partition of the current picture using a second motion vector. The first and second pixel sets are weighted for a plurality of pixels in a first portion where the first and second partitions overlap. One of the first and second motion vectors is stored in the memory as a motion vector for the first portion based on a first condition. Both the first and second motion vectors are stored in the memory as motion vectors for the first portion based on a second condition different from the first condition. The first partition is decoded using at least the weighted plurality of pixels in the first portion.

[0011] In one embodiment, an image decoding method includes predicting a first set of pixels for a first partition of a current picture using a first motion vector and predicting a second set of pixels for a second partition of the current picture using a second motion vector. The first set of pixels and the second set of pixels are weighted for a first portion of pixels where the first partition and the second partition overlap. Based on a first condition, one of the first motion vector and the second motion vector are stored in a memory as a motion vector for the first portion. Based on a second condition different from the first condition, both of the first motion vector and the second motion vector are stored in the memory as motion vectors for the first portion. The first partition is decoded using at least the weighted pixels of the first portion.

[0012] Some implementations of the embodiments of the present disclosure may improve encoding efficiency, simplify the encoding / decoding process, increase the encoding / decoding process speed, and / or efficiently select appropriate components / operations used in encoding and decoding, such as appropriate filters, block sizes, motion vectors, reference pictures, reference blocks, etc.

[0013] Further advantages and benefits of certain aspects of the present disclosure will become apparent from the specification and drawings, and while such advantages and / or benefits may be obtained by various embodiments and features described in the specification and drawings, not all of them necessarily need to be provided in order to obtain one or more advantages and / or benefits.

[0014] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, a recording medium, or any combination thereof. [Brief description of the drawings]

[0015] [Figure 1] FIG. 1 is a block diagram showing a functional configuration of an encoding device according to an embodiment. [Diagram 2] FIG. 2 is a flowchart showing an example of the overall encoding process performed by the encoding device. [Diagram 3] FIG. 3 is a conceptual diagram showing an example of block division. [Figure 4A] FIG. 4A is a conceptual diagram showing an example of a slice configuration. [Figure 4B] FIG. 4B is a conceptual diagram showing an example of a tile configuration. [Figure 5A] FIG. 5A is a table showing the transform basis functions that correspond to various transform types. [Figure 5B] FIG. 5B is a conceptual diagram showing an example of SVT (Spatially Varying Transform). [Figure 6A] FIG. 6A is a conceptual diagram showing an example of the shape of a filter used in an adaptive loop filter (ALF). [Figure 6B] FIG. 6B is a conceptual diagram showing another example of the shape of the filter used in the ALF. [Figure 6C] FIG. 6C is a conceptual diagram showing another example of the shape of the filter used in the ALF. [Figure 7] FIG. 7 is a block diagram showing an example of a detailed configuration of a loop filter unit functioning as a DBF (deblocking filter). [Figure 8] FIG. 8 is a conceptual diagram showing an example of a deblocking filter having symmetric filter characteristics with respect to block boundaries. [Figure 9] FIG. 9 is a conceptual diagram for explaining block boundaries on which deblocking filter processing is performed. [Figure 10] FIG. 10 is a conceptual diagram showing an example of the Bs value. [Figure 11] FIG. 11 is a flowchart illustrating an example of processing performed in the prediction processing unit of the encoding device. [Figure 12] FIG. 12 is a flowchart showing another example of the process performed in the prediction processing unit of the encoding device. [Figure 13]FIG. 13 is a flowchart showing another example of the process performed in the prediction processing unit of the encoding device. [Figure 14] FIG. 14 is a conceptual diagram showing an example of 67 intra prediction modes in intra prediction according to the embodiment. [Figure 15] FIG. 15 is a flowchart showing an example of the flow of basic inter prediction processing. [Figure 16] FIG. 16 is a flowchart showing an example of motion vector derivation. [Figure 17] FIG. 17 is a flowchart showing another example of motion vector derivation. [Figure 18] FIG. 18 is a flowchart showing another example of motion vector derivation. [Figure 19] FIG. 19 is a flowchart showing an example of inter prediction in the normal inter mode. [Figure 20] FIG. 20 is a flowchart showing an example of inter prediction in the merge mode. [Figure 21] FIG. 21 is a conceptual diagram for explaining an example of a motion vector derivation process in the merge mode. [Figure 22] FIG. 22 is a flowchart showing an example of a frame rate up conversion (FRUC) process. [Figure 23] FIG. 23 is a conceptual diagram for explaining an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. [Figure 24] FIG. 24 is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. [Figure 25A] FIG. 25A is a conceptual diagram for explaining an example of derivation of a motion vector for each sub-block based on motion vectors of a plurality of adjacent blocks. [Figure 25B] FIG. 25B is a conceptual diagram for explaining an example of derivation of a motion vector for each sub-block in the affine mode having three control points. [Figure 26A] FIG. 26A is a conceptual diagram for explaining the affine merge mode. [Figure 26B] FIG. 26B is a conceptual diagram for explaining an affine merge mode having two control points. [Figure 26C] FIG. 26C is a conceptual diagram for explaining an affine merge mode having three control points. [Figure 27] FIG. 27 is a flowchart showing an example of a process in the affine merge mode. [Figure 28A] FIG. 28A is a conceptual diagram for explaining an affine inter mode having two control points. [Figure 28B] FIG. 28B is a conceptual diagram for explaining an affine inter mode having three control points. [Figure 29] FIG. 29 is a flowchart showing an example of processing in the affine inter mode. [Figure 30A] FIG. 30A is a conceptual diagram for explaining an affine inter mode in which a current block has three control points and an adjacent block has two control points. [Figure 30B] FIG. 30B is a conceptual diagram for explaining an affine inter mode in which a current block has two control points and an adjacent block has three control points. [Figure 31A] FIG. 31A is a flow chart showing a merge mode including decoder motion vector refinement (DMVR). [Figure 31B] FIG. 31B is a conceptual diagram for explaining an example of the DMVR process. [Diagram 32] FIG. 32 is a flowchart showing an example of generation of a predicted image. [Diagram 33] FIG. 33 is a flowchart showing another example of generation of a predicted image. [Diagram 34] FIG. 34 is a flowchart showing another example of generation of a predicted image. [Diagram 35]FIG. 35 is a flowchart illustrating an example of a predictive image correction process using overlapped block motion compensation (OBMC). [Diagram 36] FIG. 36 is a conceptual diagram for explaining an example of the predicted image correction process by the OBMC process. [Figure 37] FIG. 37 is a conceptual diagram for explaining generation of predicted images of two triangles. [Figure 38] FIG. 38 is a conceptual diagram for explaining a model assuming uniform linear motion. [Figure 39] FIG. 39 is a conceptual diagram for explaining an example of a predicted image generating method using luminance correction processing by LIC (local illumination compensation) processing. [Diagram 40] FIG. 40 is a block diagram showing an example of implementation of an encoding device. [Diagram 41] FIG. 41 is a block diagram showing a functional configuration of a decoding device according to an embodiment. As shown in FIG. [Diagram 42] FIG. 42 is a flowchart showing an example of the overall decoding process by the decoding device. [Diagram 43] FIG. 43 is a flowchart illustrating an example of processing performed in the prediction processing unit of the decoding device. [Diagram 44] FIG. 44 is a flowchart showing another example of the process performed in the prediction processing unit of the decoding device. [Diagram 45] FIG. 45 is a flowchart showing an example of inter prediction in the normal inter mode in the decoding device. [Figure 46] FIG. 46 is a block diagram showing an implementation example of a decoding device. [Figure 47] FIG. 47 is a flowchart illustrating an overall processing flow for dividing an image block into multiple partitions, including at least a first partition having a non-rectangular shape (eg, a triangle) and a second partition for further processing, according to one embodiment. [Figure 48]FIG. 48 is a diagram showing two example methods of dividing an image block into a first partition having a non-rectangular shape (e.g., a triangle) and a second partition (also having a non-rectangular shape in the examples shown). [Figure 49] FIG. 49 illustrates an example of a boundary smoothing process that involves weighting a plurality of first values ​​of a plurality of boundary pixels predicted based on a first partition and a plurality of second values ​​of a plurality of boundary pixels predicted based on a second partition. [Figure 50] FIG. 50 illustrates three further examples of boundary smoothing processes that involve weighting first values ​​of boundary pixels predicted based on a first partition and second values ​​of boundary pixels predicted based on a second partition. [Figure 51] FIG. 51 is a table diagram showing example parameters ("first index values") and a number of sets of information that are respectively encoded by the number of parameters. [Figure 52] FIG. 52 is a table showing binarization of multiple parameters (multiple index values). [Diagram 53] FIG. 53 is a flowchart showing a process for dividing an image block into a plurality of partitions, including a first partition and a second partition having a non-rectangular shape. [Figure 54] FIG. 54 is a diagram illustrating examples of dividing an image block into partitions including a first partition having a non-rectangular shape, which in the examples shown is a triangle, and a second partition. [Figure 55] FIG. 55 is a diagram showing further examples of dividing an image block into partitions including a first partition having a non-rectangular shape that in the examples shown is a polygon with at least five sides and a corner, and a second partition. [Figure 56]FIG. 56 is a flowchart illustrating a boundary smoothing process that involves weighting first values ​​of a plurality of boundary images predicted based on a first partition and second values ​​of a plurality of boundary images predicted based on a second partition. [Figure 57A] FIG. 57A is a diagram showing an example of a boundary smoothing process in which multiple weighted first values ​​for multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​for multiple boundary pixels are predicted based on a second partition. [Figure 57B] FIG. 57B is a diagram showing an example of a boundary smoothing process in which multiple weighted first values ​​for multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​for multiple boundary pixels are predicted based on a second partition. [Figure 57C] FIG. 57C is a diagram showing an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Fig. 57D] FIG. 57D is a diagram showing an example of a boundary smoothing process in which multiple weighted first values ​​of multiple boundary pixels are predicted based on a first partition, and multiple weighted second values ​​of multiple boundary pixels are predicted based on a second partition. [Figure 58] Figure 58 is a flowchart showing a method performed on the encoding device side in which an image block is divided into multiple partitions, including a first partition having a non-rectangular shape and a second partition, based on partition parameters indicating the division, and in entropy encoding, one or more parameters including the partition parameters are written into a bitstream. [Figure 59]Figure 59 is a flowchart showing a method performed on a decoding device side, which reads one or more parameters from a bitstream including a partition parameter indicating division of an image block into multiple partitions including a first partition having a non-rectangular shape and a second partition, divides the image block into multiple partitions based on the partition parameter, and decodes the first partition and the second partition. [Figure 60] Figure 60 is a table diagram showing example partition parameters ("multiple first index values") each indicating division of an image block into multiple partitions, including a first partition and a second partition having a non-rectangular shape, and multiple sets of information that may each be coded together by the multiple partition parameters. [Figure 61] FIG. 61 is a table of several example combinations of first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is to be divided into several partitions, including a first partition having a non-rectangular shape and a second partition. [Figure 62] FIG. 62 is a flowchart illustrating an example of a process flow for predicting a first sample set for a first partition of a current picture with a first motion vector, predicting a second sample set for a first portion of the first partition with a second motion vector, weighting the first and second sample sets, storing at least one of the first and second motion vectors for the first partition, encoding or decoding the first partition using the weighted samples, and further processing in accordance with one embodiment. [Figure 63] FIG. 63 is a conceptual diagram showing an example of a method for dividing an image block into a first partition and a second partition. [Figure 64] FIG. 64 is a conceptual diagram showing nearby spatially adjacent partitions and non-neighboring spatially adjacent partitions of the first partition. [Figure 65]FIG. 65 is a conceptual diagram showing uni-predictive motion vector candidates and bi-predictive motion vector candidates for an image block. [Figure 66] FIG. 66 is a conceptual diagram illustrating an example of a first portion of a first partition and a first and second sample sets. [Figure 67] FIG. 67 is a conceptual diagram illustrating a first portion of a first partition, which is a portion of the first partition that overlaps with a portion of an adjacent partition. [Figure 68] FIG. 68 is a conceptual diagram showing an example where the first motion vector and the second motion vector are uni-predictive motion vectors pointing to different pictures in different reference picture lists. [Figure 69] FIG. 69 is a conceptual diagram showing an example where the first motion vector and the second motion vector are uni-predictive motion vectors pointing to pictures in a single reference picture list. [Figure 70] FIG. 70 is a conceptual diagram showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same picture in the same reference picture list. [Figure 71] FIG. 71 is a conceptual diagram showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same picture in the same reference picture list. [Figure 72] FIG. 72 is a conceptual diagram showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same picture in the same reference picture list. [Figure 73] FIG. 73 is a block diagram showing the overall configuration of a content supply system that realizes a content distribution service. [Figure 74] FIG. 74 is a conceptual diagram showing an example of a coding structure in scalable coding. [Figure 75] FIG. 75 is a conceptual diagram showing an example of a coding structure in scalable coding. [Figure 76] FIG. 76 is a conceptual diagram showing an example of a display screen of a web page. [Figure 77]FIG. 77 is a conceptual diagram showing an example of a display screen of a web page. [Figure 78] FIG. 78 is a block diagram showing an example of a smartphone. [Figure 79] FIG. 79 is a block diagram showing an example configuration of a smartphone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the relationship and order of the steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0017] In the following, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, for example, any of the following may be implemented.

[0018] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0019] (2) In the encoding device or decoding device of the embodiment, the functions or processes performed by some of the multiple components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, deleting, etc. For example, any function or process may be replaced or combined with another function or process described in any of the aspects of the present disclosure.

[0020] (3) In the method implemented by the encoding device or decoding device of the embodiment, some of the processes included in the method may be arbitrarily changed, such as added, replaced, deleted, etc. For example, any process in the method may be replaced or combined with another process described in any of the aspects of the present disclosure.

[0021] (4) Some of the multiple components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in each aspect of the present disclosure.

[0022] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0023] (6) In a method implemented by an encoding device or a decoding device of an embodiment, any of the multiple processes included in the method may be replaced or combined with a process described in any of the aspects of the present disclosure or with any similar process.

[0024] (7) Some of the processes among the multiple processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0025] (8) The manner of implementing the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or configurations may be implemented in a device used for a purpose other than the video encoding or video decoding disclosed in the embodiment.

[0026] [Encoding device] First, a coding device according to an embodiment will be described. Fig. 1 is a block diagram showing a functional configuration of a coding device 100 according to an embodiment. The coding device 100 is a video coding device that codes a video on a block-by-block basis.

[0027] As shown in FIG. 1, the encoding device 100 is a device that encodes an image on a block-by-block basis, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128.

[0028] The encoding device 100 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The encoding device 100 may also be realized as one or more dedicated electronic circuits corresponding to the division unit 102, the subtraction unit 104, the transformation unit 106, the quantization unit 108, the entropy coding unit 110, the inverse quantization unit 112, the inverse transformation unit 114, the addition unit 116, the loop filter unit 120, the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0029] Below, the overall processing flow of the encoding device 100 will be described, and then each component included in the encoding device 100 will be described.

[0030] [Overall encoding process flow] FIG. 2 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0031] First, the division unit 102 of the encoding device 100 divides each picture included in an input image, which is a moving image, into a plurality of fixed-size blocks (e.g., 128×128 pixels) (step Sa_1). Then, the division unit 102 selects a division pattern (also called a block shape) for the fixed-size blocks (step Sa_2). That is, the division unit 102 further divides the fixed-size block into a plurality of blocks constituting the selected division pattern. Then, the encoding device 100 performs the process of steps Sa_3 to Sa_9 for each of the plurality of blocks (i.e., the block to be encoded).

[0032] That is, a prediction processing unit consisting of all or part of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 generates a prediction signal (also called a prediction block) of the block to be coded (also called a current block) (step Sa_3).

[0033] Next, the subtraction unit 104 generates a difference between the encoding target block and the prediction block as a prediction residual (also called a difference block) (step Sa_4).

[0034] Next, the transform unit 106 and the quantization unit 108 perform transform and quantization on the difference block to generate a plurality of quantized coefficients (step Sa_5). Note that a block made up of a plurality of quantized coefficients is also called a coefficient block.

[0035] Next, the entropy coding unit 110 performs coding (specifically, entropy coding) on ​​the coefficient block and the prediction parameters related to the generation of the prediction signal to generate a coded signal (step Sa_6). The coded signal is also called a coded bit stream, a compressed bit stream, or a stream.

[0036] Next, the inverse quantization unit 112 and the inverse transform unit 114 perform inverse quantization and inverse transform on the coefficient block to reconstruct a plurality of prediction residuals (that is, difference blocks) (step Sa_7).

[0037] Next, the adder 116 reconstructs the current block into a reconstructed image (also called a reconstructed block or a decoded image block) by adding the predicted block to the restored difference block (step Sa_8). In this way, a reconstructed image is generated.

[0038] When this reconstructed image is generated, the loop filter unit 120 performs filtering on the reconstructed image as necessary (step Sa_9).

[0039] Then, the encoding device 100 determines whether or not encoding of the entire picture is completed (step Sa_10), and if it determines that encoding is not completed (No in step Sa_10), repeats the process from step Sa_2.

[0040] In the above example, the encoding device 100 selects one division pattern for fixed-size blocks and encodes each block according to the division pattern, but it may also encode each block according to each of a plurality of division patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of division patterns and select, for example, the encoded signal obtained by encoding according to the division pattern with the smallest cost as the encoded signal to be output.

[0041] As shown in the figure, the processes of steps Sa_1 to Sa_10 are performed sequentially by the encoding device 100. Alternatively, some of the processes may be performed in parallel, or the order of the processes may be changed.

[0042] [Divided part] The division unit 102 divides each picture included in the input video into a plurality of blocks, and outputs each block to the subtraction unit 104. For example, the division unit 102 first divides the picture into blocks of a fixed size (for example, 128x128). Other fixed block sizes may be adopted. The fixed-size blocks may be called coding tree units (CTUs). Then, the division unit 102 divides each of the fixed-size blocks into blocks of a variable size (for example, 64x64 or less) based on, for example, recursive quadtree and / or binary tree block division. That is, the division unit 102 selects a division pattern. The variable-size blocks may be called coding units (CUs), prediction units (PUs), or transform units (TUs). Note that in various processing examples, CUs, PUs, and TUs do not need to be distinguished, and some or all of the blocks in a picture may be the processing units of CUs, PUs, and TUs.

[0043] Fig. 3 is a conceptual diagram showing an example of block division in the embodiment, in which solid lines represent block boundaries based on quadtree block division, and dashed lines represent block boundaries based on binary tree block division.

[0044] Here, the block 10 is a square block of 128x128 pixels (128x128 block). This 128x128 block 10 is first divided into four square 64x64 blocks (quadtree block division).

[0045] The top-left 64x64 block is further divided vertically into two rectangular 32x64 blocks, and the left 32x64 block is further divided vertically into two rectangular 16x64 blocks (binary tree block division). As a result, the top-left 64x64 block is divided into two 16x64 blocks 11 and 12 and a 32x64 block 13.

[0046] The top right 64x64 block is divided horizontally into two rectangular 64x32 blocks 14, 15 (binary tree block division).

[0047] The bottom left 64x64 block is divided into four square 32x32 blocks (quadtree block division). Of the four 32x32 blocks, the top left and bottom right blocks are further divided. The top left 32x32 block is divided vertically into two rectangular 16x32 blocks, and the right 16x32 block is further divided horizontally into two 16x16 blocks (binary tree block division). The bottom right 32x32 block is divided horizontally into two 32x16 blocks (binary tree block division). As a result, the bottom left 64x64 block is divided into a 16x32 block 16, two 16x16 blocks 17, 18, two 32x32 blocks 19, 20, and two 32x16 blocks 21, 22.

[0048] The bottom right 64x64 block 23 is not split.

[0049] 3, the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0050] In Fig. 3, one block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to this. For example, one block may be divided into three blocks (ternary tree block division). Division including such ternary tree block division is sometimes called MBT (multi type tree) division.

[0051] [Picture Composition Slices / Tiles] In order to decode pictures in parallel, the pictures may be configured in slice units or tile units. Pictures configured in slice units or tile units may be configured by the division unit 102.

[0052] A slice is a basic coding unit constituting a picture. A picture is made up of, for example, one or more slices. Furthermore, a slice is made up of one or more consecutive coding tree units (CTUs).

[0053] FIG. 4A is a conceptual diagram showing an example of a slice configuration. For example, a picture includes 11×8 CTUs and is divided into four slices (slices 1-4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes 29 CTUs, and slice 4 includes 22 CTUs. Here, each CTU in a picture belongs to one of the slices. The shape of a slice is obtained by dividing a picture in the horizontal direction. The boundary of a slice does not need to be an edge of a screen, and may be any boundary of a CTU in a screen. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, a raster scan order. In addition, a slice includes header information and encoded data. The header information may describe the characteristics of the slice, such as the address of the CTU at the beginning of the slice and the slice type.

[0054] A tile is a rectangular unit that makes up a picture. Each tile may be assigned a number called a TileId in raster scan order.

[0055] FIG. 4B is a conceptual diagram showing an example of a tile configuration. For example, a picture includes 11×8 CTUs and is divided into four rectangular tiles (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed in raster scan order. When tiles are used, at least one CTU is processed in raster scan order in each of multiple tiles. For example, as shown in FIG. 4B, the processing order of multiple CTUs included in tile 1 is from the left end of the first row of tile 1 to the right end of the first row of tile 1, and then from the left end of the second row of tile 1 to the right end of the second row of tile 1.

[0056] It should be noted that one tile may include one or more slices, and one slice may include one or more tiles.

[0057] [Subtraction section] The subtraction unit 104 subtracts a prediction signal (a prediction sample input from a prediction control unit 128 described below) from an original signal (original sample) input from the division unit 102 for each block divided by the division unit 102. That is, the subtraction unit 104 calculates a prediction error (also called a residual) of a block to be coded (hereinafter called a current block). Then, the subtraction unit 104 outputs the calculated prediction error (residual) to the conversion unit 106.

[0058] The original signal is an input signal to the encoding device 100, and is a signal representing an image of each picture constituting a moving image (for example, a luminance (luma) signal and two color difference (chroma) signals). Hereinafter, the signal representing the image may also be referred to as a sample.

[0059] [Conversion section] The transform unit 106 transforms the prediction error in the spatial domain into a transform coefficient in the frequency domain, and outputs the transform coefficient to the quantization unit 108. Specifically, the transform unit 106 performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction error in the spatial domain, for example. The predetermined DCT or DST may be determined in advance.

[0060] The transform unit 106 may adaptively select a transform type from among a plurality of transform types, and transform the prediction errors into transform coefficients using a transform basis function corresponding to the selected transform type. Such a transform may be called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT).

[0061] The multiple transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Figure 5A is a table showing transform basis functions corresponding to example transform types. In Figure 5A, N indicates the number of input pixels. The selection of a transform type from among the multiple transform types may depend, for example, on the type of prediction (intra prediction and inter prediction) or on the intra prediction mode.

[0062] Such information indicating whether EMT or AMT is applied (e.g., called an EMT flag or an AMT flag) and information indicating the selected transformation type are usually signaled at a CU level, but the signaling of such information does not need to be limited to the CU level and may be at other levels (e.g., a bit sequence level, a picture level, a slice level, a tile level, or a CTU level).

[0063] Furthermore, the transform unit 106 may retransform the transform coefficients (transformation results). Such retransformation may be called an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transform unit 106 performs retransformation for each subblock (e.g., 4x4 subblock) included in a block of transform coefficients corresponding to intra-prediction errors. Information indicating whether or not to apply NSST and information regarding a transform matrix used in NSST are usually signaled at a CU level. Note that signaling of these pieces of information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0064] Separable transformation and non-separable transformation may be applied to the transformation unit 106. Separable transformation is a method of performing transformation multiple times by separating the input into directions for the number of dimensions, and non-separable transformation is a method of performing transformation collectively when the input is multidimensional, regarding two or more dimensions as one dimension.

[0065] For example, one example of a non-separable transformation is one in which, if the input is a 4x4 block, it is treated as a single array with 16 elements, and the transformation process is performed on that array using a 16x16 transformation matrix.

[0066] As another example of a non-separable transformation, a 4x4 input block may be treated as a single array with 16 elements, and then a transformation (Hypercube Givens Transform) may be performed on the array by performing multiple Givens rotations.

[0067] In the transform in the transform unit 106, the type of basis for transforming into the frequency domain can be switched according to the area in the CU. One example is SVT (Spatially Varying Transform). In SVT, as shown in FIG. 5B, a CU is divided into two equal parts in the horizontal or vertical direction, and only one of the areas is transformed into the frequency domain. The type of transform basis can be set for each area, and for example, DST7 and DCT8 are used. In this example, only one of the two areas in the CU is transformed and the other is not transformed, but both areas may be transformed. In addition, the division method can be more flexible, such as not only dividing into two, but also dividing into four equal parts, or separately encoding information indicating the division and signaling it in the same way as the CU division. In addition, SVT is also called SBT (Sub-block Transform).

[0068] [Quantization section] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the transform coefficients of the current block in a predetermined scanning order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized transform coefficients of the current block (hereinafter, referred to as quantized coefficients) to the entropy coding unit 110 and the inverse quantization unit 112. The predetermined scanning order may be determined in advance.

[0069] The predetermined scanning order is an order for quantization / dequantization of transform coefficients. For example, the predetermined scanning order may be defined as an ascending order of frequency (low to high frequencies) or a descending order of frequency (high to low frequencies).

[0070] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the quantization error increases.

[0071] In addition, a quantization matrix may be used for quantization. For example, several types of quantization matrices may be used corresponding to frequency transform sizes such as 4x4 and 8x8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and chrominance. Note that quantization refers to digitizing values ​​sampled at a predetermined interval in association with a predetermined level, and in this technical field, it may be referred to using other expressions such as rounding, rounding, and scaling, or rounding, rounding, and scaling may be adopted. The predetermined interval and level may be determined in advance.

[0072] There are two methods of using a quantization matrix: one is to use a quantization matrix that is directly set on the encoding device side, and the other is to use a default quantization matrix (default matrix). By directly setting a quantization matrix on the encoding device side, it is possible to set a quantization matrix according to the characteristics of an image. However, in this case, there is a disadvantage that the amount of code increases due to the encoding of the quantization matrix.

[0073] On the other hand, there is a method that does not use a quantization matrix and quantizes the coefficients of high-frequency components and low-frequency components in the same way. Note that this method is equivalent to using a quantization matrix in which all coefficients have the same value (a flat matrix).

[0074] The quantization matrix may be specified, for example, in a Sequence Parameter Set (SPS) or a Picture Parameter Set (PPS). The SPS contains parameters used for a sequence, and the PPS contains parameters used for a picture. The SPS and PPS are sometimes simply referred to as parameter sets.

[0075] [Entropy coding part] The entropy coding unit 110 generates a coded signal (coded bit stream) based on the quantized coefficients input from the quantization unit 108. Specifically, the entropy coding unit 110, for example, binarizes the quantized coefficients, arithmetically codes the binary signal, and outputs a compressed bit stream or sequence.

[0076] [Dequantization section] The inverse quantization unit 112 inverse quantizes the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantized coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114. The predetermined scanning order may be determined in advance.

[0077] [Inverse conversion section] The inverse transform unit 114 restores the prediction error (residual) by inverse transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction error of the current block by performing an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients. Then, the inverse transform unit 114 outputs the restored prediction error to the adder unit 116.

[0078] Note that the restored prediction error usually loses information due to quantization, and therefore does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction error usually contains a quantization error.

[0079] [Addition section] The adder 116 reconstructs a current block by adding the prediction error input from the inverse transformer 114 and the prediction sample input from the prediction control unit 128. The adder 116 then outputs the reconstructed block to the block memory 118 and the loop filter unit 120. The reconstructed block is sometimes called a local decoded block.

[0080] [Block memory] The block memory 118 is a storage unit for storing, for example, blocks referenced in intra prediction and in a picture to be coded (referred to as a current picture). Specifically, the block memory 118 stores the reconstructed block output from the adder 116.

[0081] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 122 stores the reconstructed block filtered by the loop filter unit 120.

[0082] [Loop filter section] The loop filter unit 120 applies a loop filter to the block reconstructed by the adder unit 116, and outputs the filtered reconstructed block to the frame memory 122. The loop filter is a filter (in-loop filter) used in the encoding loop, and includes, for example, a deblocking filter (DF or DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF).

[0083] In ALF, a least squared error filter is applied to remove coding artifacts. For example, for each 2x2 sub-block in the current block, one filter is selected from among multiple filters based on local gradient direction and activity.

[0084] Specifically, first, sub-blocks (e.g., 2x2 sub-blocks) are classified into a plurality of classes (e.g., 15 or 25 classes). The classification of the sub-blocks is performed based on the gradient direction and activity. For example, a classification value C (e.g., C=5D+A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes.

[0085] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions), and the gradient activity value A is derived, for example, by adding gradients in multiple directions and quantizing the sum.

[0086] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0087] The shape of the filter used in the ALF is, for example, a circularly symmetric shape. FIGS. 6A to 6C are diagrams showing a number of examples of the shape of the filter used in the ALF. FIG. 6A shows a 5×5 diamond-shaped filter, FIG. 6B shows a 7×7 diamond-shaped filter, and FIG. 6C shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Note that the signaling of the information indicating the shape of the filter does not need to be limited to the picture level, and may be at other levels (for example, the sequence level, slice level, tile level, CTU level, or CU level).

[0088] The on / off of ALF may be determined, for example, at the picture level or the CU level. For example, whether or not to apply ALF for luminance may be determined at the CU level, and whether or not to apply ALF for chrominance may be determined at the picture level. Information indicating whether or not to apply ALF is usually signaled at the picture level or the CU level. Note that the signaling of information indicating whether or not to apply ALF is not limited to the picture level or the CU level, and may be at another level (for example, the sequence level, the slice level, the tile level, or the CTU level).

[0089] The coefficient sets of multiple selectable filters (e.g., up to 15 or 25 filters) are typically signaled at the picture level, although the signaling of the coefficient sets need not be limited to the picture level, but may be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or subblock level).

[0090] [Loop filter section > Deblocking filter] In the deblocking filter, the loop filter unit 120 reduces distortion at block boundaries of the reconstructed image by applying a filtering process to the block boundaries.

[0091] FIG. 7 is a block diagram showing an example of a detailed configuration of the loop filter unit 120 functioning as a deblocking filter.

[0092] The loop filter unit 120 includes a boundary determination unit 1201 , a filter determination unit 1203 , a filter processing unit 1205 , a processing determination unit 1208 , a filter characteristic determination unit 1207 , and switches 1202 , 1204 and 1206 .

[0093] The boundary determination unit 1201 determines whether or not a pixel to be deblocking-filtered (i.e., a target pixel) exists near a block boundary. Then, the boundary determination unit 1201 outputs the determination result to the switch 1202 and the process determination unit 1208.

[0094] When the boundary determination unit 1201 determines that the target pixel is located near the block boundary, the switch 1202 outputs the image before filtering to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel is not located near the block boundary, the switch 1202 outputs the image before filtering to the switch 1206.

[0095] The filter determination unit 1203 determines whether or not to perform deblocking filter processing on the target pixel based on the pixel value of at least one surrounding pixel around the target pixel. Then, the filter determination unit 1203 outputs the determination result to the switch 1204 and the processing determination unit 1208.

[0096] When the filter determination unit 1203 determines that the deblocking filter process is to be performed on the target pixel, the switch 1204 outputs the unfiltered image acquired via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that the deblocking filter process is not to be performed on the target pixel, the switch 1204 outputs the unfiltered image acquired via the switch 1202 to the switch 1206.

[0097] When the filtering unit 1205 acquires an unfiltered image via the switches 1202 and 1204, it executes deblocking filtering on the target pixel using the filter characteristics determined by the filter characteristics determination unit 1207. Then, the filtering unit 1205 outputs the filtered pixel to the switch 1206.

[0098] The switch 1206 selectively outputs pixels that have not been subjected to the deblocking filter process and pixels that have been subjected to the deblocking filter process by the filter processing unit 1205 under the control of the process determination unit 1208 .

[0099] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filter determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near a block boundary and the filter determination unit 1203 determines that the target pixel is to be subjected to deblocking filter processing, the processing determination unit 1208 causes the switch 1206 to output a pixel that has been subjected to deblocking filter processing. In addition, in cases other than the above, the processing determination unit 1208 causes the switch 1206 to output a pixel that has not been subjected to deblocking filter processing. By repeatedly outputting pixels in this manner, a filtered image is output from the switch 1206.

[0100] FIG. 8 is a conceptual diagram showing an example of a deblocking filter having symmetric filter characteristics with respect to block boundaries.

[0101] In the deblocking filter process, for example, a pixel value and a quantization parameter are used to select one of two deblocking filters with different characteristics, that is, a strong filter and a weak filter. In the strong filter, when pixels p0 to p2 and pixels q0 to q2 exist on either side of a block boundary as shown in Fig. 8, the pixel values ​​of the pixels q0 to q2 are changed to pixel values ​​q'0 to q'2 by performing the calculation shown in the following equation, for example.

[0102] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8 q'1=(p0+q0+q1+q2+2) / 4 q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0103] In the above equations, p0 to p2 and q0 to q2 are the pixel values ​​of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the opposite side of the block boundary. Also, on the right side of each equation, the coefficients by which the pixel values ​​of each pixel used in the deblocking filter process are multiplied are filter coefficients.

[0104] Furthermore, in the deblocking filter process, a clipping process may be performed so that the pixel value after the calculation is not set to exceed the threshold. In this clipping process, the pixel value after the calculation according to the above formula is clipped to "the pixel value to be calculated ±2 × threshold" using a threshold determined from the quantization parameter. This makes it possible to prevent excessive smoothing.

[0105] Fig. 9 is a conceptual diagram for explaining block boundaries on which deblocking filter processing is performed, and Fig. 10 is a conceptual diagram showing an example of the Bs value.

[0106] The block boundary where the deblocking filter process is performed is, for example, the boundary of a PU (Prediction Unit) or TU (Transform Unit) of an 8x8 pixel block as shown in Fig. 9. The deblocking filter process can be performed in units of 4 rows or 4 columns. First, for blocks P and Q shown in Fig. 9, a Bs (Boundary Strength) value is determined as shown in Fig. 10.

[0107] According to the Bs value in FIG. 10, it is determined whether or not to perform deblocking filter processing of different strengths even for block boundaries belonging to the same image. Deblocking filter processing for the color difference signal is performed when the Bs value is 2. Deblocking filter processing for the luminance signal is performed when the Bs value is 1 or more and a predetermined condition is satisfied. The predetermined condition may be determined in advance. Note that the judgment condition for the Bs value is not limited to that shown in FIG. 10, and may be determined based on other parameters.

[0108] [Prediction processing unit (intra prediction unit, inter prediction unit, prediction control unit)] 11 is a flowchart showing an example of processing performed in the prediction processing unit of the encoding device 100. Note that the prediction processing unit is made up of all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0109] The prediction processing unit generates a prediction image of the current block (step Sb_1). This prediction image is also called a prediction signal or a prediction block. The prediction signal includes, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image that has already been obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0110] The reconstructed image may be, for example, an image of a reference picture or an image of an encoded block in a current picture, which is a picture that includes the current block. The encoded block in the current picture may be, for example, a neighboring block of the current block.

[0111] FIG. 12 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0112] The prediction processing unit generates a predicted image by a first method (step Sc_1a), generates a predicted image by a second method (step Sc_1b), and generates a predicted image by a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In these prediction methods, the above-mentioned reconstructed image may be used.

[0113] Next, the prediction processing unit selects one of the multiple predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). This selection of the predicted image, that is, the selection of the method or mode for obtaining the final predicted image, may be performed by calculating a cost for each generated predicted image and based on the cost. Alternatively, the selection of the predicted image may be performed based on parameters used in the encoding process. The encoding device 100 may signal information for identifying the selected predicted image, method, or mode in an encoding signal (also called an encoded bit stream). The information may be, for example, a flag. This allows the decoding device to generate a predicted image according to the method or mode selected in the encoding device 100 based on the information. Note that in the example shown in FIG. 12, the prediction processing unit generates a predicted image in each method and then selects one of the predicted images. However, the prediction processing unit may select a method or mode based on parameters used in the encoding process described above before generating those predicted images, and generate a predicted image according to the method or mode.

[0114] For example, the first and second methods may be intra prediction and inter prediction, respectively, and the prediction processing unit may select a final predicted image for the current block from predicted images generated according to these prediction methods.

[0115] FIG. 13 is a flowchart showing another example of the processing performed by the prediction processing unit of the encoding device 100.

[0116] First, the prediction processing unit generates a predicted image by intra prediction (step Sd_1a), and generates a predicted image by inter prediction (step Sd_1b). Note that the predicted image generated by intra prediction is also called an intra predicted image, and the predicted image generated by inter prediction is also called an inter predicted image.

[0117] Next, the prediction processing unit evaluates each of the intra-predicted image and the inter-predicted image (step Sd_2). A cost may be used for this evaluation. That is, the prediction processing unit calculates the cost C of each of the intra-predicted image and the inter-predicted image. This cost C can be calculated by an RD optimization model formula, for example, C=D+λ×R. In this formula, D is the coding distortion of the predicted image, and is represented by, for example, the sum of absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. Furthermore, R is the generated code amount of the predicted image, and specifically, is the code amount required for coding the motion information for generating the predicted image. Furthermore, λ is, for example, Lagrange's undetermined multiplier.

[0118] Then, the prediction processing unit selects the predicted image with the smallest calculated cost C from the intra-predicted image and the inter-predicted image as the final predicted image of the current block (step Sd_3). That is, a prediction method or mode for generating a predicted image of the current block is selected.

[0119] [Intra prediction section] The intra prediction unit 124 generates a prediction signal (intra prediction signal) by performing intra prediction (also called intra-screen prediction) of the current block with reference to a block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 128.

[0120] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes typically includes one or more non-directional prediction modes and a plurality of directional prediction modes. The plurality of predefined modes may be predefined.

[0121] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / HEVC standard.

[0122] The multiple directional prediction modes include, for example, 33 prediction modes defined in the H.265 / HEVC standard. The multiple directional prediction modes may include 32 prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 14 is a conceptual diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the 2 non-directional prediction modes are not shown in Figure 14).

[0123] In various processing examples, a luminance block may be referenced in intra prediction of a chrominance block. That is, a chrominance component of a current block may be predicted based on a luminance component of the current block. Such intra prediction may be called a cross-component linear model (CCLM) prediction. An intra prediction mode of a chrominance block that references such a luminance block (e.g., called a CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0124] The intra prediction unit 124 may correct pixel values ​​after intra prediction based on the gradient of reference pixels in the horizontal / vertical directions. Intra prediction with such correction may be called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is usually signaled at the CU level. Note that the signaling of this information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0125] [Inter prediction section] The inter prediction unit 126 performs inter prediction (also called inter prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture, thereby generating a prediction signal (inter prediction signal). The inter prediction is performed in units of the current block or a current sub-block (e.g., 4x4 block) in the current block. For example, the inter prediction unit 126 performs motion estimation in the reference picture for the current block or the current sub-block, and finds a reference block or sub-block that most closely matches the current block or the current sub-block. Then, the inter prediction unit 126 obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, and generates an inter prediction signal of the current block or sub-block. The inter prediction unit 126 outputs the generated inter prediction signal to the prediction control unit 128.

[0126] The motion information used for motion compensation may be signaled as an inter prediction signal in various forms, for example, a motion vector may be signaled, or, as another example, a difference between a motion vector and a motion vector predictor may be signaled.

[0127] [Basic flow of inter prediction] FIG. 15 is a flowchart showing an example of a basic flow of inter prediction.

[0128] The inter prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0129] Here, in generating a predicted image, the inter prediction unit 126 generates the predicted image by determining a motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). In determining an MV, the inter prediction unit 126 selects a candidate motion vector (candidate MV) (step Se_1) and derives an MV (step Se_2) to determine the MV. The selection of the candidate MV is performed, for example, by selecting at least one candidate MV from a candidate MV list. In deriving an MV, the inter prediction unit 126 may further select at least one candidate MV from the at least one candidate MV, and determine the selected at least one candidate MV as the MV of the current block. Alternatively, the inter prediction unit 126 may determine the MV of the current block by searching the area of ​​the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. Note that searching the area of ​​the reference picture may be referred to as motion estimation.

[0130] In the above example, steps Se_1 to Se_3 are performed by the inter prediction unit 126. However, the process of step Se_1 or step Se_2 may be performed by another component included in the encoding device 100.

[0131] [Motion vector derivation flow] FIG. 16 is a flowchart showing an example of motion vector derivation.

[0132] The inter prediction unit 126 derives the MV of the current block in a mode in which motion information (e.g., MV) is coded. In this case, for example, the motion information is coded as a prediction parameter and signaled. That is, the coded motion information is included in a coded signal (also called a coded bitstream).

[0133] Alternatively, the inter prediction unit 126 derives the MV in a mode in which motion information is not coded. In this case, the motion information is not included in the coded signal.

[0134] Here, the MV derivation mode may include a normal inter mode, a merge mode, a FRUC mode, and an affine mode, which will be described later. Among these modes, the modes for encoding motion information include a normal inter mode, a merge mode, and an affine mode (specifically, an affine inter mode and an affine merge mode). The motion information may include not only the MV but also predicted motion vector selection information, which will be described later. The modes for not encoding motion information include the FRUC mode. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0135] FIG. 17 is a flowchart showing another example of motion vector derivation.

[0136] The inter prediction unit 126 derives the MV of the current block in a mode of encoding the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the encoded signal. This differential MV is the difference between the MV of the current block and its predicted MV.

[0137] Alternatively, the inter prediction unit 126 derives the MV in a mode in which the differential MV is not coded. In this case, the coded differential MV is not included in the coded signal.

[0138] Here, as described above, the modes of deriving an MV include normal inter, merge mode, FRUC mode, and affine mode, which will be described later. Among these modes, the modes for encoding a differential MV include normal inter mode and affine mode (specifically, affine inter mode). Also, the modes for not encoding a differential MV include FRUC mode, merge mode, and affine mode (specifically, affine merge mode). The inter prediction unit 126 selects a mode for deriving an MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0139] [Motion vector derivation flow] FIG. 18 is a flowchart showing another example of motion vector derivation. There are a plurality of modes of MV derivation, that is, inter prediction modes, which are roughly divided into a mode in which a differential MV is coded and a mode in which a differential motion vector is not coded. The modes in which a differential MV is not coded include a merge mode, a FRUC mode, and an affine mode (specifically, an affine merge mode). Details of these modes will be described later, but simply, the merge mode is a mode in which the MV of the current block is derived by selecting a motion vector from a surrounding coded block, and the FRUC mode is a mode in which the MV of the current block is derived by searching between coded regions. In addition, the affine mode is a mode in which the motion vector of each of a plurality of sub-blocks constituting the current block is derived as the MV of the current block, assuming an affine transformation.

[0140] Specifically, as shown in the figure, when the inter prediction mode information indicates 0 (Sf_1 is 0), the inter prediction unit 126 derives a motion vector by the merge mode (Sf_2). When the inter prediction mode information indicates 1 (Sf_1 is 1), the inter prediction unit 126 derives a motion vector by the FRUC mode (Sf_3). When the inter prediction mode information indicates 2 (Sf_1 is 2), the inter prediction unit 126 derives a motion vector by the affine mode (specifically, the affine merge mode) (Sf_4). When the inter prediction mode information indicates 3 (Sf_1 is 3), the inter prediction unit 126 derives a motion vector by a mode for encoding a differential MV (for example, normal inter mode) (Sf_5).

[0141] [MV Derivation > Normal Inter Mode] The normal inter mode is an inter prediction mode in which the MV of the current block is derived based on a block similar to the image of the current block from the region of the reference picture indicated by the candidate MV. In addition, in this normal inter mode, the differential MV is coded.

[0142] FIG. 19 is a flowchart showing an example of inter prediction in the normal inter mode.

[0143] The inter prediction unit 126 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks around the current block in time or space (step Sg_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0144] Next, the inter prediction unit 126 extracts N candidate MVs (N is an integer equal to or greater than 2) from the multiple candidate MVs acquired in step Sg_1 as motion vector predictor candidates (also called predicted MV candidates) according to a predetermined priority order (step Sg_2). Note that the priority order may be predefined for each of the N candidate MVs.

[0145] Next, the inter prediction unit 126 selects one of the N motion vector predictor candidates as a motion vector predictor (also called a prediction MV) of the current block (step Sg_3). At this time, the inter prediction unit 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor into a stream. Note that the stream is the above-mentioned encoded signal or encoded bit stream.

[0146] Next, the inter prediction unit 126 derives the MV of the current block by referring to the coded reference picture (step Sg_4). At this time, the inter prediction unit 126 further encodes the difference value between the derived MV and the predicted motion vector as a differential MV into a stream. Note that the coded reference picture is a picture consisting of a plurality of blocks reconstructed after coding.

[0147] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Sg_5). Note that the predicted image is the above-mentioned inter prediction signal.

[0148] Furthermore, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the predicted image, which is included in the coded signal, is coded as, for example, a prediction parameter.

[0149] The candidate MV list may be used in common with lists used in other modes. Furthermore, processing related to the candidate MV list may be applied to processing related to lists used in other modes. Processing related to this candidate MV list may include, for example, extraction or selection of candidate MVs from the candidate MV list, sorting of the candidate MVs, or deletion of candidate MVs.

[0150] [MV Derivation > Merge Mode] Merge mode is an inter prediction mode that derives the MV of the current block by selecting a candidate MV from a candidate MV list as the MV for that block.

[0151] FIG. 20 is a flowchart showing an example of inter prediction in the merge mode.

[0152] The inter prediction unit 126 first obtains multiple candidate MVs for the current block based on information such as MVs of multiple coded blocks around the current block in time or space (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0153] Next, the inter prediction unit 126 derives the MV of the current block by selecting one candidate MV from the multiple candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream.

[0154] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Sh_3).

[0155] Furthermore, information indicating the inter prediction mode (merge mode in the above example) used to generate the predicted image, which is included in the coded signal, is coded as, for example, a prediction parameter.

[0156] FIG. 21 is a conceptual diagram for explaining an example of a motion vector derivation process for a current picture in the merge mode.

[0157] First, a prediction MV list is generated in which prediction MV candidates are registered. The prediction MV candidates include a spatially adjacent prediction MV, which is an MV held by a plurality of coded blocks located spatially around the target block, a temporally adjacent prediction MV, which is an MV held by a nearby block projected onto the position of the target block in a coded reference picture, a joint prediction MV, which is an MV generated by combining the MV values ​​of the spatially adjacent prediction MV and the temporally adjacent prediction MV, and a zero prediction MV, which is an MV with a value of zero.

[0158] Next, one predicted MV is selected from the multiple predicted MVs registered in the predicted MV list, and is determined as the MV for the target block.

[0159] Furthermore, the variable length coding unit writes merge_idx, which is a signal indicating which predicted MV has been selected, into the stream and codes it.

[0160] Note that the predicted MVs registered in the predicted MV list described in Figure 21 are just an example, and the number may be different from the number shown in the figure, the configuration may not include some of the types of predicted MVs shown in the figure, or the configuration may include additional predicted MVs other than the types of predicted MVs shown in the figure.

[0161] The final MV may be determined by performing a decoder motion vector refinement (DMVR) process, which will be described later, using the MV of the current block derived in the merge mode.

[0162] The prediction MV candidates are the above-mentioned candidate MVs, and the prediction MV list is the above-mentioned candidate MV list. The candidate MV list may also be called a candidate list. merge_idx is MV selection information.

[0163] [MV derivation > FRUC mode] The motion information may be derived at the decoding device side without being signaled from the encoding device side. As described above, the merge mode defined in the H.265 / HEVC standard may be used. For example, the motion information may be derived by performing motion estimation at the decoding device side. In the embodiment, the motion estimation is performed at the decoding device side without using pixel values ​​of the current block.

[0164] Here, a mode in which motion estimation is performed on the decoding device side will be described. This mode in which motion estimation is performed on the decoding device side is sometimes called a pattern matched motion vector derivation (PMMVD) mode or a frame rate up-conversion (FRUC) mode.

[0165] FIG. 22 shows an example of the FRUC process in the form of a flow chart. First, a list of multiple candidates (i.e., a candidate MV list, which may be common to the merge list) each having a predicted motion vector (MV) is generated with reference to the motion vectors of coded blocks spatially or temporally adjacent to the current block (step Si_1). Next, a best candidate MV is selected from the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate MV is selected based on the evaluation value. Then, a motion vector for the current block is derived based on the motion vector of the selected candidate (step Si_4). Specifically, for example, the motion vector of the selected candidate (best candidate MV) is derived as it is as the motion vector for the current block. Also, for example, the motion vector for the current block may be derived by performing pattern matching in a peripheral area of ​​a position in a reference picture corresponding to the motion vector of the selected candidate. That is, a search is performed on the area around the best candidate MV using pattern matching and evaluation values ​​in the reference picture, and if an MV with a better evaluation value is found, the best candidate MV is updated to that MV and used as the final MV for the current block. It is also possible to configure the system without performing the process of updating to an MV with a better evaluation value.

[0166] Finally, the inter prediction unit 126 performs motion compensation on the current block using the derived MV and the coded reference picture to generate a predicted image of the current block (step Si_5).

[0167] The same processing may be performed when processing is performed in sub-block units.

[0168] The evaluation value may be calculated by various methods. For example, a reconstructed image of an area in a reference picture corresponding to the motion vector is compared with a reconstructed image of a predetermined area (which may be, for example, an area of ​​another reference picture or an area of ​​an adjacent block of the current picture, as shown below). The predetermined area may be predetermined.

[0169] Then, the difference between the pixel values ​​of the two reconstructed images may be calculated and used as an evaluation value for the motion vector. Note that the evaluation value may be calculated using other information in addition to the difference value.

[0170] Next, an example of pattern matching will be described in detail. First, one candidate MV included in a candidate MV list (e.g., a merge list) is selected as a start point of a search by pattern matching. For example, a first pattern matching or a second pattern matching can be used as the pattern matching. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0171] [MV derivation > FRUC > Bilateral matching] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, an area in another reference picture along the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the candidate. The predetermined area may be determined in advance.

[0172] FIG. 23 is a conceptual diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in FIG. 23, in the first pattern matching, two motion vectors (MV0, MV1) are derived by searching for a pair of blocks that best match among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of a current block (Cur block). Specifically, for the current block, a difference is derived between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by a candidate MV and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV by a display time interval, and an evaluation value is calculated using the obtained difference value. It is possible to select the candidate MV with the best evaluation value among multiple candidate MVs as the final MV, which may bring about good results.

[0173] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) pointing to two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located between two reference pictures in time and the temporal distances from the current picture to the two reference pictures are equal, the first pattern matching derives bidirectional motion vectors that are mirror-symmetric.

[0174] [MV derivation > FRUC > Template matching] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Therefore, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the above-mentioned candidate.

[0175] Fig. 24 is a conceptual diagram for explaining an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in Fig. 24, in the second pattern matching, a motion vector of a current block is derived by searching in a reference picture (Ref0) for a block that best matches a block adjacent to a current block (Cur block) in a current picture (Cur Pic). Specifically, for a current block, a difference is derived between a reconstructed image of both or either of the left adjacent and / or upper adjacent coded areas and a reconstructed image at the same position in a coded reference picture (Ref0) specified by a candidate MV, an evaluation value is calculated using the obtained difference value, and a candidate MV with the best evaluation value among a plurality of candidate MVs can be selected as a best candidate MV.

[0176] Information indicating whether such a FRUC mode is applied (e.g., called a FRUC flag) may be signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating an applicable pattern matching method (first pattern matching or second pattern matching) may be signaled at the CU level. Note that the signaling of such information does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0177] [MV derivation > Affine mode] Next, a description will be given of an affine mode in which a motion vector is derived for each sub-block based on the motion vectors of a plurality of adjacent blocks. This mode is sometimes called an affine motion compensation prediction mode.

[0178] FIG. 25A is a conceptual diagram for explaining an example of derivation of a motion vector for each sub-block based on the motion vectors of multiple adjacent blocks. In FIG. 25A, the current block includes 16 4x4 sub-blocks. Here, a motion vector v0 of the top left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and similarly, a motion vector v1 of the top right corner control point of the current block is derived based on the motion vectors of the adjacent sub-blocks. Then, the two motion vectors v0 and v1 may be projected according to the following formula (1A), and the motion vectors (v x ,v y ) may be derived.

[0179]

number

[0180] Here, x and y respectively indicate the horizontal and vertical positions of the sub-block, and w indicates a predetermined weighting factor, which may be determined in advance.

[0181] Such information indicating the affine mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating the affine mode does not need to be limited to the CU level, and may be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0182] In addition, such an affine mode may include several modes that differ in the method of deriving the motion vectors of the upper-left and upper-right corner control points. For example, the affine mode includes two modes: an affine inter (also called an affine normal inter) mode and an affine merge mode.

[0183] [MV derivation > Affine mode] FIG. 25B is a conceptual diagram for explaining an example of derivation of a motion vector for each sub-block in an affine mode having three control points. In FIG. 25B, the current block includes 16 4x4 sub-blocks. Here, a motion vector v0 of the upper left corner control point of the current block is derived based on the motion vector of an adjacent block, and similarly, a motion vector v1 of the upper right corner control point of the current block is derived based on the motion vector of the adjacent block, and a motion vector v2 of the lower left corner control point of the current block is derived based on the motion vector of the adjacent block. Then, the three motion vectors v0, v1, and v2 may be projected according to the following formula (1B), and the motion vectors (v x ,v y ) may be derived.

[0184]

number

[0185] Here, x and y respectively indicate the horizontal and vertical positions of the subblock center, w indicates the width of the current block, and h indicates the height of the current block.

[0186] Affine modes with different numbers of control points (e.g., two and three) may be switched and signaled at the CU level. Note that information indicating the number of control points of the affine mode used at the CU level may also be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or subblock level).

[0187] In addition, such an affine mode having three control points may include several modes with different methods of deriving the motion vectors of the upper left, upper right, and lower left corner control points. For example, the affine mode includes two modes: affine inter (also called affine normal inter) mode and affine merge mode.

[0188] [MV Derivation > Affine Merge Mode] 26A, 26B, and 26C are conceptual diagrams for explaining the affine merge mode.

[0189] In the affine merge mode, as shown in FIG. 26A, for example, among the coded blocks A (left), B (top), C (top right), D (bottom left) and E (top left) adjacent to the current block, the predicted motion vectors of each of the control points of the current block are calculated based on a plurality of motion vectors corresponding to the blocks coded in affine mode. Specifically, these blocks are inspected in the order of coded blocks A (left), B (top), C (top right), D (bottom left) and E (top left), and the first valid block coded in affine mode is identified. Based on a plurality of motion vectors corresponding to this identified block, the predicted motion vector of the control point of the current block is calculated.

[0190] For example, as shown in Figure 26B, when the block A adjacent to the left of the current block is coded in affine mode with two control points, motion vectors v3 and v4 are derived by projecting the upper left corner and upper right corner positions of the coded block including block A. Then, from the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner are calculated.

[0191] For example, as shown in Figure 26C, when the block A adjacent to the left of the current block is coded in affine mode with three control points, motion vectors v3, v4 and v5 are derived by projecting to the upper left corner, upper right corner and lower left corner positions of the coded block including block A. Then, from the derived motion vectors v3, v4 and v5, the predicted motion vector v0 of the control point of the upper left corner of the current block, the predicted motion vector v1 of the control point of the upper right corner and the predicted motion vector v2 of the control point of the lower left corner are calculated.

[0192] Note that this predicted motion vector derivation method may be used to derive predicted motion vectors for each control point of the current block in step Sj_1 of FIG. 29, which will be described later.

[0193] FIG. 27 is a flow chart illustrating an example of the affine merge mode.

[0194] In the affine merge mode, as shown in the figure, first, the inter prediction unit 126 derives prediction MVs for each of the control points of the current block (step Sk_1). The control points are the upper left and upper right corner points of the current block as shown in Figure 25A, or the upper left, upper right and lower left corner points of the current block as shown in Figure 25B.

[0195] That is, the inter prediction unit 126 examines the coded blocks in the following order, as shown in FIG. 26A: coded block A (left), block B (top), block C (top right), block D (bottom left) and block E (top left), and identifies the first valid block coded in affine mode.

[0196] Then, when block A is identified and has two control points, as shown in FIG. 26B, the inter prediction unit 126 calculates the motion vector v0 of the control point of the upper left corner of the current block and the motion vector v1 of the control point of the upper right corner from the motion vectors v3 and v4 of the upper left corner and upper right corner of the coded block including block A. For example, the inter prediction unit 126 calculates the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner by projecting the motion vectors v3 and v4 of the upper left corner and upper right corner of the coded block onto the current block.

[0197] Alternatively, when block A is specified and block A has three control points, as shown in FIG. 26C, the inter prediction unit 126 calculates the motion vector v0 of the control point of the upper left corner of the current block, the motion vector v1 of the control point of the upper right corner, and the motion vector v2 of the control point of the lower left corner from the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block including block A. For example, the inter prediction unit 126 calculates the predicted motion vector v0 of the control point of the upper left corner of the current block, the predicted motion vector v1 of the control point of the upper right corner, and the motion vector v2 of the control point of the lower left corner by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the coded block onto the current block.

[0198] Next, the inter prediction unit 126 performs motion compensation for each of the sub-blocks included in the current block. That is, for each of the sub-blocks, the inter prediction unit 126 calculates the motion vector of the sub-block as an affine MV using two predicted motion vectors v0 and v1 and the above formula (1A), or three predicted motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter prediction unit 126 performs motion compensation for the sub-block using the affine MVs and the coded reference picture (step Sk_3). As a result, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0199] [MV Derivation > Affine Intermode] FIG. 28A is a conceptual diagram for explaining an affine inter mode having two control points.

[0200] In this affine inter mode, as shown in Figure 28A, a motion vector selected from the motion vectors of the coded blocks A, B, and C adjacent to the current block is used as a predicted motion vector v0 of the control point of the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as a predicted motion vector v1 of the control point of the upper right corner of the current block.

[0201] FIG. 28B is a conceptual diagram for explaining an affine inter mode having three control points.

[0202] In this affine inter mode, as shown in FIG. 28B, a motion vector selected from the motion vectors of the coded blocks A, B and C adjacent to the current block is used as the predicted motion vector v0 of the control point of the upper left corner of the current block. Similarly, a motion vector selected from the motion vectors of the coded blocks D and E adjacent to the current block is used as the predicted motion vector v1 of the control point of the upper right corner of the current block. Furthermore, a motion vector selected from the motion vectors of the coded blocks F and G adjacent to the current block is used as the predicted motion vector v2 of the control point of the lower left corner of the current block.

[0203] FIG. 29 is a flowchart showing an example of the affine inter mode.

[0204] As shown in the figure, in the affine inter mode, the inter prediction unit 126 first derives predicted MVs (v0, v1) or (v0, v1, v2) of two or three control points of the current block (step Sj_1). The control points are the upper left corner, upper right corner, or lower left corner of the current block, as shown in FIG. 25A or FIG. 25B.

[0205] That is, the inter prediction unit 126 derives the predicted motion vector (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the motion vector of any one of the coded blocks in the vicinity of each control point of the current block shown in Figure 28A or Figure 28B. At this time, the inter prediction unit 126 codes the predicted motion vector selection information for identifying the two selected motion vectors into a stream.

[0206] For example, the inter prediction unit 126 may use cost evaluation or the like to determine which motion vector of an encoded block adjacent to the current block to select as the predicted motion vector of the control point, and may write a flag indicating which predicted motion vector has been selected in the bitstream.

[0207] Next, the inter prediction unit 126 performs motion search (steps Sj_3 and Sj_4) while updating each of the predicted motion vectors selected or derived in step Sj_1 (step Sj_2). That is, the inter prediction unit 126 calculates the motion vector of each sub-block corresponding to the predicted motion vector to be updated as an affine MV using the above-mentioned formula (1A) or formula (1B) (step Sj_3). Then, the inter prediction unit 126 performs motion compensation for each sub-block using the affine MVs and the coded reference picture (step Sj_4). As a result, the inter prediction unit 126 determines, in the motion search loop, for example, the predicted motion vector that provides the smallest cost as the motion vector of the control point (step Sj_5). At this time, the inter prediction unit 126 further codes the difference value between the determined MV and the predicted motion vector as a differential MV into a stream.

[0208] Finally, the inter prediction unit 126 performs motion compensation on the current block using the determined MV and the encoded reference picture to generate a predicted image of the current block (step Sj_6).

[0209] [MV Derivation > Affine Intermode] When affine modes with different numbers of control points (for example, two and three) are switched and signaled at the CU level, the number of control points may differ between the coded block and the current block. Figures 30A and 30B are conceptual diagrams for explaining a method of deriving a predicted vector of a control point when the number of control points differs between the coded block and the current block.

[0210] For example, as shown in FIG. 30A, when the current block has three control points, which are the upper left corner, the upper right corner and the lower left corner, and the block A adjacent to the left of the current block is coded in affine mode with two control points, motion vectors v3 and v4 are derived that are projected to the upper left corner and the upper right corner of the coded block including block A.Then, from the derived motion vectors v3 and v4, the predicted motion vector v0 of the control point of the upper left corner of the current block and the predicted motion vector v1 of the control point of the upper right corner are calculated.Furthermore, from the derived motion vectors v0 and v1, the predicted motion vector v2 of the control point of the lower left corner is calculated.

[0211] For example, as shown in Figure 30B, if the current block has two control points at the upper left corner and the upper right corner, and the block A adjacent to the left of the current block is coded in affine mode with three control points, motion vectors v3, v4 and v5 are derived by projecting the upper left corner, upper right corner and lower left corner positions of the coded block including block A. Then, the predicted motion vector v0 of the control point at the upper left corner of the current block and the predicted motion vector v1 of the control point at the upper right corner are calculated from the derived motion vectors v3, v4 and v5.

[0212] This prediction motion vector derivation method may be used to derive the prediction motion vector for each control point of the current block in step Sj_1 of FIG.

[0213] [MV Derivation > DMVR] FIG. 31A is a flowchart showing the relationship between the merge mode and the DMVR.

[0214] The inter prediction unit 126 derives a motion vector of the current block in merge mode (step Sl_1). Next, the inter prediction unit 126 determines whether or not to search for a motion vector, that is, to perform motion search (step Sl_2). Here, when the inter prediction unit 126 determines not to perform motion search (No in step Sl_2), it determines the motion vector derived in step Sl_1 as the final motion vector for the current block (step Sl_4). That is, in this case, the motion vector of the current block is determined in merge mode.

[0215] On the other hand, when it is determined in step Sl_1 that motion search is to be performed (Yes in step Sl_2), the inter prediction unit 126 derives a final motion vector for the current block by searching the surrounding area of ​​the reference picture indicated by the motion vector derived in step Sl_1 (step Sl_3). That is, in this case, the motion vector of the current block is determined by DMVR.

[0216] FIG. 31B is a conceptual diagram for explaining an example of DMVR processing for determining an MV.

[0217] First, the optimal MVP set for the current block (e.g., in merge mode) is set as the candidate MV. Then, according to the candidate MV(L0), reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 direction. Similarly, according to the candidate MV(L1), reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 direction. A template is generated by averaging these reference pixels.

[0218] Next, the template is used to search the surrounding areas of the candidate MVs in the first reference picture (L0) and the second reference picture (L1), and the MV with the smallest cost is determined as the final MV. Note that the cost value may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search area, the candidate MV value, etc.

[0219] Typically, the encoding device and a decoding device (to be described later) basically have the same configuration and operation as the processing described here.

[0220] Any process may be used other than the process example described here as long as it is capable of searching the vicinity of the candidate MV and deriving the final MV.

[0221] [Motion compensation > BIO / OBMC] In motion compensation, there is a mode in which a predicted image is generated and the predicted image is corrected, such as BIO and OBMC, which will be described later.

[0222] FIG. 32 is a flowchart showing an example of generation of a predicted image.

[0223] The inter prediction unit 126 generates a predicted image (step Sm_1), and corrects the predicted image using, for example, any of the modes described above (step Sm_2).

[0224] FIG. 33 is a flowchart showing another example of generation of a predicted image.

[0225] The inter prediction unit 126 determines a motion vector of the current block (step Sn_1). Next, the inter prediction unit 126 generates a predicted image (step Sn_2) and determines whether or not to perform correction processing (step Sn_3). Here, if the inter prediction unit 126 determines that correction processing is to be performed (Yes in step Sn_3), it generates a final predicted image by correcting the predicted image (step Sn_4). On the other hand, if the inter prediction unit 126 determines that correction processing is not to be performed (No in step Sn_3), it outputs the predicted image as the final predicted image without correcting it (step Sn_5).

[0226] Furthermore, motion compensation has a mode in which luminance is corrected when generating a predicted image, such as LIC, which will be described later.

[0227] FIG. 34 is a flowchart showing another example of generation of a predicted image.

[0228] The inter prediction unit 126 derives a motion vector of the current block (step So_1). Next, the inter prediction unit 126 determines whether or not to perform luminance correction processing (step So_2). Here, if the inter prediction unit 126 determines to perform luminance correction processing (Yes in step So_2), it generates a predicted image while performing luminance correction (step So_3). That is, the predicted image is generated by LIC. On the other hand, if the inter prediction unit 126 determines not to perform luminance correction processing (No in step So_2), it generates a predicted image by normal motion compensation without performing luminance correction (step So_4).

[0229] [Motion Compensation > OBMC] An inter prediction signal may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks. Specifically, a prediction signal based on the motion information obtained by motion search (in the reference picture) and a prediction signal based on the motion information of adjacent blocks (in the current picture) may be weighted and added to generate an inter prediction signal for each sub-block in the current block. Such inter prediction (motion compensation) may be called OBMC (overlapped block motion compensation).

[0230] In the OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether the OBMC mode is applied (e.g., called OBMC flag) may be signaled at the CU level. Note that the signaling level of these pieces of information does not need to be limited to the sequence level and the CU level, and may be other levels (e.g., the picture level, slice level, tile level, CTU level, or sub-block level).

[0231] An example of the OBMC mode will now be described more specifically. Figures 35 and 36 are a flowchart and a conceptual diagram for explaining an overview of the predictive image correction process in the OBMC process.

[0232] First, as shown in Fig. 36, a predicted image (Pred) is obtained by normal motion compensation using a motion vector (MV) assigned to a processing target (current) block. In Fig. 36, the arrow "MV" indicates a reference picture, and indicates what the current block of the current picture refers to in order to obtain a predicted image.

[0233] Next, the motion vector (MV_L) already derived for the coded left adjacent block is applied (reused) to the block to be coded to obtain a predicted image (Pred_L). The motion vector (MV_L) is indicated by an arrow "MV_L" pointing from the current block to the reference picture. The first correction of the predicted image is then performed by superimposing the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between the adjacent blocks.

[0234] Similarly, a motion vector (MV_U) already derived for the coded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The motion vector (MV_U) is indicated by an arrow "MV_U" pointing from the current block to the reference picture. The predicted image Pred_U is then superimposed on the predicted images (e.g., Pred and Pred_L) that have been corrected the first time, thereby performing a second correction of the predicted image. This has the effect of blending the boundaries between the adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block, with the boundaries with the adjacent blocks blended (smoothed).

[0235] Note that the above example is a two-pass correction method using the left adjacent and above adjacent blocks, but the correction method may also be a three-pass or more pass correction method using the right adjacent and / or below adjacent blocks.

[0236] The area in which overlapping is performed does not have to be the entire pixel area of ​​the block, but may be only a part of the area near the block boundary.

[0237] Here, the OBMC predicted image correction process for obtaining one predicted image Pred by superimposing additional predicted images Pred_L and Pred_U from one reference picture has been described. However, when the predicted image is corrected based on multiple reference pictures, the same process may be applied to each of the multiple reference pictures. In such a case, the OBMC image correction based on multiple reference pictures is performed to obtain a corrected predicted image from each reference picture, and then the obtained multiple corrected predicted images are further superimposed to obtain a final predicted image.

[0238] In addition, in OBMC, the unit of the target block may be a prediction block unit, or a sub-block unit obtained by further dividing the prediction block.

[0239] As a method of determining whether or not to apply OBMC processing, for example, there is a method of using obmc_flag, which is a signal indicating whether or not to apply OBMC processing. As a specific example, the encoding device may determine whether or not the target block belongs to an area with complex motion. If the target block belongs to an area with complex motion, the encoding device sets a value of 1 as obmc_flag and applies OBMC processing to perform encoding, and if the target block does not belong to an area with complex motion, the encoding device sets a value of 0 as obmc_flag and performs encoding of the block without applying OBMC processing. On the other hand, the decoding device decodes obmc_flag described in a stream (e.g., a compressed sequence), and switches whether or not to apply OBMC processing depending on the value to perform decoding.

[0240] In the above example, the inter prediction unit 126 generates one rectangular predicted image for the rectangular current block. However, the inter prediction unit 126 may generate multiple predicted images of shapes other than a rectangle for the rectangular current block, and combine the multiple predicted images to generate a final rectangular predicted image. The shape other than a rectangle may be, for example, a triangle.

[0241] FIG. 37 is a conceptual diagram for explaining generation of predicted images of two triangles.

[0242] The inter prediction unit 126 generates a predicted image of a triangle by performing motion compensation on a first partition of a triangle in the current block using a first MV of the first partition. Similarly, the inter prediction unit 126 generates a predicted image of a triangle by performing motion compensation on a second partition of a triangle in the current block using a second MV of the second partition. Then, the inter prediction unit 126 generates a predicted image of the same rectangle as the current block by combining these predicted images.

[0243] In the example shown in Fig. 37, the first partition and the second partition are each triangular, but they may be trapezoidal or may have different shapes. Furthermore, in the example shown in Fig. 37, the current block is composed of two partitions, but it may be composed of three or more partitions.

[0244] Also, the first partition and the second partition may overlap each other, i.e., the first partition and the second partition may include the same pixel area. In this case, the predicted image of the current block may be generated using the predicted image of the first partition and the predicted image of the second partition.

[0245] Furthermore, in this example, a predicted image is generated by inter prediction for both of the two partitions, but a predicted image may be generated by intra prediction for at least one partition.

[0246] [Motion compensation > BIO] Next, a method for deriving a motion vector will be described. First, a mode for deriving a motion vector based on a model assuming uniform linear motion will be described. This mode is sometimes called a BIO (bi-directional optical flow) mode.

[0247] Fig. 38 is a conceptual diagram for explaining a model assuming uniform linear motion. In Fig. 38, (vx, vy) indicates a velocity vector, and τ0 and τ1 indicate the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) indicates a motion vector corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates a motion vector corresponding to the reference picture Ref1.

[0248] In this case, under the assumption of uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (2) may be adopted.

[0249]

number

[0250] Here, I(k) denotes the luminance value of reference image k (k=0,1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of this optical flow equation and Hermite interpolation, block-wise motion vectors obtained from a merge list or the like may be corrected pixel by pixel.

[0251] Note that the decoding device may derive a motion vector using a method other than the method based on a model assuming uniform linear motion. For example, a motion vector may be derived for each sub-block based on the motion vectors of multiple adjacent blocks.

[0252] [Motion Compensation > LIC] Next, an example of a mode in which a predicted image (prediction) is generated using LIC (local illumination compensation) processing will be described.

[0253] FIG. 39 is a conceptual diagram for explaining an example of a predicted image generating method using a luminance correction process by LIC processing.

[0254] First, the MV is derived from the coded reference picture to obtain the reference image corresponding to the current block.

[0255] Next, extract information indicating how the luminance value of the current block has changed between the reference picture and the current picture. This extraction is performed based on the luminance pixel values ​​of the coded left adjacent reference area (peripheral reference area) and the coded upper adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the derived MV. Then, calculate a luminance correction parameter using the information indicating how the luminance value has changed.

[0256] A luminance correction process is performed by applying the luminance correction parameters to a reference image in a reference picture specified by the MV, thereby generating a predicted image for the current block.

[0257] It should be noted that the shape of the surrounding reference region in FIG. 39 is just an example, and other shapes may be used.

[0258] Although the process of generating a predicted image from one reference picture has been described here, the process is similar when generating a predicted image from multiple reference pictures, and a luminance correction process may be performed on the reference images obtained from each reference picture in a manner similar to that described above before generating a predicted image.

[0259] As a method of determining whether or not to apply LIC processing, for example, there is a method of using lic_flag, which is a signal indicating whether or not to apply LIC processing. As a specific example, in an encoding device, it is determined whether or not the current block belongs to an area where a luminance change occurs, and if it belongs to an area where a luminance change occurs, a value of 1 is set as lic_flag and LIC processing is applied and encoding is performed, and if it does not belong to an area where a luminance change occurs, a value of 0 is set as lic_flag and encoding is performed without applying LIC processing. On the other hand, a decoding device may decode lic_flag described in a stream, and switch whether or not to apply LIC processing depending on the value and perform decoding.

[0260] Another method of determining whether to apply LIC processing is, for example, a method of determining according to whether LIC processing has been applied to surrounding blocks.As a specific example, when the current block is in merge mode, it is determined whether the surrounding coded blocks selected when deriving MV in merge mode processing have been coded by applying LIC processing.Depending on the result, it is switched to whether to apply LIC processing and performs coding.In this example, the same processing is also applied to the processing on the decoding device side.

[0261] The aspect of the LIC process (luminance correction process) has been described with reference to FIG. 39, and will be described in detail below.

[0262] First, the inter prediction unit 126 derives a motion vector for obtaining a reference image corresponding to the current block to be coded from a reference picture that is a coded picture.

[0263] Next, the inter prediction unit 126 uses the luminance pixel values ​​of the coded surrounding reference areas adjacent to the left and above the coding target block and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the motion vector to extract information indicating how the luminance values ​​have changed between the reference picture and the coding target picture, and calculates a luminance correction parameter. For example, the luminance pixel value of a pixel in the surrounding reference area in the coding target picture is set to p0, and the luminance pixel value of a pixel in the surrounding reference area in the reference picture at the equivalent position to the pixel is set to p1. The inter prediction unit 126 calculates coefficients A and B that optimize A×p1+B=p0 for multiple pixels in the surrounding reference areas as the luminance correction parameter.

[0264] Next, the inter prediction unit 126 performs luminance correction processing on a reference image in a reference picture specified by the motion vector using the luminance correction parameter to generate a predicted image for the block to be coded. For example, the luminance pixel value in the reference image is p2, and the luminance pixel value of the predicted image after the luminance correction processing is p3. The inter prediction unit 126 calculates A×p2+B=p3 for each pixel in the reference image to generate a predicted image after the luminance correction processing.

[0265] The shape of the surrounding reference area in FIG. 39 is an example, and other shapes may be used. A part of the surrounding reference area shown in FIG. 39 may be used. For example, an area including a predetermined number of pixels thinned out from each of the upper adjacent pixel and the left adjacent pixel may be used as the surrounding reference area. The surrounding reference area is not limited to an area adjacent to the encoding target block, and may be an area not adjacent to the encoding target block. The predetermined number of pixels may be determined in advance.

[0266] In the example shown in Figure 39, the surrounding reference area in the reference picture is an area specified by the motion vector of the encoding target picture from the surrounding reference area in the encoding target picture, but may be an area specified by another motion vector. For example, the other motion vector may be the motion vector of the surrounding reference area in the encoding target picture.

[0267] Although the operation of the encoding device 100 has been described above, the operation of the decoding device 200 is typically similar.

[0268] The LIC process may be applied to color difference instead of luminance. In this case, correction parameters may be derived for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.

[0269] Alternatively, the LIC process may be applied on a subblock basis. For example, the correction parameters may be derived using a surrounding reference region of the current subblock and a surrounding reference region of a reference subblock in a reference picture specified by the MV of the current subblock.

[0270] [Predictive control unit] The prediction control unit 128 selects either an intra-prediction signal (a signal output from the intra-prediction unit 124) or an inter-prediction signal (a signal output from the inter-prediction unit 126), and outputs the selected signal as a prediction signal to the subtraction unit 104 and the addition unit 116.

[0271] As shown in FIG. 1, in various exemplary encoding devices, the prediction control unit 128 may output prediction parameters to be input to the entropy coding unit 110. The entropy coding unit 110 may generate an encoded bitstream (or sequence) based on the prediction parameters input from the prediction control unit 128 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by a decoding device. The decoding device may receive and decode the encoded bitstream and perform the same prediction process as that performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., a motion vector, a prediction type, or a prediction mode used in the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or indicating the prediction process performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0272] [Example of an encoding device implementation] 40 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, several components of the encoding device 100 shown in FIG. 1 are implemented by the processor a1 and the memory a2 shown in FIG.

[0273] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes moving images. The processor a1 may be a processor such as a CPU. The processor a1 may also be a collection of multiple electronic circuits. For example, the processor a1 may play the role of multiple components among multiple components of the encoding device 100 shown in FIG. 1 etc.

[0274] The memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode a moving image is stored. The memory a2 may be an electronic circuit and may be connected to the processor a1. The memory a2 may be included in the processor a1. The memory a2 may be a collection of multiple electronic circuits. The memory a2 may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium, etc. The memory a2 may be a non-volatile memory or a volatile memory.

[0275] For example, the memory a2 may store a video to be encoded, or a bit string corresponding to the encoded video, or may store a program for the processor a1 to encode the video.

[0276] Also, for example, the memory a2 may play the role of a component for storing information among the multiple components of the encoding device 100 shown in Fig. 1 etc. For example, the memory a2 may play the role of the block memory 118 and the frame memory 122 shown in Fig. 1. More specifically, the memory a2 may store reconstructed blocks, reconstructed pictures, etc.

[0277] It should be noted that not all of the components shown in Fig. 1 and the like may be implemented, and not all of the processes described above may be performed, in the encoding device 100. Some of the components shown in Fig. 1 and the like may be included in another device, and some of the processes described above may be executed by another device.

[0278] [Decryption device] Next, a description will be given of a decoding device capable of decoding a coded signal (coded bit stream) outputted from, for example, the above coding device 100. Fig. 41 is a block diagram showing a functional configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a video decoding device that decodes a video on a block-by-block basis.

[0279] As shown in FIG. 41, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220.

[0280] The decoding device 200 is realized by, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The decoding device 200 may also be realized as one or more dedicated electronic circuits corresponding to the entropy decoding unit 202, the inverse quantization unit 204, the inverse transform unit 206, the addition unit 208, the loop filter unit 212, the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0281] Below, the overall processing flow of the decoding device 200 will be described, and then each component included in the decoding device 200 will be described.

[0282] [Overall flow of decryption process] FIG. 42 is a flowchart showing an example of the overall decoding process by the decoding device 200.

[0283] First, the entropy decoding unit 202 of the decoding device 200 identifies a division pattern of fixed-size blocks (e.g., 128×128 pixels) (step Sp_1). This division pattern is the division pattern selected by the encoding device 100. Then, the decoding device 200 performs the processes of steps Sp_2 to Sp_6 on each of the multiple blocks constituting the division pattern.

[0284] That is, the entropy decoding unit 202 decodes (specifically, entropy decodes) the coded quantized coefficients and prediction parameters of the block to be decoded (also called the current block) (step Sp_2).

[0285] Next, the inverse quantization unit 204 and the inverse transform unit 206 perform inverse quantization and inverse transform on the multiple quantized coefficients to reconstruct multiple prediction residuals (that is, difference blocks) (step Sp_3).

[0286] Next, a prediction processing unit consisting of all or a part of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction signal (also called a prediction block) of the current block (step Sp_4).

[0287] Next, the adder 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted block to the difference block (step Sp_5).

[0288] Then, when this reconstructed image is generated, the loop filter unit 212 performs filtering on the reconstructed image (step Sp_6).

[0289] Then, the decoding device 200 determines whether or not the decoding of the entire picture is completed (step Sp_7), and if it determines that the decoding is not completed (No in step Sp_7), it repeats the process from step Sp_1.

[0290] As illustrated, the processes of steps Sp_1 to Sp_7 are sequentially performed by the decoding device 200. Alternatively, some of the processes may be performed in parallel, or the order of the processes may be changed.

[0291] [Entropy Decoding Part] The entropy decoding unit 202 entropy decodes the coded bitstream. Specifically, for example, the entropy decoding unit 202 arithmetically decodes the coded bitstream into a binary signal. Then, the entropy decoding unit 202 debinarizes the binary signal. The entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 on a block-by-block basis. The entropy decoding unit 202 may output prediction parameters included in the coded bitstream (see FIG. 1) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 in the embodiment. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can execute the same prediction process as the process executed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the coding device side.

[0292] [Dequantization section] The inverse quantization unit 204 inverse quantizes the quantized coefficients of a block to be decoded (hereinafter, referred to as a current block) that is input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 inverse quantizes each quantized coefficient of the current block based on a quantization parameter corresponding to the quantized coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantized coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0293] [Inverse conversion section] The inverse transform unit 206 restores the prediction error by inverse transforming the transform coefficients input from the inverse quantization unit 204 .

[0294] For example, if the information interpreted from the encoded bitstream indicates that EMT or AMT is to be applied (e.g., the AMT flag is true), the inverse transform unit 206 inverse transforms the transform coefficients of the current block based on the interpreted information indicating the transform type.

[0295] Also for example, if the information interpreted from the coded bitstream indicates to apply NSST, then inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0296] [Addition section] The adder 208 reconstructs the current block by adding the prediction error, which is an input from the inverse transformer 206, and the prediction sample, which is an input from the prediction control unit 220. The adder 208 then outputs the reconstructed block to the block memory 210 and the loop filter unit 212.

[0297] [Block memory] The block memory 210 is a storage unit for storing blocks that are referenced in intra prediction and are in a picture to be decoded (hereinafter, referred to as a current picture). Specifically, the block memory 210 stores the reconstructed blocks output from the adder 208.

[0298] [Loop filter section] The loop filter unit 212 applies a loop filter to the block reconstructed by the adder unit 208, and outputs the filtered reconstructed block to a frame memory 214, a display device, or the like.

[0299] If the information indicating ALF on / off read from the encoded bitstream indicates ALF on, one filter is selected from among multiple filters based on the local gradient direction and activity, and the selected filter is applied to the reconstructed block.

[0300] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used in inter prediction, and may be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed blocks filtered by the loop filter unit 212.

[0301] [Prediction processing unit (intra prediction unit, inter prediction unit, prediction control unit)] 43 is a flowchart showing an example of processing performed in the prediction processing unit of the decoding device 200. Note that the prediction processing unit is made up of all or some of the components of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220.

[0302] The prediction processing unit generates a prediction image of the current block (step Sq_1). This prediction image is also called a prediction signal or a prediction block. The prediction signal includes, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction processing unit generates a prediction image of the current block using a reconstructed image that has already been obtained by generating a prediction block, generating a difference block, generating a coefficient block, restoring the difference block, and generating a decoded image block.

[0303] The reconstructed image may be, for example, an image of a reference picture, or an image of a decoded block in a current picture, which is a picture that includes the current block. The decoded block in the current picture may be, for example, a neighboring block of the current block.

[0304] FIG. 44 is a flowchart showing another example of the process performed by the prediction processing unit of the decoding device 200.

[0305] The prediction processing unit determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode may be determined based on prediction parameters, for example.

[0306] When the prediction processing unit determines the first method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the first method (step Sr_2a). When the prediction processing unit determines the second method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the second method (step Sr_2b). When the prediction processing unit determines the third method as the mode for generating the predicted image, the prediction processing unit generates the predicted image according to the third method (step Sr_2c).

[0307] The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter-prediction method, an intra-prediction method, and other prediction methods. These prediction methods may use the above-mentioned reconstructed image.

[0308] [Intra prediction section] The intra prediction unit 216 generates a prediction signal (intra prediction signal) by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on an intra prediction mode interpreted from the encoded bit stream. Specifically, the intra prediction unit 216 generates an intra prediction signal by performing intra prediction with reference to samples (e.g., luminance values, chrominance values) of blocks adjacent to the current block, and outputs the intra prediction signal to the prediction control unit 220.

[0309] Note that, when an intra prediction mode that references a luminance block in intra prediction of a chrominance block is selected, the intra prediction unit 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0310] Furthermore, when information interpreted from the encoded bitstream indicates the application of PDPC, the intra prediction unit 216 corrects pixel values ​​after intra prediction based on the gradients of reference pixels in the horizontal / vertical directions.

[0311] [Inter prediction section] The inter prediction unit 218 predicts the current block by referring to a reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks (e.g., 4x4 blocks) in the current block. For example, the inter prediction unit 218 generates an inter prediction signal of the current block or sub-block by performing motion compensation using motion information (e.g., motion vectors) interpreted from the encoded bit stream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter prediction signal to the prediction control unit 220.

[0312] If the information interpreted from the encoded bitstream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction signal using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks.

[0313] Also, if the information interpreted from the encoded bitstream indicates that the FRUC mode is applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) interpreted from the encoded bitstream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0314] In addition, when the BIO mode is applied, the inter prediction unit 218 derives a motion vector based on a model assuming uniform linear motion. In addition, when information interpreted from the encoded bitstream indicates that an affine motion compensation prediction mode is applied, the inter prediction unit 218 derives a motion vector on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0315] [MV Derivation > Normal Inter Mode] If the information interpreted from the encoded bitstream indicates that normal inter mode is to be applied, the inter prediction unit 218 derives an MV based on the information interpreted from the encoded bitstream, and performs motion compensation (prediction) using the MV.

[0316] FIG. 45 is a flowchart showing an example of inter prediction in the normal inter mode in the decoding device 200.

[0317] The inter prediction unit 218 of the decoding device 200 performs motion compensation for each block. The inter prediction unit 218 obtains multiple candidate MVs for the current block based on information such as MVs of multiple decoded blocks around the current block in time or space (step Ss_1). That is, the inter prediction unit 218 creates a candidate MV list.

[0318] Next, the inter prediction unit 218 extracts N candidate MVs (N is an integer equal to or greater than 2) from the multiple candidate MVs acquired in step Ss_1 as motion vector predictor candidates (also called prediction MV candidates) according to a predetermined priority order (step Ss_2). Note that the priority order may be predefined for each of the N prediction MV candidates.

[0319] Next, the inter prediction unit 218 decodes the predicted motion vector selection information from the input stream (i.e., the encoded bit stream), and uses the decoded predicted motion vector selection information to select one predicted MV candidate from the N predicted MV candidates as the predicted motion vector (also called predicted MV) of the current block (step Ss_3).

[0320] Next, the inter prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the differential value, which is the decoded differential MV, to the selected predicted motion vector (step Ss_4).

[0321] Finally, the inter prediction unit 218 performs motion compensation on the current block using the derived MV and the decoded reference picture to generate a predicted image of the current block (step Ss_5).

[0322] [Predictive control unit] The prediction control unit 220 selects either the intra prediction signal or the inter prediction signal, and outputs the selected signal as a prediction signal to the adder unit 208. Overall, the configurations, functions, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding device side may correspond to the configurations, functions, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding device side.

[0323] [Example of implementation of a decryption device] Fig. 46 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, several components of the decoding device 200 shown in Fig. 41 are implemented by the processor b1 and the memory b2 shown in Fig. 46.

[0324] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes encoded video (i.e., an encoded bitstream). The processor b1 may be a processor such as a CPU. The processor b1 may also be a collection of multiple electronic circuits. For example, the processor b1 may play the role of multiple components among multiple components of the decoding device 200 shown in FIG. 41 etc.

[0325] The memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode the encoded bit stream is stored. The memory b2 may be an electronic circuit and may be connected to the processor b1. The memory b2 may be included in the processor b1. The memory b2 may be a collection of multiple electronic circuits. The memory b2 may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium, etc. The memory b2 may be a non-volatile memory or a volatile memory.

[0326] For example, the memory b2 may store a video image or a coded bitstream, and may store a program for the processor b1 to decode the coded bitstream.

[0327] Also, for example, the memory b2 may play the role of a component for storing information among the multiple components of the decoding device 200 shown in Fig. 41 etc. Specifically, the memory b2 may play the role of the block memory 210 and the frame memory 214 shown in Fig. 41. More specifically, the memory b2 may store reconstructed blocks, reconstructed pictures, etc.

[0328] Note that not all of the components shown in Fig. 41 and the like may be implemented, and not all of the above-described processes may be performed, in the decoding device 200. Some of the components shown in Fig. 41 and the like may be included in another device, and some of the above-described processes may be executed by another device.

[0329] [Definition of each term] As an example, each term may be defined as follows:

[0330] A picture is an array of luma samples in monochrome format, or two corresponding arrays of luma samples and chroma samples in 4:2:0, 4:2:2 and 4:4:4 color formats. A picture may be a frame or a field.

[0331] A frame is a composition of a top field from which a number of sample rows occur: 0, 2, 4, . . . and a bottom field from which a number of sample rows occur: 1, 3, 5, . . .

[0332] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit.

[0333] A tile is a rectangular region of multiple coding tree blocks in a particular tile column and a particular tile row in a picture. A tile may also be a rectangular region of a frame that is intended to be decoded and coded independently, although loop filters across tile edges may still be applied.

[0334] A block is an MxN (N rows and M columns) array of samples, or an MxN array of transform coefficients. A block may be a square or rectangular region of pixels consisting of one luma and two chroma matrices.

[0335] A CTU (coding tree unit) may be a coding tree block of luma samples for a picture with three sample arrangements, or two corresponding coding tree blocks of chroma samples, or a coding tree block of samples for either monochrome pictures or pictures coded with three separated color planes and a syntax structure used for coding the samples.

[0336] A superblock may comprise one or two mode information blocks, or may be a square block of 64x64 pixels that can be recursively divided into four 32x32 blocks and further divided.

[0337] [Non-rectangular division] In the prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126 of the encoding device (see FIG. 1), as well as in the prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218 of the decoding device (see FIG. 41), conventionally, the partitions (or the variable size blocks or the sub-blocks) obtained from the division of each block and from which motion information (e.g., motion vectors) is obtained are always rectangular, as shown in FIG. 3. The inventors have found that generating partitions with non-rectangular shapes, such as triangular shapes, leads to improved image quality and coding efficiency in various implementations, depending on the image content in the picture. In the following, various embodiments are described in which at least one partition divided from an image block for prediction purposes has a non-rectangular shape. Note that these embodiments are equally applicable to the encoding device side (prediction control unit 128 connected to the intra prediction unit 124 and the inter prediction unit 126) and to the decoding device side (prediction control unit 220 connected to the intra prediction unit 216 and the inter prediction unit 218), and may be implemented in the encoding device of Figure 1 or the decoding device of Figure 41.

[0338] Figure 47 is a flowchart showing an example of a process for dividing an image block into multiple partitions including at least a first partition having a non-rectangular shape (e.g., a triangle) and a second partition, and then encoding (or decoding) the image block as a reconstructed combination of the first and second partitions.

[0339] In step S1001, an image block is divided into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. For example, as shown in FIG. 48, the image block may be divided from the top left corner of the image block to the bottom right corner of the image block to create a first partition and a second partition, both of which have a non-rectangular shape (e.g., a triangle). Alternatively, the image block may be divided from the top right corner of the image block to the bottom left corner of the image block to create a first partition and a second partition, both of which have a non-rectangular shape (e.g., a triangle). Various examples of non-rectangular division are described later with reference to FIG. 48 and FIG. 53-FIG. 55.

[0340] In step S1002, the process predicts a first motion vector for the first partition and predicts a second motion vector for the second partition. For example, predicting the first and second motion vectors may include selecting a first motion vector from a first set of motion vector candidates and selecting a second motion vector from a second set of motion vector candidates.

[0341] In step S1003, a motion compensation process is performed to obtain a first partition using the first motion vector derived in step S1002 above, and to obtain a second partition using the second motion vector derived in step S1002 above.

[0342] In step S1004, a prediction process is performed on the image block as a (reconstructed) combination of the first and second partitions. The prediction process includes a boundary smoothing process to smooth the boundary between the first and second partitions. For example, the boundary smoothing process involves weighting a plurality of first values ​​of a plurality of boundary pixels predicted based on the first partition and a plurality of second values ​​of a plurality of boundary pixels predicted based on the second partition. Various implementations of the boundary smoothing process are described later with reference to Figures 49, 50, 56, and 57A-57D.

[0343] In step S1005, the process encodes or decodes the image block using one or more parameters including a partition parameter indicating partitioning of the image block into a first partition having a non-rectangular shape and a second partition. As summarized in the table of FIG. 51, for example, the partition parameter ("first index value") may encode, for example, a partition direction to be applied for partitioning (e.g., from upper left to lower right or from upper right to lower left as shown in FIG. 48) together with the first and second motion vectors derived in step S1002 described above. Details of such partition syntax operations involving one or more parameters including the partition parameter are described in detail below with reference to FIG. 51, FIG. 52, and FIG. 58-FIG. 61.

[0344] FIG. 53 is a flow chart showing a process 2000 for dividing an image block. In step S2001, the process divides an image into a plurality of partitions including a first partition having a non-rectangular shape and a second partition that may or may not have a non-rectangular shape. As shown in FIG. 48, the image block is divided into a first partition having a triangular shape and a second partition also having a triangular shape. There are many other examples in which an image block is divided into a plurality of partitions including a first partition and a second partition, where at least the first partition has a non-rectangular shape. The non-rectangular shape may be a triangle, a trapezoid, or a polygon having at least five sides and angles.

[0345] For example, as shown in Figure 54, an image block may be divided into two triangular shaped partitions. An image block may be divided into more than two triangular shaped partitions (e.g., three triangular shaped partitions). An image block may be divided into a combination of one or more triangular shaped partitions and one or more rectangular shaped partitions. Alternatively, an image block may be divided into a combination of one or more triangular shaped partitions and one or more polygonal shaped partitions.

[0346] As further shown in Figure 55, the image block may be partitioned into L-shaped (polygonal) partitions and rectangular shaped partitions. The image block may be partitioned into pentagonal (polygonal) shaped partitions and triangular shaped partitions. The image block may be partitioned into hexagonal (polygonal) shaped partitions and pentagonal (polygonal) shaped partitions. Alternatively, the image block may be partitioned into multiple polygonal shaped partitions.

[0347] Referring again to FIG. 53, in step S2002, the process predicts a first motion vector for the first partition, for example by selecting a first partition from a first motion vector candidate set, and predicts a second motion vector for the second partition, for example by selecting a second partition from a second motion vector candidate set. For example, the first motion vector candidate set may include motion vectors of partitions adjacent to the first partition, and the second motion vector candidate set may include motion vectors of partitions adjacent to the second partition. The adjacent partitions may be one or both of spatially adjacent partitions and temporally adjacent partitions. Some examples of spatially adjacent partitions include partitions located to the left, bottom left, bottom, bottom right, right, top right, top or top left of the partition being processed. Examples of temporally adjacent partitions include co-located partitions in reference pictures of the image block.

[0348] In various implementations, the partitions adjacent to the first partition and the partitions adjacent to the second partition may be outside the image block divided into the first partition and the second partition. The first motion vector candidate set may be the same as or different from the second motion vector candidate set. Furthermore, at least one of the first motion vector candidate set and the second motion vector candidate set may be the same as another third motion vector candidate set prepared for the image block.

[0349] In some implementations, in response to determining in step S2002 that the second partition has a non-rectangular shape (e.g., a triangle) similar to the first partition, the process 2000 creates (for the non-rectangular shaped second partition) a second motion vector candidate set including a plurality of motion vectors of a plurality of partitions adjacent to the second partition, excluding the first partition (i.e., excluding the motion vector of the first partition). On the other hand, in response to determining that the second partition has a rectangular shape, but not similar to the first partition, the process 2000 creates (for the rectangular shaped second partition) a second motion vector candidate set including a plurality of motion vectors of a plurality of partitions adjacent to the second partition, including the first partition.

[0350] In step S2003, the process encodes or decodes the first partition using the first motion vector derived in step S2002 described above, and encodes or decodes the second partition using the second motion vector derived in step S2002 described above.

[0351] An image block division process such as process 2000 in Figure 53 may be performed by an image encoding device including a circuit and a memory connected to the circuit, such as that shown in Figure 1. The circuit operates to divide an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape (step S2001), predict a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encode the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0352] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a division unit 102 for receiving an original image and dividing it into a plurality of blocks in an operation, an adder unit 104 for receiving the plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128 in an operation, subtracting each prediction from its corresponding block, and outputting a residual, a transform unit 106 for performing a transform on the plurality of residuals output from the adder unit 104 in an operation, and outputting a plurality of transform coefficients, a quantization unit 108 for quantizing the plurality of transform coefficients in an operation to generate a plurality of quantized transform coefficients, an entropy coding unit 110 for coding the plurality of quantized transform coefficients in an operation to generate a bitstream, and an inter prediction unit 126 for generating a prediction of a current block based on a reference block in a coded reference picture in an operation, an intra prediction unit 124 for generating a prediction of the current block based on a coded reference block in the current picture in an operation, and a prediction control unit 128 connected to the memories 118, 122. In operation, the prediction control unit 128 divides multiple blocks into multiple partitions including a first partition having a non-rectangular shape and a second partition (Figure 53, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and encodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0353] According to another embodiment, there is provided an image decoding device including a circuit and a memory connected to the circuit, for example as shown in Fig. 41. The circuit, in operation, divides an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape (Fig. 53, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0354] Further according to an embodiment, the image decoding device shown in FIG. 41 is provided including: an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantized transform coefficients; an inverse quantization unit 204 and an inverse transform unit 206 that in operation inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an addition unit 208 that in operation adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks; an inter prediction unit 218 that in operation generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit 216 that in operation generates a prediction of the current block based on a decoded reference block in the current picture; and a prediction control unit 220 connected to the memories 210, 214. In operation, the prediction control unit 220 divides an image block into multiple partitions including a first partition and a second partition having a non-rectangular shape (Figure 53, step S2001), predicts a first motion vector for the first partition and a second motion vector for the second partition (step S2002), and decodes the first partition using the first motion vector and the second partition using the second motion vector (step S2003).

[0355] [Boundary smoothing] As described above in Figure 47, step S1004, according to various embodiments, performing prediction processing on an image block as a (reconstructed) combination of a first partition having a non-rectangular shape and a second partition may involve application of a boundary smoothing process along the boundary between the first partition and the second partition.

[0356] For example, Figure 57B shows an example of a boundary smoothing process that involves weighting a plurality of first values ​​of a plurality of boundary pixels that are first predicted based on a first partition and a plurality of second values ​​of a plurality of boundary pixels that are second predicted based on a second partition.

[0357] Figure 56 is a flow chart illustrating an overall boundary smoothing process 3000 involving weighting first values ​​of a first predicted plurality of boundary pixels based on a first partition and second values ​​of a second predicted plurality of boundary pixels based on a second partition according to one embodiment. In step S3001, an image block is divided along a boundary into a first partition and a second partition, where at least the first partition has a non-rectangular shape, as shown in Figure 57A or in Figures 48, 54 and 55 described above.

[0358] In step S3002, a plurality of first values ​​(e.g., color, luminance, transparency, etc.) of a set of pixels of a first partition along the boundary ("boundary pixels" in FIG. 57A) are first predicted using information of the first partition. In step S3003, a plurality of second values ​​of a (same) set of pixels of the first partition along the boundary are second predicted using information of the second partition. In some embodiments, at least one of the first prediction and the second prediction is an inter-prediction process that predicts the plurality of first values ​​and the plurality of second values ​​based on a reference partition in a coded reference picture. With reference to FIG. 57D, in some implementations, the prediction process predicts the plurality of first values ​​of all pixels of a first partition that includes a set of pixels where the first partition and the second partition overlap ("first sample set"), and predicts the second values ​​only of a set of pixels where the first partition and the second partition overlap ("second sample set"). In other implementations, at least one of the first prediction and the second prediction is an intra prediction process that predicts the first values ​​and the second values ​​based on a coded reference partition in the current picture. In some implementations, a prediction method used for the first prediction is different from a prediction method used for the second prediction. For example, the first prediction may include an inter prediction process, and the second prediction may include an intra prediction process. Information used for the first prediction of the first values ​​or the second prediction of the second values ​​may be motion vectors of the first partition or the second partition, multiple intra prediction directions, etc.

[0359] In step S3004, the first values ​​predicted using the first partition and the second values ​​predicted using the second partition are weighted. In step S3005, the first partition is encoded or decoded using the weighted first and second values.

[0360] FIG. 57B illustrates an example of a boundary smoothing operation in which the first partition and the second partition overlap at (maximum) 5 pixels in each row or at (maximum) 5 pixels in each row. That is, the number of pixel sets in each row or column for which the first values ​​are predicted based on the first partition and the second values ​​are predicted based on the second partition is at most 5. FIG. 57C illustrates another example of a boundary smoothing operation in which the first partition and the second partition overlap at (maximum) 3 pixels in each row or column. That is, the number of pixel sets in each row or column for which the first values ​​are predicted based on the first partition and the second values ​​are predicted based on the second partition is at most 3.

[0361] 49 shows another example of a boundary smoothing operation in which the first partition and the second partition overlap at (maximum) 4 pixels in each row or column. That is, the number of pixel sets in each row or column for which the first values ​​are predicted based on the first partition and the second values ​​are predicted based on the second partition is at most 4. In the shown example, weights of 1 / 8, 1 / 4, 3 / 4 and 7 / 8 may be applied to the first values ​​of the four pixels in the set, respectively, and weights of 7 / 8, 3 / 4, 1 / 4 and 1 / 8 may be applied to the second values ​​of the four pixels in the set, respectively.

[0362] FIG. 50 further illustrates examples of boundary smoothing operations in which the first and second partitions overlap at 0 pixels in each row or column (i.e., they do not overlap), overlap at (at most) 1 pixel in each row or column, and overlap at (at most) 2 pixels in each row or column. In examples in which the first and second partitions do not overlap, zero weights are applied. In examples in which the first and second partitions overlap at 1 pixel in each row or column, 1 / 2 weights may be applied to the first values ​​of the pixels in the set predicted based on the first partition, and 1 / 2 weights may be applied to the second values ​​of the pixels in the set predicted based on the second partition. In an example where the first and second partitions overlap by two pixels in each row or column, weights of 1 / 3 and 2 / 3 may be applied to multiple first values ​​of two pixels in the set predicted based on the first partition, respectively, and weights of 2 / 3 and 1 / 3 may be applied to multiple second values ​​of two pixels in the set predicted based on the second partition, respectively.

[0363] In accordance with the embodiments described above, the number of pixels in the set where the first partition and the second partition overlap is an integer. In other implementations, the number of overlapping pixels in the set may be, for example, a non-integer or a fraction. The weights applied to the first values ​​and second values ​​of the pixel set may also be fractional or integer, depending on the application.

[0364] A boundary smoothing process such as the process 3000 of FIG. 56 may be performed by an image encoding device including a circuit such as that shown in FIG. 1 and a memory connected to the circuit. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape divided from an image block and a second partition (FIG. 56, step S3001). The boundary smoothing operation includes predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information of the first partition (step S3002), predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0365] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a division unit 102 for receiving an original image and dividing it into a plurality of blocks in an operation, an adder unit 104 for receiving the plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128 in an operation, subtracting each prediction from its corresponding block, and outputting a residual, a transform unit 106 for performing a transform on the plurality of residuals output from the adder unit 104 in an operation, and outputting a plurality of transform coefficients, a quantization unit 108 for quantizing the plurality of transform coefficients in an operation to generate a plurality of quantized transform coefficients, an entropy coding unit 110 for coding the plurality of quantized transform coefficients in an operation to generate a bitstream, and an inter prediction unit 126 for generating a prediction of a current block based on a reference block in a coded reference picture in an operation, an intra prediction unit 124 for generating a prediction of the current block based on a coded reference block in the current picture in an operation, and a prediction control unit 128 connected to the memories 118, 122. In operation, the prediction control unit 128 performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape divided from an image block and a second partition (FIG. 56, step S3001). The boundary smoothing operation includes first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and encoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0366] According to another embodiment, an image decoding device is provided, including a circuit and a memory connected to the circuit, for example as shown in FIG. 41. In operation, the circuit performs a boundary smoothing operation along a boundary between a first partition having a non-rectangular shape divided from an image block and a second partition (FIG. 56, step S3001). The boundary smoothing operation includes first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0367] According to another embodiment, an image decoding device shown in FIG. 41 is provided including an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantized transform coefficients, an inverse quantization unit 204 and an inverse transform unit 206 that in operation inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals, an adder 208 that in operation adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks, an inter prediction unit 218 that in operation generates a prediction of a current block based on a reference block in a decoded reference picture, an intra prediction unit 216 that in operation generates a prediction of the current block based on a decoded reference block in the current picture, and a prediction control unit 220 connected to the memories 210, 214. In operation, the prediction control unit 220 performs a boundary smoothing operation along the boundary between a first partition having a non-rectangular shape divided from an image block and a second partition (FIG. 56, step S3001). The boundary smoothing operation includes first predicting a plurality of first values ​​of a pixel set of the first partition along the boundary using information of the first partition (step S3002), second predicting a plurality of second values ​​of a pixel set of the first partition along the boundary using information of the second partition (step S3003), weighting the plurality of first values ​​and the plurality of second values ​​(step S3004), and decoding the first partition using the weighted plurality of first values ​​and the weighted plurality of second values ​​(step S3005).

[0368] [Entropy encoding and decoding using partition parameter syntax] As shown in Figure 47, step S1005, according to various embodiments, an image block divided into a first partition and a second partition having a non-rectangular shape may be encoded or decoded using one or more parameters including a partition parameter indicating the non-rectangular partition of the image block. In various embodiments, such partition parameters may be encoded jointly with, for example, a partition direction applied to the partition (e.g., from upper left to lower right or from upper right to lower left, see Figure 48) and the first and second motion vectors predicted in step S1002, as described more fully below.

[0369] FIG. 51 is a table of sample partition parameters ("first index values") and information sets respectively coded together by the partition parameters. The partition parameters ("first index values") range from 0 to 6 and code together the direction of dividing the image block into a first partition and a second partition, both of which are triangular (see FIG. 48), a first motion vector predicted for the first partition (FIG. 47, step S1002), and a second motion vector predicted for the second partition (FIG. 47, step S1002). In particular, partition parameter 0 codes that the division direction is from the top left corner to the bottom right corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0370] Partition parameter 1 encodes that the partition direction is from the top right corner to the bottom left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the partition direction is from the top right corner to the bottom left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the partition direction is from the top left corner to the bottom right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the division direction is from the top right corner to the bottom left corner, that the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "third" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the division direction is from the top left corner to the bottom right corner, that the first motion vector is the "third" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 6 encodes that the division direction is from the top left corner to the bottom right corner, that the first motion vector is the "fourth" motion vector listed in the first motion vector candidate set for the first partition, and that the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition.

[0371] FIG. 58 is a flow chart showing a method 4000 performed on the encoding device side. In step S4001, the process divides an image block into a plurality of partitions, including a first partition having a non-rectangular shape and a second partition, based on a partition parameter indicating the division. For example, as shown in FIG. 51 described above, the partition parameter may indicate a direction in which the image block is divided (e.g., from the upper right corner to the lower left corner, or from the upper left corner to the lower right corner). In step S4002, the process encodes the first partition and the second partition. In step S4003, the process writes one or more parameters including the partition parameter into a bitstream that can be received and decoded by the decoding device side to obtain one or more parameters and perform the same prediction process (performed on the encoding device side) for the first partition and the second partition on the decoding device side. The one or more parameters including the partition parameter encode various information, such as the non-rectangular shape of the first partition, the shape of the second partition, the division direction used to divide the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., jointly or separately.

[0372] FIG. 59 is a flow chart showing a method 5000 performed on the decoding device side. In step S5001, the process reads one or more parameters from a bitstream, the one or more parameters including partition parameters indicating partitioning of an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition. The one or more parameters including the partition parameters read from the bitstream may encode various information required for the decoding device side to perform the same prediction process as that performed on the encoding device side, such as the non-rectangular shape of the first partition, the shape of the second partition, the partition direction used to divide the image block to obtain the first partition and the second partition, the first motion vector of the first partition, the second motion vector of the second partition, etc., together or separately. In step S5002, the process 5000 divides the image block into a plurality of partitions based on the partition parameters read from the bitstream. In step S5003, the process decodes the first partition and the second partition as divided from the image block.

[0373] FIG. 60 is a table of sample partition parameters ("first index value") and information sets encoded together by partition parameters, respectively, similar in characteristics to the sample table described above in FIG. 51. In FIG. 60, the partition parameters ("first index value") range from 0 to 6, and encode the shape of the first and second partitions divided from the image block, the direction of dividing the image block into the first and second partitions, the first motion vector predicted for the first partition (FIG. 47, step S1002), and the second motion vector predicted for the second partition (FIG. 47, step S1002) together. In particular, the partition parameter 0 encodes that neither the first nor the second partition has a triangular shape, and thus the division direction information is "N / A", the first motion vector information is "N / A", and the second motion vector information is "N / A".

[0374] Partition parameter 1 encodes that the first partition and the second partition are triangular, the division direction is from the upper left corner to the lower right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 2 encodes that the first partition and the second partition are triangular, the division direction is from the upper right corner to the lower left corner, the first motion vector is the "first" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 3 encodes that the first partition and the second partition are triangular, the division direction is from the top right corner to the bottom left corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "first" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 4 encodes that the first partition and the second partition are triangular, the division direction is from the top left corner to the bottom right corner, the first motion vector is the "second" motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the "second" motion vector listed in the second motion vector candidate set for the second partition. Partition parameter 5 encodes that the first partition and the second partition are triangular, the division direction is from the top right corner to the bottom left corner, the first motion vector is the “second” motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the “third” motion vector listed in the second motion vector candidate set for the second partition.Partition parameter 6 encodes that the first partition and the second partition are triangular, the division direction is from the top left corner to the bottom right corner, the first motion vector is the “third” motion vector listed in the first motion vector candidate set for the first partition, and the second motion vector is the “first” motion vector listed in the second motion vector candidate set for the second partition.

[0375] According to some implementations, the partition parameters (index values) may be binarized according to a binarization scheme selected depending on the value of at least one or more parameters. Figure 52 shows an example binarization scheme for binarizing the index values ​​(partition parameter values).

[0376] 61 is a table of example combinations of first and second parameters, where one of the first and second parameters is a partition parameter indicating that the image block is to be divided into a plurality of partitions including a first partition and a second partition having a non-rectangular shape. In this example, the partition parameter may be used to indicate that the image block is to be divided without co-encoding other information encoded by one or more of the other parameters.

[0377] In the first example in Figure 61, the first parameter is used to indicate the image block size, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape. Such a combination of the first and second parameters may be used to indicate, for example, 1) no triangular shape partition when the image block size is larger than 64x64, or 2) no triangular shape partition when the width to height ratio of the image block is larger than 4 (e.g., 64x4).

[0378] In the second example of Figure 61, the first parameter is used to indicate a prediction mode, and the second parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape. Such a combination of the first parameter and the second parameter may be used, for example, to indicate that 1) there is no triangular partition when the image block is coded in intra mode.

[0379] In the third example of Figure 61, the first parameter is used as a partition parameter (flag) to indicate that at least one of the partitions divided from the image block has a triangular shape, and the second parameter is used to indicate a prediction mode. Such a combination of the first parameter and the second parameter may be used, for example, to indicate that the image block should be inter-coded if 1) at least one of the partitions divided from the image block has a triangular shape.

[0380] In the fourth example of Figure 61, the first parameter indicates the motion vector of the neighboring block, and the second parameter is used as a partition parameter indicating the direction of dividing the image block into two triangles. Such a combination of the first and second parameters may be used to indicate, for example, that 1) if the motion vector of the neighboring block is diagonal, the direction of dividing the image block into two triangles is from the upper left corner to the lower right corner.

[0381] In the fifth example of Figure 61, the first parameter indicates the intra prediction direction of the neighboring block, and the second parameter is used as a partition parameter indicating the direction of dividing the image block into two triangles. Such a combination of the first parameter and the second parameter may be used to indicate, for example, that 1) if the intra prediction direction of the neighboring block is reverse diagonal, the direction of dividing the image block into two triangles is from the upper right corner to the lower left corner.

[0382] It should be understood that the tables of one or more parameters including partition parameters and which information is coded together or separately as shown in Figures 51, 60, and 61 are presented as examples only, and that numerous other ways of coding various information together or separately as part of the partition syntax operations described above are within the scope of this disclosure. For example, the partition parameters may indicate that the first partition is a triangle, a trapezoid, or a polygon having at least five sides and angles. The partition parameters may indicate that the second partition has a non-rectangular shape, such as a triangle, a trapezoid, and a polygon having at least five sides and angles. The partition parameters may indicate one or more pieces of information about the partition, such as the non-rectangular shape of the first partition, the shape of the second partition (which may be non-rectangular or rectangular), and the partition direction (e.g., from the top left corner of the image block to its bottom right corner, and from the top right corner of the image block to its bottom left corner) applied to divide the image block into multiple partitions. The partition parameters jointly encode further information, such as a first motion vector of the first partition, a second motion vector of the second partition, an image block size, a prediction mode, a motion vector of a neighboring block, an intra-prediction direction of the neighboring block, etc. Alternatively, any of the further information may be encoded separately by one or more parameters other than the partition parameters.

[0383] A partition syntax operation such as process 4000 of Figure 58 may be performed by an image coding device including a circuit and a memory connected to the circuit, such as that shown in Figure 1. In operation, the circuit performs the partition syntax operation including dividing an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape based on a partition parameter indicating the division (Figure 58, step S4001), encoding the first partition and the second partition (S4002), and writing one or more parameters including the partition parameter to a bitstream (S4003).

[0384] According to another embodiment, an image coding apparatus as shown in FIG. 1 is provided, including a division unit 102 for receiving an original image and dividing it into a plurality of blocks in an operation, an adder unit 104 for receiving the plurality of blocks from the division unit and a plurality of predictions from a prediction control unit 128 in an operation, subtracting each prediction from its corresponding block, and outputting a residual, a transform unit 106 for performing a transform on the plurality of residuals output from the adder unit 104 in an operation, and outputting a plurality of transform coefficients, a quantization unit 108 for quantizing the plurality of transform coefficients in an operation to generate a plurality of quantized transform coefficients, an entropy coding unit 110 for coding the plurality of quantized transform coefficients in an operation to generate a bitstream, and an inter prediction unit 126 for generating a prediction of a current block based on a reference block in a coded reference picture in an operation, an intra prediction unit 124 for generating a prediction of the current block based on a coded reference block in the current picture in an operation, and a prediction control unit 128 connected to the memories 118, 122. In operation, the prediction control unit 128 divides an image block into a plurality of partitions including a first partition and a second partition having a non-rectangular shape based on a partition parameter indicating the division (FIG. 58, step S4001), and encodes the first partition and the second partition (step S4002). In operation, the entropy coding unit 110 writes one or more parameters including the partition parameter into a bitstream (step S4003).

[0385] According to another embodiment, there is provided an image decoding device including a circuit and a memory coupled to the circuit, for example as shown in Fig. 41. In operation, the circuit performs a partition syntax operation including: reading one or more parameters from a bitstream including a partition parameter indicating partitioning of an image block into a plurality of partitions including a first partition having a non-rectangular shape and a second partition (Fig. 59, step S5001), partitioning the image block into the plurality of partitions based on the partition parameter (S5002), and decoding the first partition and the second partition (S5003).

[0386] Further according to an embodiment, the image decoding device shown in FIG. 41 is provided including: an entropy decoding unit 202 that receives and decodes an encoded bitstream in operation to obtain a plurality of quantized transform coefficients; an inverse quantization unit 204 and an inverse transform unit 206 that in operation inverse quantize the plurality of quantized transform coefficients to obtain a plurality of transform coefficients and inverse transform the plurality of transform coefficients to obtain a plurality of residuals; an addition unit 208 that in operation adds the plurality of residuals output from the inverse quantization unit 204 and the inverse transform unit 206 to a plurality of predictions output from a prediction control unit 220 to reconstruct a plurality of blocks; an inter prediction unit 218 that in operation generates a prediction of a current block based on a reference block in a decoded reference picture; an intra prediction unit 216 that in operation generates a prediction of the current block based on a decoded reference block in the current picture; and a prediction control unit 220 connected to the memories 210, 214. In operation, the entropy decoding unit 202, in some implementations in cooperation with the prediction control unit 220, reads one or more parameters from the bitstream including a partition parameter indicating division of the image block into multiple partitions including a first partition having a non-rectangular shape and a second partition (Figure 59, step S5001), divides the image block into multiple partitions based on the partition parameter (S5002), and decodes the first partition and the second partition (S5003).

[0387] According to other examples, the inter predictor may perform the following process.

[0388] All of the motion vector candidates included in the first motion vector candidate set may be multiple uni-predictive motion vectors. That is, the inter prediction unit may determine only multiple uni-predictive motion vectors as multiple motion vector candidates in the first motion vector candidate set.

[0389] The inter predictor may select only a plurality of uni-predictive motion vector candidates from the first motion vector candidate set.

[0390] Only uni-predictive motion vectors may be used to predict small blocks. Bi-predictive motion vectors may be used to predict large blocks. As an example, the prediction process may include determining a size of an image block. If the size of the image block is determined to be greater than a threshold, the prediction may include selecting a first motion vector from a first motion vector candidate set, and the first motion vector candidate set may include uni-predictive motion vectors and / or bi-predictive motion vectors. If the size of the image block is determined to be not greater than a threshold, the prediction may include selecting a first motion vector from a first motion vector candidate set, and the first motion vector candidate set may include only uni-predictive motion vectors.

[0391] [How motion vectors are stored for weighted regions] 62 is a flow chart illustrating an example of a process flow 6000 for predicting a first sample set for a first partition of a current picture with a first motion vector, obtaining a second motion vector from a second partition different from the first partition, predicting a second sample set for a first portion included in the first partition with the second motion vector, weighting the first and second sample sets, storing at least one of the first and second motion vectors for the first partition, encoding or decoding the first partition using the weighted samples, and further processing according to an embodiment. Process flow 6000 may be performed, for example, by encoding device 100 of FIG. 1, decoding device 200 of FIG. 41, etc.

[0392] In step S6001, a first sample set for a first partition of a current picture is predicted with one or more motion vectors including a first motion vector. The first partition may or may not be non-rectangular in shape. The first sample set may be, for example, a plurality of pixel blocks having corresponding motion vectors, luma pixel components, and chroma pixel components. For example, each sample may be a 4x4 pixel block associated with the first motion vector and having corresponding luma pixel components and chroma pixel components.

[0393] The first partition may be, for example, a partition of an image block that is divided into multiple partitions. FIG. 63 is a conceptual diagram illustrating an example of a method for dividing an image block into a first partition and a second partition. See also FIG. 48, FIG. 54, and FIG. 55. For example, as shown in FIG. 63, an image block may be divided into two or more partitions of various shapes. The example diagram in FIG. 63 includes an image block that is divided from the top left corner of the image block to the bottom right corner of the image block to generate a first partition and a second partition that both have a non-rectangular shape (e.g., a triangle), an image block divided into an L-shaped partition and a rectangular-shaped partition, an image block divided into a pentagonal-shaped partition and a triangular-shaped partition, an image block divided into a hexagonal-shaped partition and a pentagonal-shaped partition, and an image block divided into two polygonal-shaped partitions. The various partition shapes illustrated may be formed by dividing the image block in other ways. For example, two triangular shaped partitions may be formed by dividing the image block from the top right corner of the image block to the bottom left corner of the image block to generate a first partition and a second partition that both have a triangular shape. In some embodiments, two or more partitions of an image block may have overlapping portions.

[0394] In some embodiments, the first motion vector may be obtained by a motion estimation process. In some embodiments, the first motion vector may be parsed from a bitstream. In some embodiments, the first motion vector may be a motion vector predicted from a set of motion vector candidates for at least the first partition. The motion vector candidates in the motion vector candidate list may include motion vector candidates derived from at least spatial or temporal neighboring partitions of the first partition, or may be derived from a motion vector candidate list of image blocks for a prediction mode such as merge mode, skip mode, or inter mode.

[0395] Figure 64 is a conceptual diagram illustrating near and non-neighboring spatially adjacent partitions of a first partition of a current picture. Near spatially adjacent partitions are partitions that are adjacent to the first partition of the current picture. Non-neighboring spatially adjacent partitions are partitions that are far from the first partition of the current picture. In some embodiments, a motion vector candidate set may be derived from spatially adjacent partitions of at least the first partition of the current picture.

[0396] The first motion vector may be, for example, a uni-predictive motion vector or a bi-predictive motion vector. Figure 65 is a conceptual diagram illustrating uni-predictive motion vector candidates and bi-predictive motion vector candidates for an image block of a current picture. A uni-predictive motion vector candidate is a single motion vector of a current block in a current picture relative to a single reference picture. As shown in the upper part of Figure 65, a uni-predictive motion vector candidate is a motion vector from a block of the current picture to a block of a reference picture that precedes the current picture in display order. In some embodiments, the reference picture may be after the current picture in display order.

[0397] A bi-predictive motion vector candidate consists of two motion vectors, a first motion vector of a current block relative to a first reference picture and a second motion vector of a current block relative to a second reference picture. As shown in FIG. 65, the bi-predictive motion vector candidate on the lower left side has a first motion vector from a block of a current picture to a block of a first reference picture and a second motion vector from a block of a current picture to a block of a second reference picture. As shown, the first reference picture and the second reference picture are before the current picture in display order. The bi-predictive motion vector candidate on the lower right side of FIG. 65 has a first motion vector from a block of a current picture to a block of a first reference picture and a second motion vector from a block of a current picture to a block of a second reference picture. The first reference picture is before the current picture in display order, and the second reference picture is after the current picture in display order.

[0398] In some embodiments, some partition shapes may be associated with one type of predicted motion vector, for example, in some embodiments, triangular shaped partitions may be associated with uni-predictive motion vector candidates.

[0399] Predicting the first sample set may include a motion compensation process with one or more motion vectors, including the first motion vector, which may predict samples from the current picture (intra block), or from another picture (inter block), or a combination thereof.

[0400] In step S6002, one or more motion vectors including the second motion vector are obtained from a second partition different from the first partition. The second partition may be, for example, a spatially or temporally co-located partition of the first partition. A portion of the second partition may spatially overlap a portion of the first partition. The second motion vector may be obtained by a motion estimation process. In some embodiments, the second motion vector may be analyzed from a bitstream. In some embodiments, the second motion vector may be a motion vector predicted from a motion vector candidate set of at least the second partition. The motion vector candidates of the motion vector candidate list may include motion vector candidates derived from spatially or temporally neighboring partitions of the second partition, or may be derived from a motion vector candidate list of image blocks for a prediction mode such as merge mode, skip mode, or other inter mode. The second motion vector may be a uni-predictive motion vector or a bi-predictive motion vector.

[0401] In step S6003, a second sample set for a portion of the first partition is predicted with one or more motion vectors including the second motion vector. The first portion of the first partition is smaller than the first partition, is adjacent to an edge of the first partition, and overlaps the first sample set. Predicting the second sample set may include a motion compensation process with one or more motion vectors including the second motion vector, and the motion compensation process may predict samples from a current picture (intra block), may predict samples from another picture (inter block), or may be a combination thereof. The second sample set may be, for example, pixel blocks having corresponding motion vectors, luma pixel components, and chroma pixel components. For example, each sample of the second sample set may be a 4×4 pixel block associated with the second motion vector and having corresponding luma pixel components and chroma pixel components.

[0402] FIG. 66 is a conceptual diagram illustrating an example of a first portion of a first partition and a first and second sample sets. The first portion may be, for example, a quarter of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to an edge of the first partition, where N is an integer equal to or greater than 1, for example, 2. As shown, the right example of FIG. 66 illustrates a rectangular partition having a rectangular portion with a width that is a quarter of the width of the first partition, with a first sample set including samples outside the first portion and samples inside the first portion, and a second sample set including samples within the first portion. The center example of FIG. 66 illustrates a rectangular partition having a rectangular portion with a height that is a quarter of the height of the first partition, with a first sample set including samples outside the first portion and samples inside the first portion, and a second sample set including samples within the first portion. The example on the left in Figure 66 shows a triangular partition having a polygonal portion whose height corresponds to two samples, with a first sample set including samples outside the first portion and samples inside the first portion, and a second sample set including samples within the first portion.

[0403] The first portion may be a portion of the first partition that overlaps with an adjacent partition. FIG. 67 is a conceptual diagram showing a first portion of the first partition that is a portion of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition is shown having an overlapping portion with a spatially adjacent rectangular partition. Partitions having other shapes, such as triangular partitions, may also be employed. The overlapping portion may also overlap with a spatially or temporally adjacent partition. The adjacent partition may be a second partition.

[0404] In step S6004, a weighting process is performed on the first portion in the first partition using a subset of the first sample set in the first portion and the second sample set in the first portion. The weighting may be performed as part of a boundary smoothing process, such as an overlapped block motion compensation process (OBMC, see Figs. 35 and 36 and their descriptions), a boundary smoothing process described with reference to Figs. 47 to 61, etc. For example, a plurality of pixels of the first sample set in the first portion may be assigned different weights based on, for example, the distance of the samples from an edge of the first partition to which the first portion corresponds. A smaller weight may be assigned to samples closer to the edge than samples further from the edge. Similarly, a plurality of pixels of the second sample set in the first portion of the first partition may be assigned different weights based on, for example, the distance of the samples from an edge of the first partition to which the first portion corresponds. A larger weight may be assigned to samples closer to the edge than samples further from the edge.

[0405] In step S6005, for the first portion of the first partition, one or more motion vectors based on the first motion vector and / or the second motion vector are stored. In some embodiments, the stored one or more motion vectors may take into account the reference picture lists to which the first motion vector and the second motion vector point.

[0406] For example, FIG. 68 is a conceptual diagram showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to different pictures in different reference picture lists. In some embodiments, when the first motion vector and the second motion vector are uni-predictive motion vectors pointing to different pictures in different reference picture lists, the first motion vector and the second motion vector may be combined and stored as a bi-predictive motion vector for the first part of the first partition. As shown in FIG. 68, the current picture is POC4. The first motion vector Mv1 is a uni-predictive motion vector pointing to reference picture POC0 of reference picture list L0 including reference pictures POC0 to POC8. The second motion vector Mv2 is a uni-predictive motion vector pointing to reference picture POC16 of reference picture list L1 including reference pictures POC8 to POC16. The first motion vector Mv1 and the second motion vector Mv2 are combined and stored as a bi-predictive motion vector for the first part of the first partition.

[0407] As another example, FIG. 69 is a conceptual diagram showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to pictures in a single reference picture list. A bi-predictive motion vector points to two different reference picture lists. In this way, simply combining the first motion vector and the second motion vector does not generate a bi-predictive motion vector. In some embodiments, when the first motion vector and the second motion vector are uni-predictive motion vectors pointing to different reference pictures in a single reference picture list, if one of the reference pictures pointed to is, for example, included in another reference picture list, the motion vector corresponding to that picture (for example, the second motion vector if the reference picture pointed to by the second motion vector is included in both reference picture lists) may be represented by a motion vector pointing to the same reference picture in the other reference picture list. The first motion vector and the second motion vector are combined by replacing the motion vector corresponding to the reference picture with a replacement motion vector pointing to a reference picture in the other reference picture list, and stored as a bi-predictive motion vector for the first part of the first partition. As shown in FIG. 69, the current picture is POC4. The first motion vector Mv1 is a uni-predictive motion vector pointing to reference picture POC0 of reference picture list L0, which includes reference pictures POC0-POC8. The second motion vector Mv2 is a uni-predictive motion vector pointing to reference picture POC8 of reference picture list L0. Reference picture POC8 is also included in reference picture list L1. The second motion vector Mv2 is replaced with a replacement motion vector Mv2' pointing to reference picture POC8 in reference picture list L1. The first motion vector Mv1 and the replacement second motion vector Mv2' are combined and stored as a bi-predictive motion vector for the first part of the first partition. That is, different lists including the same pictures may be used.

[0408] As another example, FIG. 70, FIG. 71, and FIG. 72 are conceptual diagrams showing an example in which the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same reference picture in the same reference picture list. In some embodiments, when the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same reference picture in the same reference picture list, the first motion vector is stored as a uni-predictive motion vector for the first part of the first partition. As shown in FIG. 70, the current picture is POC4. The first motion vector Mv1 is a uni-predictive motion vector pointing to reference picture POC0 of the reference picture list L0 including reference pictures POC0 to POC8. The second motion vector Mv2 is a uni-predictive motion vector pointing to reference picture POC0 of the reference picture list L0. The first motion vector Mv1 is stored as a uni-predictive motion vector for the first part of the first partition.

[0409] In some embodiments, if the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same reference picture in the same reference picture list, the second motion vector is stored as the uni-predictive motion vector for the first part of the first partition. As shown in Figure 71, the current picture is POC4. The first motion vector Mv1 is a uni-predictive motion vector pointing to reference picture POC0 of reference picture list L0 including reference pictures POC0-POC8. The second motion vector Mv2 is a uni-predictive motion vector pointing to reference picture POC0 of reference picture list L0. The second motion vector Mv2 is stored as the uni-predictive motion vector for the first part of the first partition.

[0410] In some embodiments, if the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same reference picture in the same reference picture list, the second motion vector is replaced with an inverted motion vector pointing to a reference picture in another reference picture list, and then the first motion vector and the inverted motion vector are combined and stored as a bi-predictive motion vector for the first part of the first partition. As shown in FIG. 72, the current picture is POC4. The first motion vector Mv1 is a uni-predictive motion vector pointing to reference picture POC0 of the reference picture list L0 that includes reference pictures POC0-POC8. The second motion vector Mv2 is a uni-predictive motion vector pointing to reference picture POC0 of the reference picture list L0. The second motion vector Mv2 is replaced with an inverted motion vector MV2' pointing to a reference picture in the reference picture list L1, such as reference picture POC8 of the reference picture list L1 shown in the figure. The first motion vector Mv1 and the reverse motion vector Mv2' are combined and stored as a bi-predictive motion vector for the first part of the first partition. Such a reference picture of another reference picture list may be selected in various ways. For example, a reference picture in the reference picture list L1 that is at the same position as the reference picture pointed to by the second motion vector in the reference picture list L0 may be selected, or a reference picture in the reference picture list L1 that is at the same distance in display order from the current picture as the reference picture pointed to by the second motion vector in the reference picture list L0 may be selected. The reverse motion vector may be appropriately scaled according to the reference picture of the other reference picture list pointed to by the reverse motion vector.

[0411] In some embodiments, if a bi-predictive motion vector based on the first and second motion vectors can be stored without inverting or scaling one of the first and second motion vectors, a bi-predictive motion vector is stored for a first portion of the first partition (see, for example, the embodiments of Figures 68 and 69), and if a bi-predictive motion vector cannot be generated without inverting or scaling one of the first and second motion vectors, a uni-predictive motion vector is stored for the first portion of the first partition (see, for example, Figures 70 and 71, where one of the first and second motion vectors is stored as a uni-predictive motion vector for the first portion of the first partition).

[0412] As another example, if the first motion vector and the second motion vector are uni-predictive motion vectors pointing to the same reference picture, the first motion vector and the second motion vector may be averaged to generate a uni-predictive motion vector that is stored as the uni-predictive motion vector for the first portion of the first partition.

[0413] As another example, one of the first motion vector and the second motion vector may be a bi-predictive motion vector and the other may be a uni-predictive motion vector. The bi-predictive motion vector may be stored as the motion vector for the first portion of the first partition.

[0414] As another example, both the first motion vector and the second motion vector may be bi-predictive motion vectors, and the reference picture for the first motion vector may be the same as the reference picture for the second motion vector. The first motion vector and the second motion vector may be averaged to generate a bi-predictive motion vector that is stored as the bi-predictive motion vector for the first portion of the first partition.

[0415] As another example, both the first motion vector and the second motion vector may be bi-predictive motion vectors, and the reference picture of the first motion vector may be different from the reference picture of the second motion vector. The second motion vector may be scaled to the reference picture of the first motion vector and then averaged with the first motion vector to generate an average bi-predictive motion vector, which may be stored as the bi-predictive motion vector for the first portion of the first partition.

[0416] As another example, the first motion vector may be stored as the motion vector of the first part of the first partition, for example, always or if a certain condition exists. As another example, the second motion vector may be stored as the motion vector of the first part of the first partition, for example, always or if a certain condition exists. In step S6006, the first partition is encoded or decoded using at least the weighted samples of the first part of the first partition. For example, weighted pixel component values ​​(e.g., luma values, chroma values) of samples of the first sample set may be combined with weighted pixel component values ​​of corresponding samples of the second sample set to determine pixel component values ​​used in the encoding or decoding process. See, for example, Figures 47 to 61 and their corresponding descriptions.

[0417] The disclosed embodiment introduces a method for storing motion vectors of a weighted region (e.g., the first part of the first partition in the disclosed example). This may improve the accuracy of motion vector prediction for the next partition to be processed, which may improve processing efficiency. Note that the motion vectors of the weighted region may be stored for blocks of 4 pixels in general. For example, each sample may be a 4x4 pixel block in which the motion vectors are stored. The uni-predictive or bi-predictive motion vectors for the weighted region may be stored without being noticeable, i.e., without increasing the amount of memory used to store the motion vectors for the weighted region in any case.

[0418] It should be noted that the terms "encoding" and "processing" used in the description of the encoding method and encoding process performed by the image encoding device are interchangeable with the term "decoding" used in the description of the decoding method and decoding process performed by the image decoding device. Not all of the processes and components described are necessarily required, and in some embodiments, one or more of the described processes may be omitted.

[0419] One or more aspects disclosed herein may be implemented in combination with at least a part of other aspects in the present disclosure. Also, some processes shown in the flowcharts of one or more aspects disclosed herein, some configurations of devices, some syntax, etc. may be implemented in combination with other aspects.

[0420] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can usually be realized by an MPU (micro processing unit) and a memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memories. It is also possible to realize each functional block by hardware (dedicated circuitry). Various combinations of hardware and software may be adopted.

[0421] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Also, the processor that executes the above program may be single or multiple. That is, centralized processing or distributed processing may be performed.

[0422] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, which are also included within the scope of the aspects of the present disclosure.

[0423] Further, here, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems implementing the application examples will be described. Such a system may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device including both. Other configurations of such a system can be appropriately changed depending on the case.

[0424] [Usage example] 73 is a diagram showing the overall configuration of an appropriate content supply system ex100 for realizing a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0425] In this content supply system ex100, devices such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104, and base stations ex106 to ex110. The content supply system ex100 may be configured to connect any of the above devices in combination. In various implementations, the devices may be directly or indirectly connected to each other via a telephone network or short-distance wireless communication, etc., without the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game machine ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. Furthermore, the streaming server ex103 may be connected to a terminal in a hotspot in an airplane ex117, etc., via a satellite ex116.

[0426] Instead of the base stations ex106 to ex110, wireless access points or hot spots may be used. The streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to an airplane ex117 without going through a satellite ex116.

[0427] The camera ex113 is a device capable of taking still images and videos, such as a digital camera. The smartphone ex115 is a smartphone, a mobile phone, or a PHS (Personal Handy-phone System) that supports the mobile communication system, such as 2G, 3G, 3.9G, 4G, and 5G in the future.

[0428] The home appliance ex114 is a refrigerator, or an appliance included in a home fuel cell cogeneration system.

[0429] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live distribution or the like. In live distribution, a terminal (such as a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal in an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content photographed by a user using the terminal, may multiplex the video data obtained by encoding with sound data obtained by encoding sound corresponding to the video, and may transmit the obtained data to the streaming server ex103. That is, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0430] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game machine ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal in an airplane ex117, or the like, capable of decoding the encoded data. Each device that receives the distributed data may decode and play the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0431] [Distributed processing] The streaming server ex103 may be a plurality of servers or computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be realized by a CDN (Contents Delivery Network), and content distribution may be realized by a network that connects a large number of edge servers distributed around the world. In the CDN, a physically close edge server may be dynamically assigned according to the client. The content is cached and distributed to the edge server, thereby reducing delays. In addition, when some types of errors occur or communication conditions change due to an increase in traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the part of the network where a failure has occurred, thereby realizing high-speed and stable distribution.

[0432] In addition to the distributed processing of the distribution itself, the encoding processing of the captured data may be performed by each terminal, may be performed by the server side, or may be shared among the terminals. As an example, in the encoding processing, a processing loop is generally performed twice. In the first loop, the complexity of the image or the amount of code is detected for each frame or scene. In the second loop, processing is performed to maintain the image quality and improve the encoding efficiency. For example, the terminal performs the first encoding processing, and the server side that receives the content performs the second encoding processing, thereby improving the quality and efficiency of the content while reducing the processing load on each terminal. In this case, if there is a request to receive and decode almost in real time, the data encoded once by the terminal can be received and played back by other terminals, making it possible to perform more flexible real-time distribution.

[0433] As another example, the camera ex113 etc. extracts features (amount of features or characteristics) from an image, compresses data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning of the image (or the importance of the content), for example by determining the importance of an object from the features and switching the quantization precision. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server recompresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a high processing load such as CABAC (context-adaptive binary arithmetic coding).

[0434] As another example, in a stadium, a shopping mall, a factory, etc., there may be a plurality of video data in which almost the same scene has been shot by a plurality of terminals. In this case, using the plurality of terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, coding processing is assigned to each of them, for example, in units of GOPs (Group of Pictures), in units of pictures, or in units of tiles obtained by dividing a picture, for distributed processing. This reduces delays and realizes better real-time performance.

[0435] Since the multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referenced. The server may also receive the encoded data from each terminal and change the reference relationship between the multiple data, or correct or replace the pictures themselves and re-encode them. This makes it possible to generate a stream with improved quality and efficiency for each piece of data.

[0436] Furthermore, the server may distribute the video data after performing transcoding to change the encoding method of the video data. For example, the server may convert the MPEG encoding method to the VP encoding method (e.g., VP9), or convert H.264 to H.265.

[0437] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, in the following, descriptions such as "server" or "terminal" are used to indicate the entity performing the processing, but some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0438] [3D, multi-angle] It is becoming increasingly common to integrate and use images or videos of different scenes or the same scene taken from different angles by multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other. The videos taken by each device can be integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points included in the videos match.

[0439] The server may not only encode 2D video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. If the server can obtain the relative positional relationship between the shooting terminals, the server may generate a 3D shape of the scene based on not only 2D video but also images of the same scene captured from different angles. The server may separately encode 3D data generated by point cloud or the like, or may generate images to be transmitted to the receiving terminal by selecting or reconstructing images from images captured by multiple terminals based on the results of recognizing or tracking people or objects using the 3D data.

[0440] In this way, the user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting terminal, or can enjoy content in which a video of a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0441] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and the left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without mutual reference. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0442] In the case of an AR image, the server may superimpose virtual object information in the virtual space on camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may obtain or hold virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect the two-dimensional image to create superimposed data. Alternatively, the decoding device may transmit the movement of the user's viewpoint to the server in addition to the request for virtual object information. The server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data held by the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data typically has an α value indicating the transparency in addition to RGB, and the server may set the α value of the part other than the object created from the three-dimensional data to 0 or the like, and encode the part in a state where the part is transparent. Alternatively, the server may generate data in which a predetermined value of RGB value is set to the background like a chromakey, and the part other than the object is the background color. The predetermined value of RGB value may be determined in advance.

[0443] Similarly, the decoding process of the distributed data may be performed by the client (e.g., a terminal), by the server, or may be shared between the client and the server. As an example, a certain terminal may once send a reception request to the server, and the content corresponding to the request may be received by another terminal, decoded, and the decoded signal may be transmitted to a device having a display. By distributing the processing and selecting the appropriate content regardless of the performance of the communication-enabled terminal itself, data with good image quality can be reproduced. As another example, while large-sized image data is received by a TV or the like, a part of the area, such as tiles into which the picture is divided, may be decoded and displayed on the viewer's personal terminal. This allows the viewer to check his / her own area of ​​responsibility or the area he / she wants to check in more detail while sharing the overall picture.

[0444] It may be possible to seamlessly receive content using delivery system standards such as MPEG-DASH in situations where multiple short-range, medium-range, or long-range wireless communication is available indoors and outdoors. A user may freely select and switch in real time between a decoding device or a display device, such as a user's terminal or a display device placed indoors or outdoors. In addition, decoding can be performed while switching between a decoding device and a display device using the user's location information, etc. This makes it possible to map and display information on a part of the wall or ground of a neighboring building where a displayable device is embedded while the user is moving to a destination. It is also possible to switch the bit rate of the received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be accessed from the receiving terminal in a short time, or copied to an edge server in a content delivery service.

[0445] [Scalable Coding] Regarding the switching of contents, a scalable stream compressed and coded by applying the video coding method shown in each of the above embodiments, as shown in FIG. 74, will be used for explanation. The server may have multiple streams with the same content but different qualities as individual streams, but may be configured to switch contents by taking advantage of the characteristics of a temporal / spatial scalable stream realized by coding in layers as shown in the figure. In other words, the decoding side can freely switch and decode low-resolution content and high-resolution content by determining which layer to decode according to an internal factor such as performance and an external factor such as the state of the communication band. For example, if a user wants to continue watching a video that he or she was watching on the smartphone ex115 while on the move on a device such as an Internet TV after returning home, the device only needs to decode the same stream up to a different layer, thereby reducing the burden on the server side.

[0446] Furthermore, as described above, pictures are coded for each layer, and in addition to the configuration in which scalability is realized in an enhancement layer above the base layer, the enhancement layer may include meta-information based on image statistics and the like. The decoding side may generate high-quality content by super-resolving pictures in the base layer based on the meta-information. The super-resolution may improve the signal-to-noise ratio while maintaining and / or increasing the resolution. The meta-information includes information for specifying linear or non-linear filter coefficients to be used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least squares calculation to be used in the super-resolution process.

[0447] Alternatively, a configuration may be provided in which a picture is divided into tiles or the like according to the meaning of an object or the like in an image. The decoding side selects tiles to be decoded to decode only a part of the area. Furthermore, by storing the attributes of the object (person, car, ball, etc.) and the position in the video (coordinate position in the same image, etc.) as meta information, the decoding side can identify the position of a desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 75, the meta information may be stored using a data storage structure different from pixel data, such as a supplemental enhancement information (SEI) message in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0448] Meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. The decoding side can obtain the time when a specific person appears in a video, and by combining the picture-by-picture information with the time information, it can identify the picture in which an object exists and determine the position of the object within the picture.

[0449] [Web page optimization] FIG. 76 is a diagram showing an example of a display screen of a web page in a computer ex111 or the like. FIG. 77 is a diagram showing an example of a display screen of a web page in a smartphone ex115 or the like. As shown in FIG. 76 and FIG. 77, a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture that each content has as a link image, or may display an image such as a gif animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the image until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image enters the screen.

[0450] When a link image is selected by a user, the display device performs decoding, for example, giving top priority to the base layer. If the HTML constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, in order to ensure real-time performance, before selection or when the communication band is very tight, the display device decodes and displays only forward-reference pictures (I pictures, P pictures, and B pictures with forward reference only), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of decoding the content to the start of display). Furthermore, the display device may intentionally ignore the reference relationship of pictures, roughly decode all B pictures and P pictures with forward reference, and perform normal decoding as the number of received pictures increases over time.

[0451] [Automatic driving] Furthermore, when transmitting and receiving still image or video data such as 2D or 3D map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0452] In this case, since a car, drone, or airplane including a receiving terminal moves, the receiving terminal can realize seamless reception and decoding while switching between base stations ex106 to ex110 by transmitting location information of the receiving terminal. Also, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information according to a user's selection, a user's situation, and / or a communication band state.

[0453] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0454] [Distribution of personal content] Furthermore, the content supply system ex100 allows not only high-quality, long-duration content from video distributors, but also low-quality, short-duration content from individuals via unicast or multicast distribution. Such personal content is expected to continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, by using the following configuration.

[0455] During shooting, in real time or after accumulating, the server performs recognition processing such as shooting errors, scene search, semantic analysis, and object detection from the original image data or the encoded data. Then, based on the recognition results, the server manually or automatically performs editing such as correcting focus deviation or camera shake, deleting less important scenes such as scenes that are less bright than other pictures or out of focus, emphasizing object edges, and changing color. The server encodes the edited data based on the editing results. It is also known that if the shooting time is too long, the viewer rating will decrease, and the server may automatically clip not only scenes with less importance as described above but also scenes with little movement based on the image processing results so that the content will be within a specific time range depending on the shooting time. Alternatively, the server may generate a digest based on the result of the semantic analysis of the scene and encode it.

[0456] In some cases, personal content may contain content that infringes copyright, moral rights, or portrait rights, and the range of sharing may exceed the intended range, which may be inconvenient for individuals. Therefore, for example, the server may change the image to an unfocused image of a person's face on the periphery of the screen, or the inside of a house, and encode it. Furthermore, the server may recognize whether the image to be encoded contains a face of a person other than a person registered in advance, and if so, may perform processing such as applying a mosaic to the face. Alternatively, as pre-processing or post-processing of encoding, the user may specify a person or background area that he or she wishes to process in the image from the viewpoint of copyright, etc. The server may replace the specified area with another image, or may perform processing such as blurring the focus. If it is a person, the person can be tracked in the video and the image of the person's face can be replaced.

[0457] Since viewing of personal content with a small amount of data requires real-time performance, the decoding device may receive the base layer as a top priority and decode and play it, depending on the bandwidth. The decoding device may receive an enhancement layer during this time, and when the content is played back two or more times, such as when the playback is looped, play back high-quality video including the enhancement layer. With a stream that has been scalably encoded in this way, it is possible to provide an experience in which the video is rough when not selected or when viewing begins, but the stream gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can be provided even if a rough stream that is played the first time and a second stream that is encoded with reference to the first video are configured as a single stream.

[0458] [Other application examples] Moreover, these encoding or decoding processes are generally processed in an LSIex500 possessed by each terminal. The LSI (large scale integration circuitry)ex500 (see FIG. 73) may be a one-chip or multiple-chip configuration. Note that software for encoding or decoding moving images may be incorporated into some kind of recording medium (such as a CD-ROM, a flexible disk, or a hard disk) that can be read by the computer ex111 or the like, and the encoding or decoding process may be performed using the software. Furthermore, if the smartphone ex115 has a camera, video data captured by the camera may be transmitted. The video data at this time may be data encoded and processed by the LSIex500 possessed by the smartphone ex115.

[0459] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether the terminal supports the encoding method of the content or has the ability to execute a specific service. If the terminal does not support the encoding method of the content or does not have the ability to execute a specific service, the terminal may download a codec or application software, and then acquire and play the content.

[0460] Furthermore, at least one of the video encoding device (image encoding device) or video decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is carried and transmitted over broadcasting radio waves using a satellite or the like, there is a difference in that it is more suitable for multicast compared to the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to the encoding process and decoding process.

[0461] [Hardware configuration] FIG. 78 is a diagram showing further details of the smartphone ex115 shown in FIG. 73. FIG. 79 is a diagram showing a configuration example of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of taking videos and still images, and a display unit ex458 for displaying the video captured by the camera unit ex465 and the decoded data of the video received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting audio or sound, an audio input unit ex456 such as a microphone for inputting audio, a memory unit ex467 capable of storing encoded data such as captured video or still images, recorded audio, received video or still images, and e-mail, or decoded data, and a slot unit ex464 which is an interface unit with a SIMex468 for identifying a user and authenticating access to various data including a network. In addition, an external memory may be used instead of the memory unit ex467.

[0462] A main control unit ex460 that can comprehensively control the display unit ex458 and operation unit ex466 etc. is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0463] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 into an operational state and supplies power to each unit from the battery pack.

[0464] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, a ROM, and a RAM. During a telephone call, a voice signal collected by a voice input unit ex456 is converted into a digital voice signal by a voice signal processing unit ex454, and then subjected to spectrum spreading processing by a modulation / demodulation unit ex452, and then subjected to digital-to-analog conversion processing and frequency conversion processing by a transmission / reception unit ex451, and the resulting signal is transmitted via an antenna ex450. In addition, the received data is amplified and subjected to frequency conversion processing and analog-to-digital conversion processing, and then subjected to spectrum inverse spreading processing by a modulation / demodulation unit ex452, and then converted into an analog voice signal by a voice signal processing unit ex454, and then output from a voice output unit ex457. During a data communication mode, text, still images, or video data can be sent under the control of the main control unit ex460 via an operation input control unit ex462 based on the operation of an operation unit ex466 or the like of the main unit. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and codes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 by the moving image coding method shown in each of the above embodiments, and sends the coded video data to the multiplexing / separation unit ex453. The audio signal processing unit ex454 codes the audio signal collected by the audio input unit ex456 while the video or still image is being captured by the camera unit ex465, and sends the coded audio data to the multiplexing / separation unit ex453. The multiplexing / separation unit ex453 multiplexes the coded video data and the coded audio data by a predetermined method, and performs modulation and conversion processing in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, and transmits the data via the antenna ex450. The predetermined method may be determined in advance.

[0465] In the case of receiving a video attached to an e-mail or a chat, or a video linked to a web page, in order to decode the multiplexed data received via the antenna ex450, the multiplexing / separation unit ex453 separates the multiplexed data into a bit stream of video data and a bit stream of audio data by separating the multiplexed data, and supplies the encoded video data to the video signal processing unit ex455 via the synchronization bus ex470, and supplies the encoded audio data to the audio signal processing unit ex454. The video signal processing unit ex455 decodes the video signal by a video decoding method corresponding to the video encoding method shown in each of the above embodiments, and the video or still image included in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and the audio is output from the audio output unit ex457. As real-time streaming becomes more and more popular, audio playback may not be socially appropriate depending on the user's situation. Therefore, as an initial setting, it is preferable to have a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0466] Although the smartphone ex115 has been described as an example here, other implementation formats are possible, such as a transmitting terminal having only an encoder and a receiving terminal having only a decoder, in addition to a transmitting / receiving terminal having both an encoder and a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed with video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed in the multiplexed data. Also, video data itself may be received or transmitted instead of the multiplexed data.

[0467] Although the main control unit ex460 including the CPU controls the encoding or decoding process, various terminals are often equipped with a GPU. Therefore, a configuration may be used in which a wide area is processed collectively by utilizing the performance of the GPU using a memory shared by the CPU and GPU, or a memory whose addresses are managed so that they can be used in common. This can shorten the encoding time, ensure real-time performance, and achieve low latency. In particular, it is efficient to perform the processing of motion search, deblocking filter, SAO (Sample Adaptive Offset), and transformation and quantization collectively in units such as pictures by the GPU, rather than by the CPU.

Claims

1. The circuit, a memory connected to the circuit; The circuit, in operation, predicting a first set of pixels for a first partition of the image block using a first motion vector, the first motion vector being a uni-predictive motion vector; predicting a second set of pixels for a second partition of the image block using a second motion vector, the second motion vector being a unipredictive motion vector; weighting the first set of pixels and the second set of pixels for a first portion of pixels where the first partition and the second partition overlap; storing the first motion vector and the second motion vector in the memory as bi-predictive motion vectors and as motion vectors of the first portion; encoding the first partition using at least a plurality of pixels of the weighted first portion; Image encoding device.

2. The circuit, a memory connected to the circuit; The circuit, in operation, predicting a first set of pixels for a first partition of the image block using a first motion vector, the first motion vector being a uni-predictive motion vector; predicting a second set of pixels for a second partition of the image block using a second motion vector, the second motion vector being a unipredictive motion vector; weighting the first set of pixels and the second set of pixels for a first portion of pixels where the first partition and the second partition overlap; storing the first motion vector and the second motion vector in the memory as bi-predictive motion vectors and as motion vectors of the first portion; Decoding the first partition using at least the weighted pixels of the first portion. Image decoding device.

3. The circuit, a memory connected to the circuit; The circuit, in operation, predicting a first set of pixels for a first partition of the image block using a first motion vector, the first motion vector being a uni-predictive motion vector; predicting a second set of pixels for a second partition of the image block using a second motion vector, the second motion vector being a unipredictive motion vector; weighting the first set of pixels and the second set of pixels for a first portion of pixels where the first partition and the second partition overlap; storing the first motion vector and the second motion vector in the memory as bi-predictive motion vectors and as motion vectors of the first portion; encoding the first partition using at least a plurality of pixels of the weighted first portion; generating a bitstream including information used to encode the image block using the first partition; Bitstream generator.

Citation Information

Patent Citations

  • Method and apparatus for video encoding and decoding of geometrically partitioned bidirectional predictive mode partitions

    JP2011501508A

  • Smoothing of overlapping regions resulting from geometric motion subdivision.

    JP2013520877A

  • Adaptive overlapping block motion compensation

    JP2015502094A