Symbolizing device, decoding device, symbolizing method, decoding method, and bit stream generating device

The encoding device optimizes video coding by selectively referencing a video parameter set in the sequence parameter set based on the presence of a multi-layer or single-layer structure, addressing challenges in encoding efficiency and image quality while reducing processing and circuit demands.

JP7693551B2Active Publication Date: 2025-06-17PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021545592
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-09-10
Publication Date
2025-06-17
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in improving encoding efficiency, image quality, reducing processing volume, and circuit scale, while also appropriately selecting elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks.

Method used

The proposed solution involves an encoding device that generates an encoded bitstream with a configuration that includes a sequence parameter set (SPS) that refers to a video parameter set (VPS) when a multi-layer structure is present, and an SPS that does not refer to a VPS when a single-layer structure is present, thereby optimizing encoding parameters and reducing unnecessary data transmission.

Benefits of technology

This approach enhances coding efficiency, improves image quality, reduces processing and circuit requirements, and optimizes the selection of encoding components and operations, leading to improved overall processing efficiency and encoding speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693551000004
    Figure 0007693551000004
  • Figure 0007693551000005
    Figure 0007693551000005
  • Figure 0007693551000006
    Figure 0007693551000006
Patent Text Reader

Abstract

An encoding device (100) comprises a circuit and memory connected to the circuit, wherein the circuit, in an operation, generates an encoded bitstream, generates the encoded bitstream to include a sequence parameter set that references a video parameter set when generating the encoded bitstream to include a multilayered structure, and generates the encoded bitstream to include a sequence parameter set that does not reference the video parameter set when generating the encoded bitstream to not include the multilayered structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to video coding, and particularly to systems, components, and methods in video encoding and decoding, etc.

Background Art

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, there is always a need to provide improvements and optimizations to video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure relates to further progress, improvements, and optimizations in video coding.

[0003] Non-Patent Document 1 relates to an example of a conventional standard regarding the above-described video coding technology.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Regarding the encoding method as described above, for improving encoding efficiency, improving image quality, reducing processing volume, reducing circuit scale, or appropriately selecting elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, etc., a new method is desired to be proposed.

[0006] The present disclosure provides a configuration or method that can contribute to, for example, one or more of improvement in coding efficiency, improvement in image quality, reduction in processing amount, reduction in circuit scale, improvement in processing speed, and appropriate selection of elements or operations. Note that the present disclosure may include a configuration or method that can contribute to benefits other than the above.

Means for Solving the Problems

[0007] For example, an encoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit and, in operation, the circuit CVS (Symbolic Video Sequence) generates Generate the SPS that includes a VPS identifier indicating the identifier of the VPS (Video Parameter Set) referenced by the SPS (Sequence Parameter Set). When the CVS has one or more layers, the VPS identifier is set to a value greater than 0, and the SPS having the VPS identifier with a value of zero indicates (i) that the SPS does not reference the VPS, and (ii) that the CVS includes only one layer, and none of the VPSs have an identifier of the VPS with a value of zero .

[0008] In video coding technology, new methods are desired for improving coding efficiency, improving image quality, reducing circuit scale, etc.

[0009] Each embodiment in the present disclosure, or each of some of its configurations or methods, enables, for example, at least any one of improvement in coding efficiency, improvement in image quality, reduction in encoding / decoding processing amount, reduction in circuit scale, or improvement in encoding / decoding processing speed. Alternatively, each embodiment in the present disclosure, or each of some of its configurations or methods, enables appropriate selection of components / operations such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, etc. in encoding and decoding. Note that the present disclosure also includes disclosure of configurations or methods that can provide benefits other than the above. For example, a configuration or method for improving coding efficiency while suppressing an increase in processing amount.

[0010] Further advantages and effects in one aspect of the present disclosure will be apparent from the specification and drawings. Such advantages and / or effects can be obtained by each of some embodiments and features described in the specification and drawings, but it is not necessarily required that all are provided in order to obtain one or more advantages and / or effects.

[0011] Note that these general or specific aspects may be implemented in a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

Advantages of the Invention

[0012] A configuration or method according to an aspect of the present disclosure can contribute to, for example, one or more of improvement in coding efficiency, improvement in image quality, reduction in processing amount, reduction in circuit scale, improvement in processing speed, and appropriate selection of elements or operations. Note that a configuration or method according to an aspect of the present disclosure may contribute to benefits other than those described above.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23A

Figure 23B

Figure 23C

Figure 23D

Figure 23E

Figure 23F

Figure 23G

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38A

Figure 38B

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46A

Figure 46B

Figure 47A

Figure 47B

Figure 47C

Figure 48A

Figure 48B

Figure 49A

Figure 49B

Figure 50

Figure 51

Figure 52A

Figure 52B

Figure 52C

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58A

Figure 58B

Figure 59

Figure 60

Figure 61

Figure 62

Figure 63

Figure 64

Figure 65

Figure 66A

Figure 66B

Figure 67

Figure 68

Figure 69

Figure 70

Figure 71

Figure 72

Figure 73

Figure 74

Figure 75

Figure 76

Figure 77

Figure 78

Figure 79

Figure 80A

Figure 80B

Figure 81

Figure 82

Figure 83

Figure 84

Figure 85

Figure 86

Figure 87

Figure 88

Figure 89

Figure 90

Figure 91

Figure 92

Figure 93

Figure 94

Figure 95

Figure 96

Figure 97

Figure 98

Figure 99

Figure 100

Figure 101

Figure 102

Figure 103

Figure 104

Figure 105

Figure 106

Figure 107

DETAILED DESCRIPTION OF THE INVENTION

[0014] [Introduction] There is scalable encoding for encoding moving images. In scalable encoding, an encoded bitstream is encoded so as to have a scalability function. An encoded bitstream having a scalability function has a stream structure including a multi-layer structure. On the other hand, an encoded bitstream not having a scalability function has a stream structure including a single-layer structure.

[0015] An encoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, the circuit generates an encoded bitstream. When generating the encoded bitstream including a multi-layer structure, the circuit generates the encoded bitstream including a sequence parameter set that refers to a video parameter set. When generating the encoded bitstream without including a multi-layer structure therein, the circuit generates the encoded bitstream including a sequence parameter set that does not refer to the video parameter set.

[0016] Accordingly, when the encoding device generates an encoded bitstream that does not include a multi-layer structure, since the video parameter set is not referred to in the sequence parameter set, there may be a case where it is not necessary to transmit the video parameter set. Therefore, since the encoding device may not transmit the video parameter set, it may not be necessary to include the video parameter set in the encoded bitstream, and there is a possibility of improving the processing efficiency along with the encoding efficiency.

[0017] Here, for example, when generating the encoded bitstream including a multi-layer structure therein, the circuit generates the encoded bitstream including the ID of the sequence parameter set in which the value of 1 is described as the sequence parameter set that refers to the video parameter set. When generating the encoded bitstream without including a multi-layer structure therein, the circuit generates the encoded bitstream including the ID of the sequence parameter set in which the value of 0 is described as the sequence parameter set that does not refer to the video parameter set.

[0018] In this way, the encoding device can describe whether to refer to the video parameter set in the sequence parameter set by describing a value of 1 or 0 in the ID of the sequence parameter set.

[0019] Further, for example, the video parameter set is a video parameter set referred to by the sequence parameter set and does not have a value of 0.

[0020] Thereby, the encoding device can suppress a contradiction such that although the sequence parameter set refers to the video parameter set, the referred video parameter set is an encoded bitstream that does not include a multi-layer structure. Thus, the encoding device can suppress causing confusion on the decoding side due to the content indicated by the video parameter set and the sequence parameter set being different.

[0021] Also, a decoding device according to an aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, the circuit decodes an encoded bitstream, and when decoding the encoded bitstream, checks whether a sequence parameter set refers to a video parameter set. When the sequence parameter set refers to the video parameter set, a multi-layer is included in the encoded bitstream, and when the sequence parameter set does not refer to the video parameter set, a multi-layer is not included in the encoded bitstream.

[0022] Thereby, the decoding device can check whether a multi-layer structure is included in the encoded bitstream to be decoded only by checking whether the sequence parameter set refers to the video parameter set. Thus, there is a possibility that the decoding device can improve the processing efficiency when decoding the encoded bitstream.

[0023] Here, for example, when decoding the encoded bitstream, if the sequence parameter set refers to the video parameter set, the circuit determines the value of a variable indicating a list of extended layers to be decoded, and if the sequence parameter set does not refer to the video parameter set, the circuit determines the value of the variable indicating the list to be 0.

[0024] Also, for example, when the sequence parameter set refers to a video parameter set, a value of 1 is described in the ID of the sequence parameter set, and when the sequence parameter set does not refer to a video parameter set, a value of 0 is described in the ID of the sequence parameter set.

[0025] Thereby, the decoding device can confirm whether or not the encoded bit stream to be decoded includes a multi-layer structure only by checking the ID of the sequence parameter set. Therefore, there is a possibility that the decoding device can improve the processing efficiency when decoding the encoded bit stream.

[0026] Also, for example, the video parameter set is a video parameter set referred to by the sequence parameter set and does not have a value of 0.

[0027] Thereby, there is a possibility that the decoding device can prevent confusion caused by the difference in the content indicated by the video parameter set and the sequence parameter set.

[0028] Also, for example, when decoding the encoded bit stream, when the sequence parameter set refers to a video parameter set, the use of reference pictures that cross layers is permitted, and when the sequence parameter set does not refer to a video parameter set, the use of reference pictures that cross layers is prohibited.

[0029] Thereby, when decoding an encoded bit stream that does not include a multi-layer structure, the decoding device does not need to use reference pictures that cross layers, so there is a possibility that the processing efficiency when decoding the encoded bit stream can be improved.

[0030] Further, for example, in the encoding method according to one aspect of the present disclosure, when generating an encoded bitstream and including a multi-layer structure in the encoded bitstream, the encoded bitstream is generated including a sequence parameter set that refers to a video parameter set. When generating the encoded bitstream without including a multi-layer structure in the encoded bitstream, the encoded bitstream is generated including a sequence parameter set that does not refer to the video parameter set.

[0031] Thereby, when the encoding method generates an encoded bitstream that does not include a multi-layer structure, since the video parameter set is not referred to in the sequence parameter set, there may be a case where it is not necessary to transmit the video parameter set. Therefore, since the encoding method may not transmit the video parameter set, it is not necessary to include the video parameter set in the encoded bitstream, and there is a possibility that the encoding efficiency and the processing efficiency can be improved.

[0032] Further, for example, in the decoding method according to one aspect of the present disclosure, when decoding an encoded bitstream, when decoding the encoded bitstream, it is confirmed whether the sequence parameter set refers to a video parameter set. When the sequence parameter set refers to the video parameter set, a multi-layer structure is included in the encoded bitstream. When the sequence parameter set does not refer to the video parameter set, a multi-layer structure is not included in the encoded bitstream.

[0033] Thereby, the decoding method can confirm whether a multi-layer structure is included in the encoded bitstream to be decoded only by confirming whether the sequence parameter set refers to the video parameter set. Therefore, there is a possibility that the processing efficiency when decoding the encoded bitstream can be improved.

[0034] Furthermore, these general or specific aspects may be implemented in a system, apparatus, method, integrated circuit, computer program, or non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, apparatus, method, integrated circuit, computer program, and recording medium.

[0035] [Definition of Terms] Each term may be defined as follows by way of example.

[0036] (1) Image A unit of data composed of a set of pixels, consisting of a picture or a block smaller than a picture, including still images as well as moving images.

[0037] (2) Picture A processing unit of an image composed of a set of pixels, which may also be called a frame or a field.

[0038] (3) Block A processing unit of a set including a specific number of pixels, regardless of the name as in the following examples. Also, regardless of the shape, for example, a rectangle consisting of M×N pixels, a square consisting of M×M pixels, as well as a triangle, a circle, and other shapes are included.

[0039] (Examples of Blocks) · Slice / Tile / Brick · CTU / Super Block / Basic Division Unit · VPDU / Hardware Processing Division Unit · CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transformation Block Unit (TU) / Unit · Sub-block

[0040] (4) Pixel / Sample The smallest unit point constituting an image, including not only pixels at integer positions but also pixels at fractional positions generated based on pixels at integer positions.

[0041] (5) Pixel value / Sample value It is a unique value that a pixel has, including not only luminance values, color difference values, RGB gradations, but also depth values or binary values of 0 and 1.

[0042] (6) Flag In addition to 1 bit, it also includes cases of multiple bits. For example, it may be a parameter or index of 2 bits or more. Also, it may be not only binary two - valued but also multi - valued using other number systems.

[0043] (7) Signal It is something encoded and symbolized to transmit information, including not only discrete digital signals but also analog signals that take continuous values.

[0044] (8) Stream / Bitstream It refers to a data sequence or flow of digital data. A stream / bitstream may be composed of a single stream or may be divided into multiple layers and composed of multiple streams. Also, in addition to the case of being transmitted serially through a single transmission path, it includes the case of being transmitted by packet communication through multiple transmission paths.

[0045] (9) Difference In the case of a scalar quantity, in addition to the simple difference (x - y), as long as the operation of difference is included, it is sufficient, including the absolute value of the difference (|x - y|), the square difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a and b are constants), and the offset difference (x - y + a: a is an offset).

[0046] (10) Sum In the case of a scalar quantity, in addition to the simple sum (x + y), as long as the operation of sum is included, it is sufficient, including the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted sum (ax + by: a and b are constants), and the offset sum (x + y + a: a is an offset).

[0047] (11) Based on This also includes cases where elements other than the target elements based on are taken into account. Also, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0048] (12) used, using This also includes cases where elements other than the elements to be used are taken into account. Also, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0049] (13) prohibit, forbid It can be rephrased as not permitted. Also, not prohibited or permitted does not necessarily mean an obligation.

[0050] (14) limit, restriction / restrict / restricted It can be rephrased as not permitted. Also, not prohibited or permitted does not necessarily mean an obligation. Furthermore, it suffices if a part is prohibited quantitatively or qualitatively, and it also includes cases where it is prohibited comprehensively.

[0051] (15) chroma Adjectives represented by the symbols Cb and Cr that specify that a sample array or a single sample represents one of two colour difference signals related to primary colours. Instead of the term chroma, the term chrominance can also be used.

[0052] (16) luma Adjectives represented by a symbol or subscript Y or L that specify that a sample array or a single sample represents a monochrome signal related to primary colours. Instead of the term luma, the term luminance can also be used.

[0053] [Explanation regarding the description] In the drawings, the same reference numerals denote the same or similar components. Also, the sizes and relative positions of the components in the drawings are not necessarily drawn to scale.

[0054] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that the embodiments described below are all examples showing comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0055] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device to which the processes and / or configurations described in each aspect of the present disclosure are applicable. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented, for example.

[0056] (1) Any one of the plurality of components of the encoding device or decoding device of the embodiment described in each aspect of the present disclosure may be replaced or combined with any other component described in any aspect of the present disclosure.

[0057] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. of functions or processes performed by some of the plurality of components of the encoding device or decoding device may be made. For example, any function or process may be replaced or combined with any other function or process described in any aspect of the present disclosure.

[0058] (3) In the method implemented by the encoding device or decoding device of the embodiment, for some of the plurality of processes included in the method, arbitrary changes such as addition, replacement, and deletion may be made. For example, any process in the method may be replaced or combined with other processes described in any of the aspects of the present disclosure.

[0059] (4) Some of the plurality of components constituting the encoding device or decoding device of the embodiment may be combined with the components described in any of the aspects of the present disclosure, or may be combined with components having a part of the functions described in any of the aspects of the present disclosure, or may be combined with components that implement a part of the processes implemented by the components described in each aspect of the present disclosure.

[0060] (5) A component having a part of the functions of the encoding device or decoding device of the embodiment, or a component that implements a part of the processes of the encoding device or decoding device of the embodiment may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having a part of the functions described in any of the aspects of the present disclosure, or a component that implements a part of the processes described in any of the aspects of the present disclosure.

[0061] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the plurality of processes included in the method may be replaced or combined with the processes described in any of the aspects of the present disclosure, or with any similar processes.

[0062] (7) Some of the plurality of processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0063] (8) The manner of implementing the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for purposes different from the moving image encoding or moving image decoding disclosed in the embodiments.

[0064] [System Configuration] FIG. 1 is a schematic diagram showing an example of the configuration of a transmission system according to the present embodiment.

[0065] The transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, an encoding device 100, a network Nw, and a decoding device 200 as shown in FIG. 1.

[0066] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding.

[0067] Note that the original image before being encoded, which is input to the encoding device 100, is also called the original image, original signal, or original sample. Also, the image may be a moving image or a still image. Further, the image is a higher concept such as a sequence, picture, and block, and is not subject to spatial and temporal region limitations unless otherwise specified. Also, the image consists of an array of pixels or pixel values, and the signal representing the image, or the pixel values, are also called samples. Also, the stream may be called a bit stream, encoded bit stream, compressed bit stream, or encoded signal. Furthermore, the encoding device may be called an image encoding device or a moving image encoding device, and the encoding method by the encoding device 100 may be called an encoding method, image encoding method, or moving image encoding method.

[0068] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Also, the network Nw may be replaced by a storage medium that records a stream such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0069] The decoding device 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method by the encoding device 100.

[0070] Note that the decoding device may be referred to as an image decoding device or a moving image decoding device, and the decoding method by the decoding device 200 may be referred to as a decoding method, an image decoding method, or a moving image decoding method.

[0071] [Data Structure] FIG. 2 is a diagram showing an example of the hierarchical structure of data in a stream. The stream includes, for example, a video sequence. This video sequence includes, for example, as shown in FIG. 2(a), a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and a plurality of pictures.

[0072] The VPS includes encoding parameters common to a plurality of layers in a moving image composed of a plurality of layers, and encoding parameters related to the plurality of layers or individual layers included in the moving image.

[0073] The SPS includes parameters used for a sequence, that is, the encoding parameters that the decoding device 200 refers to for decoding the sequence. For example, the encoding parameters may indicate the width or height of a picture. Note that there may be multiple SPSs.

[0074] The PPS includes parameters used for a picture, that is, the encoding parameters that the decoding device 200 refers to for decoding each picture in the sequence. For example, the encoding parameters may include a reference value of the quantization width used for decoding the picture and a flag indicating the application of weighted prediction. Note that there may be multiple PPSs. Also, the SPS and the PPS may simply be called a parameter set in some cases.

[0075] As shown in FIG. 2(b), a picture may include a picture header and one or more slices. The picture header includes the encoding parameters that the decoding device 200 refers to for decoding the one or more slices.

[0076] As shown in FIG. 2(c), a slice includes a slice header and one or more blocks. The slice header includes the encoding parameters that the decoding device 200 refers to for decoding the one or more blocks.

[0077] As shown in FIG. 2(d), a block includes one or more CTUs (Coding Tree Units).

[0078] Note that a picture may not include slices and instead may include tile groups. In this case, a tile group includes one or more tiles. Also, a slice may be included in a block.

[0079] The CTU is also called a superblock or a basic division unit. Such a CTU includes a CTU header and one or more CUs (Coding Units), as shown in Fig. 2(e). The CTU header includes coding parameters that a decoding device 200 refers to for decoding one or more CUs.

[0080] A CU may be divided into a plurality of smaller CUs. Also, as shown in Fig. 2(f), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual described later. Note that a CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but in, for example, the SBT described later, it may include a plurality of TUs smaller than the CU. Also, a CU may be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is, for example, a fixed unit that can be processed in one stage when performing pipeline processing in hardware.

[0081] Note that the stream may not have some of the layers shown in Fig. 2. Also, the order of these layers may be swapped, and any layer may be replaced by another layer. Also, a picture that is the target of the processing currently performed by a device such as the encoding device 100 or the decoding device 200 is called a current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded, and if the processing is decoding, the current picture is synonymous with the picture to be decoded. Also, a block such as a CU or a CU that is the target of the processing currently performed by a device such as the encoding device 100 or the decoding device 200 is called a current block. If the processing is encoding, the current block is synonymous with the block to be encoded, and if the processing is decoding, the current block is synonymous with the block to be decoded.

[0082] [Configuration of Picture, Slice / Tile] To decode pictures in parallel, a picture may be composed of slices or tiles.

[0083] A slice is a basic encoding unit that makes up a picture. A picture is composed of, for example, one or more slices. Also, a slice consists of one or more consecutive CTUs.

[0084] Figure 3 is a diagram showing an example of the composition of a slice. For example, a picture contains 11×8 CTUs and is divided into 4 slices (Slice 1 - 4). Slice 1 consists of, for example, 16 CTUs, Slice 2 consists of, for example, 21 CTUs, Slice 3 consists of, for example, 29 CTUs, and Slice 4 consists of, for example, 22 CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of a slice is in the form of dividing the picture horizontally. The boundary of a slice does not have to be at the edge of the screen and can be anywhere among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs within a slice is, for example, the raster scan order. Also, a slice includes a slice header and encoded data. The slice header may describe the characteristics of that slice, such as the CTU address at the start of the slice and the slice type.

[0085] A tile is a unit of a rectangular area that makes up a picture. A number called TileId may be assigned to each tile in the raster scan order.

[0086] FIG. 4 is a diagram showing an example of the configuration of tiles. For example, a picture includes 11×8 CTUs and is divided into tiles (tiles 1-4) in four rectangular regions. When tiles are used, the processing order of CTUs is changed compared to when tiles are not used. When tiles are not used, a plurality of CTUs in a picture are processed, for example, in raster scan order. When tiles are used, in each of the plurality of tiles, at least one CTU is processed, for example, in raster scan order. For example, as shown in FIG. 4, the processing order of the plurality of CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0087] Note that one tile may include one or more slices, and one slice may include one or more tiles.

[0088] Note that a picture may be configured in tile set units. A tile set may include one or more tile groups and may include one or more tiles. A picture may be configured by only any one of a tile set, a tile group, and a tile. For example, the order of scanning a plurality of tiles in raster order for each tile set is defined as the basic encoding order of the tiles. A collection of one or more tiles with continuous basic encoding order within each tile set is defined as a tile group. Such a picture may be configured by a division unit 102 (see FIG. 7) described later.

[0089] [Scalable Encoding] FIGS. 5 and 6 are diagrams showing an example of the configuration of a scalable stream.

[0090] As shown in FIG. 5, the symbolization device 100 may generate a temporally / spatially scalable stream by dividing each of a plurality of pictures into one of a plurality of layers and symbolizing them. For example, the symbolization device 100 realizes scalability in which an enhancement layer exists above a base layer by symbolizing pictures for each layer. Symbolizing each such picture is called scalable symbolization. Thereby, the decoding device 200 can switch the image quality of the image displayed by decoding the stream. That is, the decoding device 200 determines up to which layer to decode according to internal factors such as its own performance and external factors such as the state of the communication band. As a result, the decoding device 200 can freely switch and decode the same content into low-resolution content and high-resolution content. For example, a user of the stream watches the moving image of the stream up to a certain point using a smartphone while moving, and watches the continuation of the moving image using a device such as an Internet TV after returning home. Note that a decoding device 200 with the same or different performance is incorporated in each of the above-described smartphone and device. In this case, if the device decodes up to the upper layer of the stream, the user can watch a high-quality moving image after returning home. Thereby, the symbolization device 100 does not need to generate a plurality of streams with the same content but different image qualities, and can reduce the processing load.

[0091] Furthermore, the enhancement layer may include meta information based on statistical information of the image or the like. The decoding device 200 may generate a moving image with improved image quality by super-resolving the picture of the base layer based on the meta information. Super-resolution may be either an improvement in the signal-to-noise ratio at the same resolution or an enlargement of the resolution. The meta information may include information for specifying linear or non-linear filter coefficients used for super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least-squares operation used for super-resolution processing.

[0092] Alternatively, according to the meaning of each object in the picture, etc., the picture may be divided into tiles or the like. In this case, the decoding device 200 may decode only a partial area of the picture by selecting the tile to be decoded. Also, the attributes of the object (such as a person, a car, a ball, etc.) and the position in the picture (such as the coordinate position in the same picture) may be stored as meta information. In this case, the decoding device 200 can identify the position of the desired object based on the meta information and determine the tile including the object. For example, as shown in FIG. 6, the meta information is stored using a data storage structure different from the pixel data, such as SEI in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0093] Also, the meta information may be stored in units composed of a plurality of pictures, such as a stream, a sequence, or a random access unit. Thereby, the decoding device 200 can obtain the time when a specific person appears in the moving image, etc., and by using the time and the information in picture units, can identify the picture in which the object exists and the position of the object in that picture.

[0094] [Encoding device] Next, the encoding device 100 according to the embodiment will be described. FIG. 7 is a block diagram showing an example of the functional configuration of the encoding device 100 according to the embodiment. The encoding device 100 encodes an image in block units.

[0095] As shown in FIG. 7, the encoding device 100 is a device that encodes an image in block units, and includes a division unit 102, a subtraction unit 104, a conversion unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse conversion unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. Note that each of the intra prediction unit 124 and the inter prediction unit 126 is configured as a part of the prediction processing unit.

[0096] [Implementation Example of Encoding Device] FIG. 8 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, a plurality of components of the encoding device 100 shown in FIG. 7 are implemented by the processor a1 and the memory a2 shown in FIG. 8.

[0097] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes an image. The processor a1 may be a processor such as a CPU. Also, the processor a1 may be an aggregate of a plurality of electronic circuits. Also, for example, the processor a1 may play the roles of a plurality of components of the encoding device 100 shown in FIG. 7, excluding the components for storing information.

[0098] The memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode an image is stored. The memory a2 may be an electronic circuit and may be connected to the processor a1. Also, the memory a2 may be included in the processor a1. Also, the memory a2 may be an aggregate of a plurality of electronic circuits. Also, the memory a2 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or a recording medium, etc. Also, the memory a2 may be a non-volatile memory or a volatile memory.

[0099] For example, the memory a2 may store the image to be encoded or the stream corresponding to the encoded image. Also, the memory a2 may store a program for the processor a1 to encode an image.

[0100] Also, for example, the memory a2 may serve as a component for storing information among the plurality of components of the encoding device 100 shown in FIG. 7. Specifically, the memory a2 may serve as the block memory 118 and the frame memory 122 shown in FIG. 7. More specifically, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.) may be stored in the memory a2.

[0101] Note that in the encoding device 100, not all of the plurality of components shown in FIG. 7 need to be implemented, and not all of the plurality of processes described above need to be performed. A part of the plurality of components shown in FIG. 7 may be included in another device, and a part of the plurality of processes described above may be executed by another device.

[0102] Hereinafter, after explaining the overall processing flow of the encoding device 100, each component included in the encoding device 100 will be described.

[0103] [Overall Flow of Encoding Process] FIG. 9 is a flowchart showing an example of the overall encoding process by the encoding device 100.

[0104] First, the splitting unit 102 of the encoding device 100 splits the picture included in the original image into a plurality of blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the splitting unit 102 selects a splitting pattern for the block of the fixed size (step Sa_2). That is, the splitting unit 102 further splits the block of the fixed size into a plurality of blocks constituting the selected splitting pattern. Then, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 for each of the plurality of blocks.

[0105] The prediction processing unit including the intra prediction unit 124 and the inter prediction unit 126 and the prediction control unit 128 generate a predicted image of the current block (step Sa_3). Note that the predicted image is also called a prediction signal, a predicted block, or a prediction sample.

[0106] Next, the subtraction unit 104 generates the difference between the current block and the predicted image as the prediction residual (step Sa_4). Note that the prediction residual is also called the prediction error.

[0107] Next, the conversion unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing conversion and quantization on the predicted image (step Sa_5).

[0108] Next, the entropy encoding unit 110 generates a stream by performing encoding (specifically, entropy encoding) on the plurality of quantization coefficients and the prediction parameters related to the generation of the predicted image (step Sa_6).

[0109] Next, the inverse quantization unit 112 and the inverse conversion unit 114 restore the prediction residual by performing inverse quantization and inverse conversion on the plurality of quantization coefficients (step Sa_7).

[0110] Next, the addition unit 116 reconstructs the current block by adding the predicted image to the restored prediction residual (step Sa_8). Thereby, a reconstructed image is generated. Note that the reconstructed image is also called a reconstructed block, and in particular, the reconstructed image generated by the encoding device 100 is also called a local decoding block or a local decoded image.

[0111] When this reconstructed image is generated, the loop filter unit 120 performs filtering on the reconstructed image as necessary (step Sa_9).

[0112] Then, the encoding device 100 determines whether the encoding of the entire picture is completed (step Sa_10). If it is determined that the encoding is not completed (No in step Sa_10), the processing from step Sa_2 is repeatedly executed.

[0113] In the above example, the encoding device 100 selects one splitting pattern for a block of fixed size and encodes each block according to the selected splitting pattern. However, each block may be encoded according to each of a plurality of splitting patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of splitting patterns, and for example, select the stream obtained by encoding according to the splitting pattern with the smallest cost as the finally output stream.

[0114] Also, the processes of steps Sa_1 to Sa_10 may be sequentially performed by the encoding device 100, a plurality of some of those processes may be performed in parallel, or the order may be changed.

[0115] The encoding process by such an encoding device 100 is a hybrid encoding using predictive encoding and transform encoding. Also, the predictive encoding is performed by an encoding loop including a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. That is, the prediction processing unit including the intra prediction unit 124 and the inter prediction unit 126 constitutes a part of the encoding loop.

[0116] [Splitting Unit] The splitting unit 102 splits each picture included in the original image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128x128 pixels). Such blocks of fixed size are sometimes referred to as Coding Tree Units (CTUs). Then, the splitting unit 102 splits each of the fixed-size blocks into blocks of variable size (e.g., 64x64 pixels or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects a splitting pattern. Such blocks of variable size are sometimes referred to as Coding Units (CUs), Prediction Units (PUs), or Transformation Units (TUs). Note that in various implementation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and some or all of the blocks in the picture may be processing units of CUs, PUs, or TUs.

[0117] FIG. 10 is a diagram showing an example of block splitting in the embodiment. In FIG. 10, the solid lines represent block boundaries by quadtree block splitting, and the dashed lines represent block boundaries by binary tree block splitting.

[0118] Here, the block 10 is a square block of 128x128 pixels. This block 10 is first split into four square blocks of 64x64 pixels (quadtree block splitting).

[0119] The upper-left square block of 64x64 pixels is further vertically split into two rectangular blocks each consisting of 32x64 pixels, and the left rectangular block of 32x64 pixels is further vertically split into two rectangular blocks each consisting of 16x64 pixels (binary tree block splitting). As a result, the upper-left square block of 64x64 pixels is split into two rectangular blocks 11 and 12 of 16x64 pixels and a rectangular block 13 of 32x64 pixels.

[0120] The upper right 64x64 pixel square block is horizontally divided into two rectangular blocks 14 and 15 each consisting of 64x32 pixels (binary tree block division).

[0121] The lower left 64x64 pixel square block is divided into four square blocks each consisting of 32x32 pixels (quad-tree block division). Among the four square blocks each consisting of 32x32 pixels, the upper left block and the lower right block are further divided. The upper left 32x32 pixel square block is vertically divided into two rectangular blocks each consisting of 16x32 pixels, and the right 16x32 pixel rectangular block is further horizontally divided into two square blocks each consisting of 16x16 pixels (binary tree block division). The lower right 32x32 pixel square block is horizontally divided into two rectangular blocks each consisting of 32x16 pixels (binary tree block division). As a result, the lower left 64x64 pixel square block is divided into a 16x32 pixel rectangular block 16, two 16x16 pixel square blocks 17 and 18, two 32x32 pixel square blocks 19 and 20, and two 32x16 pixel rectangular blocks 21 and 22.

[0122] The block 23 consisting of 64x64 pixels in the lower right is not divided.

[0123] As described above, in FIG. 10, the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.

[0124] Note that in FIG. 10, one block is divided into four or two blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, one block may be divided into three blocks (ternary tree block division). Such division including ternary tree block division is sometimes called MBT (multi type tree) division.

[0125] FIG. 11 is a diagram showing an example of the functional configuration of the dividing unit 102. As shown in FIG. 11, the dividing unit 102 may include a block division determination unit 102a. The block division determination unit 102a may perform the following processing as an example.

[0126] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the above-described division pattern based on the block information. The dividing unit 102 divides the original image according to the division pattern, and outputs one or more blocks obtained by the division to the subtraction unit 104.

[0127] Further, the block division determination unit 102a outputs, for example, parameters indicating the above-described division pattern to the conversion unit 106, the inverse conversion unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy encoding unit 110. The conversion unit 106 may convert the prediction residual based on the parameters, and the intra prediction unit 124 and the inter prediction unit 126 may generate a prediction image based on the parameters. Further, the entropy encoding unit 110 may perform entropy encoding on the parameters.

[0128] The parameters related to the division pattern may be written into the stream as follows as an example.

[0129] FIG. 12 is a diagram showing an example of the division pattern. The division pattern includes, for example, a four-way division (QT) in which a block is divided into two in each of the horizontal and vertical directions, a three-way division (HT or VT) in which a block is divided in the same direction at a ratio of 1:2:1, a two-way division (HB or VB) in which a block is divided in the same direction at a ratio of 1:1, and no division (NS).

[0130] Note that in the case of four-way division and no division, the division pattern does not have a block division direction, and in the case of two-way division and three-way division, the division pattern has division direction information.

[0131] Figures 13A and 13B are diagrams showing an example of a syntax tree of a split pattern. In the example of Figure 13A, first, information (S: Split flag) indicating whether to perform splitting exists. Next, information (QT: QT flag) indicating whether to perform four-way splitting exists. Next, information (TT: TT flag or BT: BT flag) indicating whether to perform three-way splitting or two-way splitting exists, and finally, information (Ver: Vertical flag or Hor: Horizontal flag) indicating the split direction exists. Note that for each of one or more blocks obtained by splitting according to such a split pattern, splitting may be repeatedly applied in the same manner. That is, as an example, the determination of whether to perform splitting, whether to perform four-way splitting, whether the splitting method is horizontal or vertical, and whether to perform three-way splitting or two-way splitting may be recursively performed, and the determined results may be encoded into a stream according to the encoding order disclosed in the syntax tree shown in Figure 13A.

[0132] Also, in the syntax tree shown in Figure 13A, the information is arranged in the order of S, QT, TT, Ver, but it may also be arranged in the order of S, QT, Ver, BT. That is, in the example of Figure 13B, first, information (S: Split flag) indicating whether to perform splitting exists. Next, information (QT: QT flag) indicating whether to perform four-way splitting exists. Next, information (Ver: Vertical flag or Hor: Horizontal flag) indicating the split direction exists, and finally, information (BT: BT flag or TT: TT flag) indicating whether to perform two-way splitting or three-way splitting exists.

[0133] Note that the split pattern described here is an example, and a pattern other than the described split pattern may be used, or only a part of the described split pattern may be used.

[0134] [Subtraction unit] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units divided by the division unit 102 and input from the division unit 102. That is, the subtraction unit 104 calculates the prediction residual of the current block. Then, the subtraction unit 104 outputs the calculated prediction residual to the conversion unit 106.

[0135] The original image is an input signal of the encoding device 100 and is, for example, a signal representing the image of each picture constituting a moving image (for example, a luminance signal and two color difference signals).

[0136] [Conversion Unit] The conversion unit 106 converts the prediction residual in the spatial domain into conversion coefficients in the frequency domain and outputs the conversion coefficients to the quantization unit 108. Specifically, the conversion unit 106 performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction residual in the spatial domain.

[0137] Note that the conversion unit 106 may adaptively select a conversion type from a plurality of conversion types and convert the prediction residual into conversion coefficients using a transform basis function corresponding to the selected conversion type. Such a conversion is sometimes referred to as an EMT (explicit multiple core transform) or an AMT (adaptive multiple transform). Also, the transform basis function is sometimes simply referred to as a basis.

[0138] The plurality of conversion types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these conversion types may be respectively denoted as DCT2, DCT5, DCT8, DST1, and DST7. FIG. 14 is a table showing the transform basis functions corresponding to each conversion type. In FIG. 14, N indicates the number of input pixels. The selection of the conversion type from among these plurality of conversion types may depend on, for example, the type of prediction (intra prediction and inter prediction, etc.) or the intra prediction mode.

[0139] Information indicating whether or not to apply such EMT or AMT (for example, what is called an EMT flag or an AMT flag) and information indicating the selected conversion type are usually signaled at the CU level. Note that the signaling of these pieces of information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, block level, or CTU level).

[0140] Also, the conversion unit 106 may re-convert the conversion coefficients (that is, the conversion results). Such re-conversion may be called AST (adaptive secondary transform) or NSST (non-separable secondary transform). For example, the conversion unit 106 performs re-conversion for each sub-block (for example, a sub-block of 4x4 pixels) included in the block of conversion coefficients corresponding to the intra prediction residual. Information indicating whether or not to apply NSST and information regarding the conversion matrix used for NSST are usually signaled at the CU level. Note that the signaling of these pieces of information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, block level, or CTU level).

[0141] A separable conversion and a non-separable conversion may be applied to the conversion unit 106. The separable conversion is a method in which conversion is performed a plurality of times by separating for each direction by the number of dimensions of the input, and the non-separable conversion is a method in which when the input is multi-dimensional, two or more dimensions are regarded as one dimension and conversion is performed together.

[0142] For example, as an example of the non-separable conversion, when the input is a block of 4×4 pixels, it is regarded as an array having 16 elements, and a conversion process is performed on the array with a 16×16 conversion matrix.

[0143] Also, in a further example of non-separable transformation, after regarding a 4×4 pixel input block as an array having 16 elements, a transformation (Hypercube Givens Transform) that performs a plurality of Givens rotations on the array may be performed.

[0144] In the transformation in the transformation unit 106, the type of transformation basis function for transformation into the frequency domain can also be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform).

[0145] FIG. 15 is a diagram showing an example of SVT.

[0146] In SVT, as shown in FIG. 15, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The type of transformation may be set for each region. For example, DST7 and DCT8 are used. For example, among the two regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used for the region at position 0. Alternatively, DST7 can be used for the region at position 1 among the two regions. Similarly, among the two regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 can be used for the region at position 0. Alternatively, DST7 can be used for the region at position 1 among the two regions. In the example shown in FIG. 15 like this, only one of the two regions within the CU is transformed and the other is not, but transformation may be performed for each of the two regions. Also, the splitting method may be not only bisecting but also quartering. Also, it can be made more flexible, such as encoding the information indicating the splitting method and signaling it in the same way as CU splitting. Note that SVT is sometimes also called SBT (Sub-block Transform).

[0147] The aforementioned AMT and EMT may be referred to as MTS (Multiple Transform Selection). When applying MTS, a transform type such as DST7 or DCT8 can be selected, and the information indicating the selected transform type may be encoded as index information for each CU. On the other hand, there is a process called IMTS (Implicit MTS) as a process of selecting a transform type to be used for orthogonal transformation without encoding index information based on the shape of the CU. When applying IMTS, for example, if the shape of the CU is rectangular, orthogonal transformation is performed using DST7 for the short side of the rectangle and DCT2 for the long side. Also, for example, if the shape of the CU is square, if MTS is effective within the sequence, DCT2 is used, and if MTS is ineffective, DST7 is used for orthogonal transformation. DCT2 and DST7 are just examples, and other transform types may be used, or different combinations of the transform types to be used are also possible. IMTS may be used only for blocks of intra prediction, or may be used for both blocks of intra prediction and blocks of inter prediction.

[0148] In the above, as a selection process for selectively switching the conversion type used in orthogonal conversion, three processes of MTS, SBT, and IMTS have been described. However, all three selection processes may be valid, or only some of the selection processes may be selectively made valid. Whether to make each individual selection process valid can be identified by flag information in the header such as SPS. For example, if all three selection processes are valid, one is selected from the three selection processes for orthogonal conversion in CU units. Note that as long as at least one of the following four functions [1] to [4] can be realized, a selection process different from the above three selection processes may be used for the selection process of selectively switching the conversion type, or each of the above three selection processes may be replaced with a different process. Function [1] is a function of performing orthogonal conversion on the entire range within a CU and encoding information indicating the conversion type used in the conversion. Function [2] is a function of performing orthogonal conversion on the entire range of a CU and determining the conversion type based on a predetermined rule without encoding the information indicating the conversion type. Function [3] is a function of performing orthogonal conversion on a partial region of a CU and encoding information indicating the conversion type used in the conversion. Function [4] is a function of performing orthogonal conversion on a partial region of a CU and determining the conversion type based on a predetermined rule without encoding the information indicating the conversion type used in the conversion.

[0149] Note that whether each of MTS, IMTS, and SBT is applied or not may be determined for each processing unit. For example, it may be determined whether to apply it in sequence units, picture units, block units, slice units, CTU units, or CU units.

[0150] Note that the tool for selectively switching the conversion type in the present disclosure may also be described as a method of adaptively selecting a basis used in the conversion process, a selection process, or a process of selecting a basis. Also, the tool for selectively switching the conversion type may be described as a mode of adaptively selecting the conversion type.

[0151] FIG. 16 is a flowchart showing an example of the processing by the conversion unit 106.

[0152] For example, the conversion unit 106 determines whether to perform an orthogonal transformation (step St_1). Here, if the conversion unit 106 determines to perform an orthogonal transformation (Yes in step St_1), it selects a conversion type to be used for the orthogonal transformation from a plurality of conversion types (step St_2). Next, the conversion unit 106 performs an orthogonal transformation by applying the selected conversion type to the prediction residual of the current block (step St_3). Then, the conversion unit 106 causes the entropy encoding unit 110 to encode the information indicating the selected conversion type by outputting the information to the entropy encoding unit 110 (step St_4). On the other hand, if the conversion unit 106 determines not to perform an orthogonal transformation (No in step St_1), it causes the entropy encoding unit 110 to encode the information indicating that no orthogonal transformation is performed by outputting the information to the entropy encoding unit 110 (step St_5). Note that the determination of whether to perform an orthogonal transformation in step St_1 may be made based on, for example, the size of the conversion block, the prediction mode applied to the CU, and the like. Also, the information indicating the conversion type to be used for the orthogonal transformation may not be encoded, and the orthogonal transformation may be performed using a predefined conversion type.

[0153] FIG. 17 is a flowchart showing another example of the processing by the conversion unit 106. Note that the example shown in FIG. 17 is an example of orthogonal transformation when the method of selectively switching the conversion type to be used for the orthogonal transformation is applied, similar to the example shown in FIG. 16.

[0154] As an example, the first conversion type group may include DCT2, DST7, and DCT8. As another example, the second conversion type group may include DCT2. Also, the conversion types included in the first conversion type group and the second conversion type group may partially overlap or may be all different conversion types.

[0155] Specifically, the conversion unit 106 determines whether the conversion size is less than or equal to a predetermined value (step Su_1). Here, if it is determined that the conversion size is less than or equal to the predetermined value (Yes in step Su_1), the conversion unit 106 orthogonally transforms the prediction residual of the current block using the conversion types included in the first conversion type group (step Su_2). Next, the conversion unit 106 outputs information indicating which conversion type among one or more conversion types included in the first conversion type group is used to the entropy encoding unit 110, thereby causing the information to be encoded (step Su_3). On the other hand, if the conversion unit 106 determines that the conversion size is not less than or equal to the predetermined value (No in step Su_1), the conversion unit 106 orthogonally transforms the prediction residual of the current block using the second conversion type group (step Su_4).

[0156] In step Su_3, the information indicating the conversion type used for the orthogonal transformation may be information indicating a combination of the conversion type applied in the vertical direction and the conversion type applied in the horizontal direction of the current block. Also, the first conversion type group may include only one conversion type, and the information indicating the conversion type used for the orthogonal transformation may not be encoded. The second conversion type group may include a plurality of conversion types, and information indicating the conversion type used for the orthogonal transformation among one or more conversion types included in the second conversion type group may be encoded.

[0157] Also, the conversion type may be determined based only on the conversion size. Note that as long as it is a process of determining the conversion type used for the orthogonal transformation based on the conversion size, it is not limited to the determination of whether the conversion size is less than or equal to a predetermined value.

[0158] [Quantization Unit] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106. Specifically, the quantization unit 108 scans a plurality of conversion coefficients of the current block in a predetermined scanning order, and quantizes the scanned conversion coefficients based on quantization parameters (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit 108 outputs the plurality of quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0159] The predetermined scanning order is an order for quantization / inverse quantization of conversion coefficients. For example, the predetermined scanning order is defined in ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).

[0160] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, as the value of the quantization parameter increases, the quantization step also increases. That is, as the value of the quantization parameter increases, the error of the quantization coefficient (quantization error) increases.

[0161] Also, a quantization matrix may be used for quantization. For example, several types of quantization matrices may be used corresponding to frequency conversion sizes such as 4x4 and 8x8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and color difference. Note that quantization refers to digitizing values sampled at predetermined intervals by associating them with predetermined levels, and in this technical field, expressions such as rounding, rounding, or scaling may also be used.

[0162] As methods of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device 100 side and a method of using a default quantization matrix (default matrix). On the encoding device 100 side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a demerit that the amount of code increases due to the encoding of the quantization matrix. Instead of using the default quantization matrix or the encoded quantization matrix as it is, a quantization matrix used for quantization of the current block may be generated based on the default quantization matrix or the encoded quantization matrix.

[0163] On the other hand, there is also a method of not using a quantization matrix and quantizing both the coefficients of the high-frequency components and the coefficients of the low-frequency components in the same way. Note that this method is equivalent to a method of using a quantization matrix (flat matrix) in which all the coefficients have the same value.

[0164] The quantization matrix may be encoded, for example, at the sequence level, picture level, slice level, block level, or CTU level.

[0165] When using a quantization matrix, the quantization unit 108 scales, for example, for each transform coefficient, a quantization width obtained from quantization parameters or the like using the value of the quantization matrix. The quantization process performed without using a quantization matrix may be a process of quantizing transform coefficients based on a quantization width obtained from quantization parameters or the like. In the quantization process performed without using a quantization matrix, a predetermined value common to all the transform coefficients in the block may be multiplied by the quantization width.

[0166] FIG. 18 is a block diagram showing an example of the functional configuration of the quantization unit 108.

[0167] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108a, a predictive quantization parameter generation unit 108b, a quantization parameter generation unit 108c, a quantization parameter storage unit 108d, and a quantization processing unit 108e.

[0168] FIG. 19 is a flowchart showing an example of quantization by the quantization unit 108.

[0169] As an example, the quantization unit 108 may perform quantization for each CU based on the flowchart shown in FIG. 19. Specifically, the quantization parameter generation unit 108c determines whether to perform quantization (step Sv_1). Here, if it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generation unit 108c generates the quantization parameter of the current block (step Sv_2) and stores the quantization parameter in the quantization parameter storage unit 108d (step Sv_3).

[0170] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameter generated in step Sv_2 (step Sv_4). Then, the predictive quantization parameter generation unit 108b acquires the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). The predictive quantization parameter generation unit 108b generates the predictive quantization parameter of the current block based on the acquired quantization parameter (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameter of the current block generated by the quantization parameter generation unit 108c and the predictive quantization parameter of the current block generated by the predictive quantization parameter generation unit 108b (step Sv_7). By calculating this difference, a differential quantization parameter is generated. The differential quantization parameter generation unit 108a causes the entropy encoding unit 110 to encode the differential quantization parameter by outputting the differential quantization parameter to the entropy encoding unit 110 (step Sv_8).

[0171] Note that the differential quantization parameter may be encoded at the sequence level, picture level, slice level, block level, or CTU level. Also, the initial value of the quantization parameter may be encoded at the sequence level, picture level, slice level, block level, or CTU level. At this time, the quantization parameter may be generated using the initial value of the quantization parameter and the differential quantization parameter.

[0172] Note that the quantization unit 108 may include a plurality of quantizers, and dependent quantization may be applied to quantize the transform coefficients using a quantization method selected from a plurality of quantization methods.

[0173] [Entropy Encoding Unit] FIG. 20 is a block diagram showing an example of the functional configuration of the entropy encoding unit 110.

[0174] The entropy encoding unit 110 generates a stream by performing entropy encoding on the quantized coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130. For this entropy encoding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy encoding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization that converts multi-value signals such as quantized coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, and the like. The context control unit 110b derives a context value, that is, the occurrence probability of a binary signal, according to the characteristics of the syntax element or the surrounding situation. Examples of methods for deriving this context value include bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic coding unit 110c performs arithmetic coding on the binarized signal using the derived context value.

[0175] FIG. 21 is a diagram showing the flow of CABAC in the entropy encoding unit 110.

[0176] First, in CABAC in the entropy encoding unit 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding unit 110c and setting of initial context values are performed. Then, the binarization unit 110a and the binary arithmetic coding unit 110c sequentially perform binarization and arithmetic coding on each of the plurality of quantized coefficients of, for example, a CTU. At this time, the context control unit 110b updates the context value every time arithmetic coding is performed. Then, as post-processing, the context control unit 110b saves the context value. This saved context value is used, for example, as the initial value of the context value for the next CTU.

[0177] [Inverse quantization unit] The inverse quantization unit 112 inverse quantizes the quantization coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 inverse quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0178] [Inverse transform unit] The inverse transform unit 114 restores the prediction residual by inverse-transforming the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 performs an inverse transform corresponding to the transform by the transform unit 106 on the transform coefficients to restore the prediction residual of the current block. Then, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.

[0179] Note that since information is usually lost due to quantization, the restored prediction residual does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual usually includes a quantization error.

[0180] [Addition unit] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 and the prediction image input from the prediction control unit 128. As a result, a reconstructed image is generated. Then, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.

[0181] [Block memory] The block memory 118 is, for example, a storage unit for storing blocks within the current picture that are blocks referred to in intra prediction. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.

[0182] [Frame memory] The frame memory 122 is a storage unit for storing reference pictures used for, for example, inter prediction, and is sometimes called a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.

[0183] [Loop filter unit] The loop filter unit 120 performs loop filter processing on the reconstructed image output from the addition unit 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter (in-loop filter) used within the encoding loop, and includes, for example, an adaptive loop filter (ALF), a deblocking filter (DF or DBF), and a sample adaptive offset (SAO).

[0184] FIG. 22 is a block diagram showing an example of the functional configuration of the loop filter unit 120.

[0185] The loop filter unit 120 includes, for example, as shown in FIG. 22, a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 120b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. Also, the ALF processing unit 120c applies the above-described ALF processing to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is a process for improving image quality by reducing ringing (a phenomenon in which pixel values fluctuate like waves around the edge) and correcting pixel value shifts. This SAO processing includes, for example, edge offset processing and band offset processing. Note that the loop filter unit 120 does not necessarily include all the processing units disclosed in FIG. 22, and may include only some of the processing units. Also, the loop filter unit 120 may be configured to perform the above-described respective processes in an order different from the processing order disclosed in FIG. 22.

[0186] [Loop Filter Section > Adaptive Loop Filter] In the ALF, a least-squares error filter for removing encoding distortion is applied, and for example, for each 2x2 pixel sub-block within the current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.

[0187] Specifically, first, sub-blocks (for example, 2x2 pixel sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of sub-blocks is performed based on, for example, the direction and activity of the gradient. In a specific example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the sub-blocks are classified into a plurality of classes.

[0188] The gradient direction value D is derived, for example, by comparing gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). Also, the gradient activity value A is derived, for example, by adding gradients in a plurality of directions and quantizing the addition result.

[0189] Based on the results of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0190] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is utilized. FIGS. 23A to 23C are diagrams showing a plurality of examples of the shape of the filter used in the ALF. FIG. 23A shows a 5x5 diamond-shaped filter, FIG. 23B shows a 7x7 diamond-shaped filter, and FIG. 23C shows a 9x9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Note that the signaling of information indicating the shape of the filter is not necessarily limited to the picture level and may be at other levels (for example, sequence level, slice level, block level, CTU level, or CU level).

[0191] The ON / OFF of ALF may be determined, for example, at the picture level or the CU level. For example, for luminance, it may be determined whether to apply ALF at the CU level, and for chrominance difference, it may be determined whether to apply ALF at the picture level. The information indicating the ON / OFF of ALF is usually signaled at the picture level or the CU level. Note that the signaling of the information indicating the ON / OFF of ALF is not necessarily limited to the picture level or the CU level, and it may be at other levels (e.g., sequence level, slice level, block level, or CTU level).

[0192] Also, as described above, one filter is selected from a plurality of filters and ALF processing is performed on the sub-block. For each of the plurality of filters (e.g., up to 15 or 25 filters), the set of coefficients consisting of a plurality of coefficients used for the filter is usually signaled at the picture level. Note that the signaling of the coefficient set is not necessarily limited to the picture level, and it may be at other levels (e.g., sequence level, slice level, block level, CTU level, CU level, or sub-block level).

[0193] [Loop Filter > Cross Component Adaptive Loop Filter] FIG. 23D is a diagram showing an example in which Y samples (first component) are used for CCALF of Cb and CCALF of Cr (a plurality of components different from the first component). FIG. 23E is a diagram showing a diamond-shaped filter.

[0194] One example of CC-ALF operates by applying a linear diamond filter (Figs. 23D, 23E) to the luminance channels of each chrominance component. For example, the filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled with a variable block size and signaled with a context-encoded flag received for each block of samples. The block size and the CC-ALF enable flag are received at the slice level for each chrominance component. The syntax and semantics of CC-ALF are provided in the Appendix. Contributions support block sizes of 16x16, 32x32, 64x64, 128x128 (in chrominance samples).

[0195] [Loop Filter > Joint Chroma Cross Component Adaptive Loop Filter] Fig. 23F is a diagram showing an example of JC-CCALF. Fig. 23G is a diagram showing an example of the weight_index candidates of JC-CCALF.

[0196] One example of JC-CCALF uses only one CCALF filter to generate one CCALF filter output as a chrominance adjustment signal for only one color component and applies a properly weighted version of the same chrominance adjustment signal to the other color components. In this way, the complexity of the existing CCALF is approximately halved.

[0197] The weight values are encoded into a sign flag and a weight index. The weight index (denoted as weight_index) is encoded into 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight. It cannot be made equal to 0. The magnitude of JcCcWeight is determined as follows.

[0198] · If weight_index is 4 or less, JcCcWeight is equal to weight_index >> 2.

[0199] · Otherwise, JcCcWeight is equal to 4 / (weight_index - 4).

[0200] The block-level on / off control of the ALF filtering for Cb and Cr is separate. This is the same as CCALF, and two separate sets of block-level on / off control flags are coded. Here, different from CCALF, since the on / off control block sizes for Cb and Cr are the same, only one block size variable is coded.

[0201] [Loop Filter Section > Deblocking Filter] In the deblocking filter process, the loop filter section 120 reduces the distortion occurring at the block boundary of the reconstructed image by performing filtering at the block boundary.

[0202] FIG. 24 is a block diagram showing an example of the detailed configuration of the deblocking filter processing section 120a.

[0203] The deblocking filter processing section 120a includes, for example, a boundary determination section 1201, a filter determination section 1203, a filter processing section 1205, a processing determination section 1208, a filter characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0204] The boundary determination section 1201 determines whether or not the pixel to be deblocking-filtered (i.e., the target pixel) exists near the block boundary. Then, the boundary determination section 1201 outputs the determination result to the switches 1202 and the processing determination section 1208.

[0205] When the boundary determination unit 1201 determines that the target pixel exists near the block boundary, the switch 1202 outputs the image before the filter process to the switch 1204. Conversely, when the boundary determination unit 1201 determines that the target pixel does not exist near the block boundary, the switch 1202 outputs the image before the filter process to the switch 1206. Note that the image before the filter process is an image composed of the target pixel and at least one surrounding pixel around the target pixel.

[0206] The filter determination unit 1203 determines whether to perform deblocking filter processing on the target pixel based on the pixel values of at least one surrounding pixel around the target pixel. Then, the filter determination unit 1203 outputs the determination result to the switch 1204 and the process determination unit 1208.

[0207] When the filter determination unit 1203 determines that deblocking filter processing is to be performed on the target pixel, the switch 1204 outputs the image before the filter process obtained via the switch 1202 to the filter processing unit 1205. Conversely, when the filter determination unit 1203 determines that deblocking filter processing is not to be performed on the target pixel, the switch 1204 outputs the image before the filter process obtained via the switch 1202 to the switch 1206.

[0208] When the filter processing unit 1205 obtains the image before the filter process via the switches 1202 and 1204, it performs deblocking filter processing having the filter characteristics determined by the filter characteristic determination unit 1207 on the target pixel. Then, the filter processing unit 1205 outputs the pixel after the filter process to the switch 1206.

[0209] The switch 1206 selectively outputs the pixel that has not been subjected to deblocking filter processing and the pixel that has been subjected to deblocking filter processing by the filter processing unit 1205 according to the control by the process determination unit 1208.

[0210] The processing determination unit 1208 controls the switch 1206 based on the determination results of the boundary determination unit 1201 and the filter determination unit 1203 respectively. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filter determination unit 1203 determines that the target pixel is to be subjected to deblocking filter processing, the processing determination unit 1208 causes the deblocking filter-processed pixel to be output from the switch 1206. Also, in cases other than the above, the processing determination unit 1208 causes the non-deblocking filter-processed pixel to be output from the switch 1206. By repeatedly outputting such pixels, the image after the filter processing is output from the switch 1206. Note that the configuration shown in FIG. 24 is an example of the configuration in the deblocking filter processing unit 120a, and the deblocking filter processing unit 120a may have other configurations.

[0211] FIG. 25 is a diagram showing an example of a deblocking filter having filter characteristics symmetric with respect to a block boundary.

[0212] In the deblocking filter processing, for example, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, namely the strong filter and the weak filter, is selected. In the strong filter, as shown in FIG. 25, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the respective pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.

[0213] q’0=(p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8 q’1=(p0 + q0 + q1 + q2 + 2) / 4 q’2=(p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0214] In the above equations, p0 to p2 and q0 to q2 are the pixel values of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3 adjacent to pixel q2 on the side opposite to the block boundary. Further, in the right side of each of the above equations, the coefficient multiplied by the pixel value of each pixel used in the deblocking filter process is the filter coefficient.

[0215] Furthermore, in the deblocking filter process, clip processing may be performed so that the pixel value after the operation does not change beyond the threshold value. In this clip processing, the pixel value after the operation by the above equation is clipped to "the pixel value before the operation ± 2 × the threshold value" using the threshold value determined from the quantization parameter. This can prevent excessive smoothing.

[0216] FIG. 26 is a diagram for explaining an example of a block boundary where deblocking filter processing is performed. FIG. 27 is a diagram showing an example of the BS value.

[0217] The block boundary where deblocking filter processing is performed is, for example, the boundary of a CU, PU, or TU of an 8×8 pixel block as shown in FIG. 26. The deblocking filter processing is performed, for example, in units of 4 rows or 4 columns. First, for blocks P and Q shown in FIG. 26, a Bs (Boundary Strength) value is determined as shown in FIG. 27.

[0218] According to the Bs value in FIG. 27, even for block boundaries belonging to the same image, it may be determined whether to perform deblocking filter processing with different strengths. The deblocking filter processing for the chrominance signal is performed when the Bs value is 2. The deblocking filter processing for the luminance signal is performed when the Bs value is 1 or more and a predetermined condition is satisfied. Note that the determination condition of the Bs value is not limited to that shown in FIG. 27 and may be determined based on other parameters.

[0219] [Prediction unit (intra prediction unit · inter prediction unit · prediction control unit)] FIG. 28 is a flowchart showing an example of the processing performed by the prediction unit of the encoding apparatus 100. As an example, the prediction unit may include all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.

[0220] The prediction unit generates a predicted image of the current block (step Sb_1). The predicted image includes, for example, an intra predicted image (intra prediction signal) or an inter predicted image (inter prediction signal). Specifically, the prediction unit generates a predicted image of the current block using the reconstructed image that has already been obtained by performing generation of a predicted image for other blocks, generation of a prediction residual, generation of quantization coefficients, restoration of the prediction residual, and addition of the predicted image.

[0221] The reconstructed image may be, for example, an image of a reference picture, or an image of an encoded block (i.e., the other blocks described above) in the current picture that includes the current block. The encoded block in the current picture is, for example, an adjacent block of the current block.

[0222] FIG. 29 is a flowchart showing another example of the processing performed by the prediction unit of the encoding apparatus 100.

[0223] The prediction unit generates a predicted image in a first method (step Sc_1a), generates a predicted image in a second method (step Sc_1b), and generates a predicted image in a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter prediction method, an intra prediction method, and other prediction methods, respectively. In these prediction methods, the above-described reconstructed image may be used.

[0224] Next, the prediction unit evaluates the predicted images generated in each of steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit calculates the cost C for the predicted images generated in each of steps Sc_1a, Sc_1b, and Sc_1c, and evaluates those predicted images by comparing the costs C of those predicted images. Note that the cost C is calculated by an equation of an R-D optimization model, for example, C = D + λ × R. In this equation, D is the encoding distortion of the predicted image, and is represented by, for example, the sum of absolute differences between the pixel values of the current block and the pixel values of the predicted image. Also, R is the bit rate of the stream. Also, λ is, for example, the Lagrange multiplier.

[0225] Next, the prediction unit selects any one of the predicted images generated in each of steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects a method or mode for obtaining the final predicted image. For example, the prediction unit selects the predicted image with the smallest cost C based on the cost C calculated for those predicted images. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be performed based on parameters used in the encoding process. The encoding device 100 may signal information for specifying the selected predicted image, method, or mode to the stream. The information may be, for example, a flag. Thereby, the decoding device 200 can generate a predicted image according to the method or mode selected in the encoding device 100 based on that information. Note that in the example shown in FIG. 29, the prediction unit selects one of the predicted images after generating the predicted images by each method. However, the prediction unit may select a method or mode based on the parameters used in the above-described encoding process before generating those predicted images, and generate a predicted image according to that method or mode.

[0226] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction unit may select the final prediction image for the current block from the prediction images generated according to these prediction methods.

[0227] FIG. 30 is a flowchart showing another example of the processing performed by the prediction unit of the encoding apparatus 100.

[0228] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a) and generates a prediction image by inter prediction (step Sd_1b). Note that the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0229] Next, the prediction unit evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above-described cost C may be used for this evaluation. Then, the prediction unit may select, as the final prediction image of the current block, the prediction image for which the smallest cost C has been calculated from the intra prediction image and the inter prediction image (step Sd_3). That is, a prediction method or mode for generating the prediction image of the current block is selected.

[0230] [Intra Prediction Unit] The intra prediction unit 124 generates a prediction image (i.e., an intra prediction image) of the current block by performing intra prediction (also referred to as in-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118. Specifically, the intra prediction unit 124 generates an intra prediction image by performing intra prediction with reference to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control unit 128.

[0231] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0232] The one or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC standard.

[0233] The plurality of directional prediction modes include, for example, the 33-direction prediction mode defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may further include a 32-direction prediction mode (a total of 65 directional prediction modes) in addition to the 33 directions. FIG. 31 is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent the 33 directions defined in the H.265 / HEVC standard, and the dashed arrows represent the additional 32 directions (the 2 non-directional prediction modes are not shown in FIG. 31).

[0234] In various implementation examples, in the intra prediction of a chrominance block, a luminance block may be referenced. That is, based on the luminance component of the current block, the chrominance component of the current block may be predicted. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that references such a luminance block (for example, called the CCLM mode) may be added as one of the intra prediction modes of the chrominance block.

[0235] The intra prediction unit 124 may correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. Such intra prediction with such correction is sometimes called PDPC (position dependent intra prediction combination). Information indicating the presence or absence of the application of PDPC (for example, called a PDPC flag) is usually signaled at the CU level. Note that the signaling of this information does not have to be limited to the CU level and may be at other levels (for example, sequence level, picture level, slice level, block level or CTU level).

[0236] FIG. 32 is a flowchart showing an example of the processing by the intra prediction unit 124.

[0237] The intra prediction unit 124 selects one intra prediction mode from a plurality of intra prediction modes (step Sw_1). Then, the intra prediction unit 124 generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM consists of, for example, six intra prediction modes. Two of the six intra prediction modes may be the Planar prediction mode and the DC prediction mode, and the remaining four modes may be directional prediction modes. Then, the intra prediction unit 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0238] Here, when it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode among the MPM (step Sw_6). Note that the MPM flag set to 1 and the information indicating the intra prediction mode are each encoded by the entropy encoding unit 110 as prediction parameters.

[0239] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit 124 sets the MPM flag to 0 (step Sw_7). Or, the intra prediction unit 124 does not set the MPM flag. Then, the intra prediction unit 124 generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). Note that the MPM flag set to 0 and the information indicating the intra prediction mode are each encoded by the entropy encoding unit 110 as prediction parameters. The information indicating the intra prediction mode indicates, for example, any value between 0 and 60.

[0240] [Inter Prediction Unit] The inter prediction unit 126 generates a predicted image (inter prediction image) by performing inter prediction (also called inter-frame prediction) of the current block with reference to a reference picture stored in the frame memory 122 and different from the current picture. The inter prediction is performed in units of the current block or a current sub-block within the current block. The sub-block is included in the block and is a unit smaller than the block. The size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may be switched in units such as slices, bricks, or pictures.

[0241] For example, the inter prediction unit 126 performs motion estimation within the reference picture for the current block or current sub-block, and finds the reference block or sub-block that most closely matches the current block or current sub-block. Then, the inter prediction unit 126 obtains motion information (e.g., a motion vector) for compensating the motion or change from the reference block or sub-block to the current block or sub-block. Based on the motion information, the inter prediction unit 126 performs motion compensation (or motion prediction) to generate an inter prediction image of the current block or sub-block. The inter prediction unit 126 outputs the generated inter prediction image to the prediction control unit 128.

[0242] The motion information used for motion compensation may be signaled as an inter prediction image in various forms. For example, the motion vector may be signaled. As another example, the difference between the motion vector and the motion vector predictor may be signaled.

[0243] [Reference Picture List] FIG. 33 is a diagram showing an example of each reference picture, and FIG. 34 is a conceptual diagram showing an example of a reference picture list. The reference picture list is a list showing one or more reference pictures stored in the frame memory 122. In FIG. 33, a rectangle represents a picture, an arrow represents a reference relationship between pictures, the horizontal axis represents time, I, P, and B in the rectangle represent an intra prediction picture, a single prediction picture, and a bi-prediction picture, respectively, and the numbers in the rectangle represent the decoding order. As shown in FIG. 33, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As shown in FIG. 34, the reference picture list is a list representing candidates for reference pictures, and for example, one picture (or slice) may have one or more reference picture lists. For example, if the current picture is a single prediction picture, one reference picture list is used, and if the current picture is a bi-prediction picture, two reference picture lists are used. In the examples of FIGS. 33 and 34, the picture B3 which is the current picture currPic has two reference picture lists, the L0 list and the L1 list. When the current picture currPic is the picture B3, the candidates for the reference pictures of the current picture currPic are I0, P1, and B2, and each reference picture list (that is, the L0 list and the L1 list) shows these pictures. The inter prediction unit 126 or the prediction control unit 128 designates which picture in each reference picture list is actually referred to by the reference picture index refidxLx. In FIG. 34, the reference pictures P1 and B2 are designated by the reference picture indexes refIdxL0 and refIdxL1.

[0244] Such a reference picture list may be generated in units of sequence, picture, slice, block, CTU, or CU. Further, among the reference pictures shown in the reference picture list, a reference picture index indicating the reference picture referred to in inter prediction may be coded at sequence level, picture level, slice level, block level, CTU level, or CU level. Also, a common reference picture list may be used in a plurality of inter prediction modes.

[0245] [Basic Flow of Inter Prediction] FIG. 35 is a flowchart showing the basic flow of inter prediction.

[0246] The inter prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0247] Here, in generating the predicted image, the inter prediction unit 126 generates the predicted image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Also, in determining the MV, the inter prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Note that MVs derived in the past may be added as candidate MVs to the candidate MV list. Also, in deriving the MV, the inter prediction unit 126 may determine the selected at least one candidate MV as the MV of the current block by selecting at least one more candidate MV from the at least one candidate MV. Alternatively, the inter prediction unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by the candidate MV for each of the selected at least one candidate MVs. Note that searching the region of the reference picture may be referred to as motion estimation.

[0248] Also, in the above example, steps Se_1 to Se_3 are performed by the inter prediction unit 126, but processing such as step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.

[0249] Note that a candidate MV list may be created for each process in each inter prediction mode, or a common candidate MV list may be used in a plurality of inter prediction modes. Also, the processes of steps Se_3 and Se_4 respectively correspond to the processes of steps Sa_3 and Sa_4 shown in FIG. 9. Also, the process of step Se_3 corresponds to the process of step Sd_1b in FIG. 30.

[0250] [Flow of MV Derivation] FIG. 36 is a flowchart showing an example of MV derivation.

[0251] The inter prediction unit 126 may derive the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information may be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.

[0252] Alternatively, the inter prediction unit 126 may derive the MV in a mode that does not encode motion information. In this case, the motion information is not included in the stream.

[0253] Here, the modes of MV derivation include a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, etc., which will be described later. Among these modes, the modes that encode motion information include a normal inter mode, a normal merge mode, and an affine mode (specifically, an affine inter mode and an affine merge mode). Note that the motion information may include not only the MV but also prediction MV selection information, which will be described later. Also, the modes that do not encode motion information include a FRUC mode. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and derives the MV of the current block using the selected mode.

[0254] FIG. 37 is a flowchart showing another example of MV derivation.

[0255] The inter prediction unit 126 may derive the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. This differential MV is the difference between the MV of the current block and its predicted MV. Note that the predicted MV is a predicted motion vector.

[0256] Alternatively, the inter prediction unit 126 may derive the MV in a mode where the differential MV is not encoded. In this case, the encoded differential MV is not included in the stream.

[0257] Here, as described above, the modes for deriving the MV include the normal inter, normal merge mode, FRUC mode, and affine mode, etc. Among these modes, the modes for encoding the differential MV include the normal inter mode and the affine mode (specifically, the affine inter mode), etc. Also, the modes for not encoding the differential MV include the FRUC mode, the normal merge mode, and the affine mode (specifically, the affine merge mode), etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0258] [Mode of MV derivation] FIG. 38A and FIG. 38B are diagrams showing an example of classification of each mode of MV derivation. For example, as shown in FIG. 38A, according to whether motion information is encoded or not and whether differential MV is encoded or not, the mode of MV derivation is roughly classified into three modes. The three modes are the inter mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter mode is a mode that performs motion search and encodes motion information and differential MV. For example, as shown in FIG. 38B, the inter mode includes an affine inter mode and a normal inter mode. The merge mode is a mode that does not perform motion search, selects an MV from surrounding encoded blocks, and derives the MV of the current block using that MV. This merge mode is basically a mode that encodes motion information and does not encode differential MV. For example, as shown in FIG. 38B, the merge mode includes a normal merge mode (sometimes also called a regular merge mode or a normal merge mode), an MMVD (Merge with Motion Vector Difference) mode, a CIIP (Combined inter merge / intra prediction) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, in the MMVD mode among each mode included in the merge mode, differentially MV is encoded exceptionally. Note that the above-mentioned affine merge mode and affine inter mode are modes included in the affine mode. The affine mode is a mode that assumes an affine transformation and derives the MV of each of a plurality of sub-blocks constituting the current block as the MV of the current block. The FRUC mode is a mode that derives the MV of the current block by performing a search between encoded regions and does not encode either motion information or differential MV. Note that the details of each of these modes will be described later.

[0259] Note that the classification of each mode shown in FIGS. 38A and 38B is an example and is not limited to this. For example, when the differential MV is encoded in the CIIP mode, the CIIP mode is classified as an inter mode.

[0260] [MV Derivation > Normal Inter Mode] The normal inter mode is an inter prediction mode for deriving the MV of the current block by finding a block similar to the image of the current block from the area of the reference picture indicated by the candidate MV. Also, in this normal inter mode, the differential MV is encoded.

[0261] FIG. 39 is a flowchart showing an example of inter prediction by the normal inter mode.

[0262] The inter prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially around the current block (step Sg_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0263] Next, the inter prediction unit 126 extracts each of N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as a prediction MV candidate according to a predetermined priority order (step Sg_2). Note that the priority order is determined in advance for each of the N candidate MVs.

[0264] Next, the inter prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV of the current block (step Sg_3). At this time, the inter prediction unit 126 encodes the prediction MV selection information for identifying the selected prediction MV into the stream. That is, the inter prediction unit 126 outputs the prediction MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0265] Next, the inter prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter prediction unit 126 further encodes the difference value between the derived MV and the predicted MV as a differential MV into the stream. That is, the inter prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. Note that the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0266] Finally, the inter prediction unit 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are executed for each block. For example, when the processes of steps Sg_1 to Sg_5 are executed for each of all the blocks included in a slice, the inter prediction using the normal inter mode for that slice ends. Also, when the processes of steps Sg_1 to Sg_5 are executed for each of all the blocks included in a picture, the inter prediction using the normal inter mode for that picture ends. Note that even if the processes of steps Sg_1 to Sg_5 are not executed for all the blocks included in a slice and are executed for some of the blocks, the inter prediction using the normal inter mode for that slice may end. Similarly, even if the processes of steps Sg_1 to Sg_5 are executed for some of the blocks included in a picture, the inter prediction using the normal inter mode for that picture may end.

[0267] Note that the predicted image is the above-described inter prediction signal. Also, information indicating the inter prediction mode (normal inter mode in the above example) used for generating the predicted image, which is included in the encoded signal, is encoded as a prediction parameter, for example.

[0268] Note that the candidate MV list may be used in common with the lists used in other modes. Also, the processing related to the candidate MV list may be applied to the processing related to the lists used in other modes. The processing related to this candidate MV list is, for example, extraction or selection of candidate MVs from the candidate MV list, rearrangement of candidate MVs, or deletion of candidate MVs.

[0269] [MV Derivation > Normal Merge Mode] The normal merge mode is an inter prediction mode that derives an MV by selecting a candidate MV from the candidate MV list as the MV of the current block. Note that the normal merge mode is a merge mode in a narrow sense and may be simply called the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.

[0270] FIG. 40 is a flowchart showing an example of inter prediction by the normal merge mode.

[0271] First, the inter prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially around the current block (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0272] Next, the inter prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream. That is, the inter prediction unit 126 outputs the MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0273] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processes of steps Sh_1 to Sh_3 are executed for each block, for example. For example, when the processes of steps Sh_1 to Sh_3 are executed for each of all the blocks included in a slice, the inter prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_1 to Sh_3 are executed for each of all the blocks included in a picture, the inter prediction using the normal merge mode for that picture ends. Note that when the processes of steps Sh_1 to Sh_3 are not executed for all the blocks included in a slice but are executed for some of the blocks, the inter prediction using the normal merge mode for that slice may end. Similarly, when the processes of steps Sh_1 to Sh_3 are executed for some of the blocks included in a picture, the inter prediction using the normal merge mode for that picture may end.

[0274] In addition, information indicating the inter prediction mode (normal merge mode in the above example) used for generating the predicted image, which is included in the stream, is encoded as prediction parameters, for example.

[0275] FIG. 41 is a diagram for explaining an example of the MV derivation process of the current picture in the normal merge mode.

[0276] First, the inter prediction unit 126 generates a candidate MV list in which candidate MVs are registered. Examples of candidate MVs include a spatial adjacent candidate MV which is an MV that a plurality of encoded blocks spatially adjacent to the current block have, a temporal adjacent candidate MV which is an MV that a nearby block obtained by projecting the position of the current block in the encoded reference picture has, a combined candidate MV which is an MV generated by combining the MV values of the spatial adjacent candidate MV and the temporal adjacent candidate MV, and a zero candidate MV which is an MV with a value of zero.

[0277] Next, the inter prediction unit 126 determines one candidate MV as the MV of the current block by selecting one candidate MV from among the plurality of candidate MVs registered in the candidate MV list.

[0278] Furthermore, the entropy encoding unit 110 describes and encodes the merge_idx, which is a signal indicating which candidate MV has been selected, into the stream.

[0279] Note that the candidate MVs registered in the candidate MV list described in FIG. 41 are merely examples, and the number may be different from that in the figure, the configuration may not include some of the types of candidate MVs in the figure, or the configuration may include candidate MVs other than the types of candidate MVs in the figure.

[0280] The final MV may be determined by performing DMVR (dynamic motion vector refreshing), which will be described later, using the MV of the current block derived in the normal merge mode. Note that in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same manner as the normal merge mode, but encodes the differential MV. Such MMVD may be classified into the merge mode together with the normal merge mode as shown in FIG. 38B. Note that the differential MV in the MMVD mode does not have to be the same as the differential MV used in the inter mode. For example, the derivation of the differential MV in the MMVD mode may be a process with a smaller processing amount compared to the derivation of the differential MV in the inter mode.

[0281] Also, a CIIP (Combined inter merge / intra prediction) mode may be performed in which the predicted image generated by inter prediction and the predicted image generated by intra prediction are superimposed to generate the predicted image of the current block.

[0282] Note that the candidate MV list may also be referred to as the candidate list. Also, merge_idx is MV selection information.

[0283] [MV Derivation > HMVP Mode] FIG. 42 is a diagram for explaining an example of the MV derivation process of the current picture in the HMVP mode.

[0284] In the normal merge mode, the MV of a current block, for example, a CU, is determined by selecting one candidate MV from a candidate MV list generated with reference to an encoded block (for example, a CU). Here, other candidate MVs may be registered in the candidate MV list. A mode in which such other candidate MVs are registered is called the HMVP mode.

[0285] In the HMVP mode, candidate MVs are managed using a FIFO (First-In First-Out) buffer for HMVP separately from the candidate MV list in the normal merge mode.

[0286] In the FIFO buffer, motion information such as the MVs of blocks processed in the past is stored in order from the newest. In the management of this FIFO buffer, each time a block is processed, the MV of the newest block (that is, the CU processed immediately before) is stored in the FIFO buffer, and instead, the MV of the oldest CU (that is, the CU processed earliest) in the FIFO buffer is deleted from the FIFO buffer. In the example shown in FIG. 42, HMVP1 is the MV of the newest block, and HMVP5 is the MV of the oldest block.

[0287] Then, for example, the inter prediction unit 126 checks, in order from HMVP1, whether each MV managed in the FIFO buffer is different from all the candidate MVs already registered in the candidate MV list in the normal merge mode. Then, when the inter prediction unit 126 determines that it is different from all the candidate MVs, the MV managed in the FIFO buffer may be added as a candidate MV to the candidate MV list in the normal merge mode. At this time, one or more candidate MVs registered from the FIFO buffer may be added.

[0288] By using the HMVP mode in this way, it becomes possible to add not only the MVs of spatially or temporally adjacent blocks of the current block but also the MVs of blocks processed in the past to the candidates. As a result, the variations of the candidate MVs in the normal merge mode are widened, increasing the likelihood of improving the coding efficiency.

[0289] Note that the above-mentioned MV may be motion information. That is, the information stored in the candidate MV list and the FIFO buffer may include not only the value of the MV but also information indicating the information of the reference picture, the reference direction, the number of pictures, etc. Further, the above-mentioned block is, for example, a CU.

[0290] Note that the candidate MV list and the FIFO buffer in FIG. 42 are merely examples, and the candidate MV list and the FIFO buffer may be a list or buffer of a size different from that in FIG. 42, or may have a configuration in which candidate MVs are registered in an order different from that in FIG. 42. Further, the processing described here is common to both the encoding device 100 and the decoding device 200.

[0291] Note that the HMVP mode can also be applied to modes other than the normal merge mode. For example, motion information such as the MVs of blocks processed in the affine mode in the past may be stored in the FIFO buffer in order from the newest and used as candidate MVs. The mode in which the HMVP mode is applied to the affine mode may be called the history affine mode.

[0292] [MV Derivation > FRUC Mode] Motion information may be derived on the decoder 200 side without being signaled from the encoder 100 side. For example, motion information may be derived by performing motion search on the decoder 200 side. In this case, on the decoder 200 side, motion search is performed without using the pixel values of the current block. Modes in which such motion search is performed on the decoder 200 side include, for example, the FRUC (frame rate up-conversion) mode or the PMMVD (pattern matched motion vector derivation) mode.

[0293] An example of FRUC processing is shown in FIG. 43. First, with reference to the MVs of each encoded block spatially or temporally adjacent to the current block, a list showing those MVs as candidate MVs (that is, a candidate MV list, which may be common to the candidate MV list in the normal merge mode) is generated (step Si_1). Next, the best candidate MV is selected from among the plurality of candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate MV is selected as the best candidate MV based on that evaluation value. Then, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the selected best candidate MV, by performing pattern matching, the MV for the current block may be derived. That is, for the region around the best candidate MV, search using pattern matching and evaluation values in the reference picture is performed, and if there is an MV with a better evaluation value, the best candidate MV may be updated to that MV and used as the final MV for the current block. It is not necessary to perform the update to an MV with a better evaluation value.

[0294] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_1 to Si_5 are executed, for example, for each block. For example, when the processes of steps Si_1 to Si_5 are executed for each of all the blocks included in a slice, the inter prediction using the FRUC mode for that slice ends. Also, when the processes of steps Si_1 to Si_5 are executed for each of all the blocks included in a picture, the inter prediction using the FRUC mode for that picture ends. Note that when the processes of steps Si_1 to Si_5 are not executed for all the blocks included in a slice but are executed for some blocks, the inter prediction using the FRUC mode for that slice may end. Similarly, when the processes of steps Si_1 to Si_5 are executed for some of the blocks included in a picture, the inter prediction using the FRUC mode for that picture may end.

[0295] Processing may be performed in the same manner as in the above-described block unit even in units of sub-blocks.

[0296] The evaluation value may be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the MV is compared with the reconstructed image of a predetermined region (the region may be, for example, as shown below, a region of another reference picture or a region of an adjacent block of the current picture). Then, the difference between the pixel values of the two reconstructed images may be calculated and used as the evaluation value of the MV. Note that information other than the difference value may be used to calculate the evaluation value.

[0297] Next, pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (also referred to as the merge list) is selected as the start point of the search by pattern matching. As the pattern matching, the first pattern matching or the second pattern matching may be used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.

[0298] [MV Derivation > FRUC > Bilateral Matching] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are along the motion trajectory of the current block. Therefore, in the first pattern matching, as the predetermined region for calculating the evaluation value of the candidate MV described above, a region in another reference picture along the motion trajectory of the current block is used.

[0299] FIG. 44 is a diagram for explaining an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As shown in FIG. 44, in the first pattern matching, two MVs (MV0, MV1) are derived by searching for the pair that most matches among pairs of two blocks in two different reference pictures (Ref0, Ref1) that are along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. It is preferable that the candidate MV having the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.

[0300] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional MVs are derived.

[0301] [MV Derivation > FRUC > Template Matching] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (a block adjacent to the current block within the current picture (e.g., the upper and / or left adjacent blocks)) and the blocks in the reference picture. Therefore, in the second pattern matching, the blocks adjacent to the current block within the current picture are used as the predetermined region for calculating the evaluation value of the candidate MVs described above.

[0302] FIG. 45 is a diagram for explaining an example of pattern matching (template matching) between the template in the current picture and the blocks in the reference picture. As shown in FIG. 45, in the second pattern matching, the MV of the current block is derived by searching for the block in the reference picture (Ref0) that most matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of both or either one of the left adjacent and upper adjacent encoded regions and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. It is preferable that the candidate MV with the best evaluation value among the plurality of candidate MVs is selected as the best candidate MV.

[0303] Information indicating whether or not to apply such a FRUC mode (e.g., called a FRUC flag) may be signaled at the CU level. Also, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating an applicable pattern matching method (first pattern matching or second pattern matching) may be signaled at the CU level. Note that the signaling of this information need not be limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, block level, CTU level, or sub-block level).

[0304] [MV Derivation > Affine Mode] The affine mode is a mode that generates an MV using an affine transformation. For example, an MV may be derived for each sub-block based on the MVs of a plurality of adjacent blocks. This mode may sometimes be called an affine motion compensation prediction mode.

[0305] FIG. 46A is a diagram for explaining an example of the derivation of an MV for each sub-block based on the MVs of a plurality of adjacent blocks. In FIG. 46A, a current block includes, for example, a sub-block composed of 16 4x4 pixels. Here, a motion vector v0 of the upper left control point of the current block is derived based on the MVs of adjacent blocks. Similarly, a motion vector v1 of the upper right control point of the current block is derived based on the MVs of adjacent sub-blocks. Then, the two motion vectors v0 and v1 are projected by the following formula (1A) to derive a motion vector (vx, vy) for each sub-block within the current block.

[0306] [Equation]

[0307] Here, x and y indicate the horizontal position and vertical position of the sub-block, respectively, and w indicates a predetermined weight coefficient.

[0308] Information indicating such an affine mode (e.g., called an affine flag) may be signaled at the CU level. Note that the signaling of the information indicating this affine mode is not necessarily limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, block level, CTU level, or sub-block level).

[0309] Also, such an affine mode may include several modes in which the derivation methods of the MVs of the upper left and upper right corner control points are different. For example, the affine mode includes two modes: an affine inter (also called affine normal inter) mode and an affine merge mode.

[0310] FIG. 46B is a diagram for explaining an example of the derivation of the MV in units of sub-blocks in an affine mode using three control points. In FIG. 46B, the current block includes, for example, a sub-block composed of 16 4x4 pixels. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the MV of the adjacent block, and the motion vector v2 of the lower left corner control point of the current block is derived based on the MV of the adjacent block. Then, by the following formula (1B), the three motion vectors v0, v1, and v2 are projected to derive the motion vector (vx, vy) of each sub-block in the current block.

[0311] [Number]

[0312] Here, x and y respectively indicate the horizontal position and vertical position of the sub-block center, and w and h indicate predetermined weight coefficients. w may indicate the width of the current block, and h may indicate the height of the current block.

[0313] The affine mode using different numbers of control points (e.g., two and three) may be signaled by switching at the CU level. Note that information indicating the number of control points of the affine mode used at the CU level may be signaled at other levels (e.g., sequence level, picture level, slice level, block level, CTU level, or sub-block level).

[0314] In addition, the affine mode having such three control points may include several modes with different methods for deriving the MVs of the upper left, upper right, and lower left corner control points. For example, the affine mode having three control points includes two modes, namely, the affine inter mode and the affine merge mode, similar to the affine mode having the above-described two control points.

[0315] Note that in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.

[0316] [MV Derivation > Affine Mode > Control Point] FIGS. 47A, 47B, and 47C are conceptual diagrams for explaining an example of MV derivation of control points in the affine mode.

[0317] In the affine mode, as shown in FIG. 47A, for example, among the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, based on a plurality of MVs corresponding to the blocks encoded in the affine mode, the respective predicted MVs of the control points of the current block are calculated. Specifically, these blocks are inspected in the order of the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left), and the first valid block encoded in the affine mode is identified. Based on the plurality of MVs corresponding to this identified block, the MVs of the control points of the current block are calculated.

[0318] For example, as shown in FIG. 47B, when block A adjacent to the left of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 projected onto the upper left and upper right corner positions of the encoded block including block A are derived. Then, from the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point are calculated.

[0319] For example, as shown in FIG. 47C, when block A adjacent to the left of the current block is encoded in an affine mode having three control points, motion vectors v3, v4, and v5 projected onto the upper left, upper right, and lower left corner positions of the encoded block including block A are derived. Then, from the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point of the current block, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point are calculated.

[0320] Note that the method for deriving the MV shown in FIGS. 47A to 47C may be used for deriving the MV of each control point of the current block in step Sk_1 shown in FIG. 50 described later, or may be used for deriving the predicted MV of each control point of the current block in step Sj_1 shown in FIG. 51 described later.

[0321] FIGS. 48A and 48B are conceptual diagrams for explaining another example of deriving the control point MV in the affine mode.

[0322] FIG. 48A is a diagram for explaining an affine mode having two control points.

[0323] In this affine mode, as shown in FIG. 48A, the MV selected from each of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from each of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block.

[0324] FIG. 48B is a diagram for explaining the affine mode having three control points.

[0325] In this affine mode, as shown in FIG. 48B, the MV selected from each of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from each of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block. Further, the MV selected from each of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 of the lower left control point of the current block.

[0326] Note that the method for deriving the MVs shown in FIGS. 48A and 48B may be used for deriving the MVs of each control point of the current block in step Sk_1 shown in FIG. 50 described later, or may be used for deriving the predicted MVs of each control point of the current block in step Sj_1 of FIG. 51 described later.

[0327] Here, for example, when switching and signaling affine modes with different numbers of control points (for example, two and three) at the CU level, there may be a case where the number of control points is different between the encoded block and the current block.

[0328] FIG. 49A and FIG. 49B are conceptual diagrams for explaining an example of a method for deriving the MV of control points when the number of control points in an encoded block and a current block is different.

[0329] For example, as shown in FIG. 49A, the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left of the current block is encoded in an affine mode having two control points. In this case, motion vectors v3 and v4 projected onto the upper left corner and upper right corner positions of the encoded block including block A are derived. Then, from the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point are calculated. Further, from the derived motion vectors v0 and v1, the motion vector v2 of the lower left corner control point is calculated.

[0330] For example, as shown in FIG. 49B, the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left of the current block is encoded in an affine mode having three control points. In this case, motion vectors v3, v4, and v5 projected onto the upper left corner, upper right corner, and lower left corner positions of the encoded block including block A are derived. Then, from the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point are calculated.

[0331] Note that the method for deriving the MV shown in FIGS. 49A and 49B may be used for deriving the MV of each control point of the current block in step Sk_1 shown in FIG. 50 described later, or may be used for deriving the predicted MV of each control point of the current block in step Sj_1 of FIG. 51 described later.

[0332] [MV Derivation > Affine Mode > Affine Merge Mode] FIG. 50 is a flowchart showing an example of the affine merge mode.

[0333] In the affine merge mode, first, the inter prediction unit 126 derives the respective MVs of the control points of the current block (step Sk_1). As shown in FIG. 46A, the control points are the points at the upper left corner and the upper right corner of the current block, or as shown in FIG. 46B, the points at the upper left corner, the upper right corner, and the lower left corner of the current block. At this time, the inter prediction unit 126 may encode the MV selection information for identifying the two or three derived MVs into the stream.

[0334] For example, when using the MV derivation method shown in FIGS. 47A to 47C, the inter prediction unit 126 inspects these blocks in the order of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left) as shown in FIG. 47A, and identifies the first valid block encoded in the affine mode.

[0335] The inter prediction unit 126 derives the MVs of the control points using the first valid block encoded in the identified affine mode. For example, if block A is identified and block A has two control points, as shown in FIG. 47B, the inter prediction unit 126 calculates the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point from the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block including block A. For example, the inter prediction unit 126 calculates the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point by projecting the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block onto the current block.

[0336] Alternatively, when block A is identified and block A has three control points, as shown in FIG. 47C, the inter prediction unit 126 calculates the motion vectors v0, v1, and v2 of the upper left control point, upper right control point, and lower left control point of the current block from the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the encoded block including block A. For example, the inter prediction unit 126 projects the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the encoded block onto the current block to calculate the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block.

[0337] Note that, as shown in FIG. 49A described above, when block A is identified and block A has two control points, the MVs of the three control points may be calculated. As shown in FIG. 49B described above, when block A is identified and block A has three control points, the MVs of the two control points may be calculated.

[0338] Next, the inter prediction unit 126 performs motion compensation for each of the plurality of sub-blocks included in the current block. That is, the inter prediction unit 126 calculates the MV of each of the plurality of sub-blocks as an affine MV using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter prediction unit 126 performs motion compensation for the sub-block using those affine MVs and the encoded reference picture (step Sk_3). When the processes of steps Sk_2 and Sk_3 are executed for each of all the sub-blocks included in the current block, the process of generating a prediction image using the affine merge mode for the current block ends. That is, motion compensation is performed for the current block, and a prediction image of the current block is generated.

[0339] In step Sk_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list including candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47A to 47C, the MV derivation methods shown in FIGS. 48A and 48B, the MV derivation methods shown in FIGS. 49A and 49B, and other MV derivation methods.

[0340] Note that the candidate MV list may include candidate MVs in a mode that performs prediction in units of sub-blocks other than the affine mode.

[0341] Note that, as the candidate MV list, for example, a candidate MV list including candidate MVs in the affine merge mode having two control points and candidate MVs in the affine merge mode having three control points may be generated. Alternatively, a candidate MV list including candidate MVs in the affine merge mode having two control points and a candidate MV list including candidate MVs in the affine merge mode having three control points may be generated respectively. Alternatively, a candidate MV list including candidate MVs in one of the modes of the affine merge mode having two control points and the affine merge mode having three control points may be generated. The candidate MV may be, for example, an MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be an MV of a valid block among those blocks.

[0342] Note that, as the MV selection information, an index indicating which candidate MV in the candidate MV list may be sent.

[0343] [MV Derivation > Affine Mode > Affine Inter Mode] FIG. 51 is a flowchart showing an example of the affine inter mode.

[0344] In the affine inter mode, first, the inter prediction unit 126 derives the prediction MV (v0, v1) or (v0, v1, v2) for each of two or three control points of the current block (step Sj_1). As shown in FIG. 46A or FIG. 46B, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0345] For example, when using the MV derivation method shown in FIGS. 48A and 48B, the inter prediction unit 126 derives the prediction MV (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the MV of any block among the encoded blocks near each control point of the current block shown in FIG. 48A or FIG. 48B. At this time, the inter prediction unit 126 encodes the prediction MV selection information for identifying the selected two or three prediction MVs into the stream.

[0346] For example, the inter prediction unit 126 may determine which block's MV to select as the prediction MV of the control point from the encoded blocks adjacent to the current block using cost evaluation or the like, and describe a flag indicating which prediction MV is selected in the bit stream. That is, the inter prediction unit 126 outputs the prediction MV selection information such as a flag as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0347] Next, while updating the prediction MVs selected or derived in step Sj_1 respectively (step Sj_2), the inter prediction unit 126 performs motion search (steps Sj_3 and Sj_4). That is, the inter prediction unit 126 calculates the MV of each sub-block corresponding to the updated prediction MV as an affine MV using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter prediction unit 126 performs motion compensation on each sub-block using those affine MVs and the encoded reference picture (step Sj_4). The processes of steps Sj_3 and Sj_4 are executed for all the blocks in the current block each time the prediction MV is updated in step Sj_2. As a result, in the motion search loop, the inter prediction unit 126 determines, for example, the prediction MV that obtains the smallest cost as the MV of the control point (step Sj_5). At this time, the inter prediction unit 126 further encodes the difference value between the determined MV and the prediction MV as a differential MV into the stream. That is, the inter prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0348] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).

[0349] Note that in step Sj_1, the above candidate MV list may be generated. The candidate MV list may be, for example, a list including candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47A to 47C, the MV derivation methods shown in FIGS. 48A and 48B, the MV derivation methods shown in FIGS. 49A and 49B, and other MV derivation methods.

[0350] Note that the candidate MV list may include candidate MVs in a mode that performs prediction in units of sub-blocks other than the affine mode.

[0351] Note that as the candidate MV list, a candidate MV list including an affine inter-mode candidate MV having two control points and an affine inter-mode candidate MV having three control points may be generated. Alternatively, a candidate MV list including an affine inter-mode candidate MV having two control points and a candidate MV list including an affine inter-mode candidate MV having three control points may be generated respectively. Alternatively, a candidate MV list including candidate MVs of one of the modes of an affine inter-mode having two control points and an affine inter-mode having three control points may be generated. The candidate MV may be, for example, the MVs of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be the MVs of the valid blocks among those blocks.

[0352] Note that as the prediction MV selection information, an index indicating which candidate MV in the candidate MV list may be sent.

[0353] [MV Derivation > Triangle Mode] In the above example, the inter prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter prediction unit 126 may generate a plurality of prediction images having shapes different from the rectangle for the rectangular current block, and generate a final rectangular prediction image by combining those plurality of prediction images. The shape different from the rectangle may be, for example, a triangle.

[0354] FIG. 52A is a diagram for explaining the generation of two triangular prediction images.

[0355] The inter prediction unit 126 generates a predicted image of the triangle by performing motion compensation on the first partition of the triangle within the current block using the first MV of the first partition. Similarly, the inter prediction unit 126 generates a predicted image of the triangle by performing motion compensation on the second partition of the triangle within the current block using the second MV of the second partition. Then, the inter prediction unit 126 generates a predicted image of the same rectangle as the current block by combining these predicted images.

[0356] Note that as the predicted image of the first partition, a first predicted image of a rectangle corresponding to the current block may be generated using the first MV. Also, as the predicted image of the second partition, a second predicted image of a rectangle corresponding to the current block may be generated using the second MV. The predicted image of the current block may be generated by weighted addition of the first predicted image and the second predicted image. Note that the region for weighted addition may be only a partial region sandwiching the boundary between the first partition and the second partition.

[0357] FIG. 52B is a conceptual diagram showing a first portion of a first partition that overlaps a second partition, and examples of a first sample set and a second sample set that may be weighted as part of a correction process. The first portion may be, for example, one-fourth of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero, and for example, N may be the integer 2. FIG. 52B shows a rectangular partition having a rectangular portion with a width of one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples within the first portion. The central example in FIG. 52B shows a rectangular partition having a rectangular portion with a height of one-fourth of the height of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples within the first portion. The right example in FIG. 52B shows a triangular partition having a polygonal portion with a height corresponding to two samples. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples within the first portion.

[0358] The first portion may be a portion of the first partition that overlaps an adjacent partition. FIG. 52C is a conceptual diagram showing a first portion of a first partition that is a portion of the first partition that overlaps a portion of an adjacent partition. For simplicity of explanation, a rectangular partition having a portion that overlaps a spatially adjacent rectangular partition is shown. Partitions having other shapes, such as triangular partitions, may be used, and the overlapping portion may overlap spatially or temporally adjacent partitions.

[0359] Also, an example of generating a predicted image for each of two partitions using inter prediction is shown, but a predicted image may be generated for at least one partition using intra prediction.

[0360] FIG. 53 is a flowchart showing an example of the triangle mode.

[0361] In the triangle mode, first, the inter prediction unit 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter prediction unit 126 may encode, as prediction parameters, partition information, which is information regarding the division into each partition, into the stream. That is, the inter prediction unit 126 may output the partition information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0362] Next, the inter prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks around the current block temporally or spatially (step Sx_2). That is, the inter prediction unit 126 creates a candidate MV list.

[0363] Then, the inter prediction unit 126 selects, as a first MV and a second MV, a candidate MV for the first partition and a candidate MV for the second partition, respectively, from among the plurality of candidate MVs obtained in step Sx_1 (step Sx_3). At this time, the inter prediction unit 126 may encode, as prediction parameters, MV selection information for identifying the selected candidate MVs into the stream. That is, the inter prediction unit 126 may output the MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0364] Next, the inter prediction unit 126 generates a first prediction image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Similarly, the inter prediction unit 126 generates a second prediction image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).

[0365] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing weighted addition of the first predicted image and the second predicted image (step Sx_6).

[0366] In the example shown in FIG. 52A, the first partition and the second partition are each triangular, but they may be trapezoidal, or they may have different shapes from each other. Further, in the example shown in FIG. 52A, the current block is composed of two partitions, but it may be composed of three or more partitions.

[0367] Also, the first partition and the second partition may overlap. That is, the first partition and the second partition may include the same pixel region. In this case, a predicted image of the current block may be generated using the predicted image in the first partition and the predicted image in the second partition.

[0368] Also, in this example, an example in which predicted images are generated by inter prediction for both of the two partitions is shown, but predicted images may be generated by intra prediction for at least one partition.

[0369] Note that the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different or may be the same candidate MV list.

[0370] Note that the partition information may include at least an index indicating a division direction for dividing the current block into a plurality of partitions. The MV selection information may include an index indicating the selected first MV and an index indicating the selected second MV. One index may indicate a plurality of pieces of information. For example, one index that collectively indicates a part or all of the partition information and a part or all of the MV selection information may be encoded.

[0371] [MV Derivation > ATMVP Mode] FIG. 54 is a diagram showing an example of the ATMVP mode in which MVs are derived in sub-block units.

[0372] The ATMVP mode is a mode classified into the merge mode. For example, in the ATMVP mode, candidate MVs in sub-block units are registered in the candidate MV list used for the normal merge mode.

[0373] Specifically, in the ATMVP mode, first, as shown in FIG. 54, in the encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, the temporal MV reference block associated with the current block is specified. Next, for each sub-block in the current block, the MV used at the time of encoding the region corresponding to the sub-block in the temporal MV reference block is specified. The MV thus specified is included in the candidate MV list as a candidate MV for the sub-block of the current block. When the candidate MVs of such each sub-block are selected from the candidate MV list, motion compensation using the candidate MV as the MV of the sub-block is executed for the sub-block. Thereby, a predicted image of each sub-block is generated.

[0374] Note that, in the example shown in FIG. 54, a block adjacent to the lower left of the current block is used as the peripheral MV reference block, but other blocks may be used. Also, the size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may be switched in units such as slices, blocks, or pictures.

[0375] [Motion Search > DMVR] FIG. 55 is a diagram showing the relationship between the merge mode and DMVR.

[0376] The inter prediction unit 126 derives the MV of the current block in merge mode (step Sl_1). Next, the inter prediction unit 126 determines whether to perform MV search, that is, motion search (step Sl_2). Here, if the inter prediction unit 126 determines not to perform motion search (No in step Sl_2), it determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in merge mode.

[0377] On the other hand, if it is determined in step Sl_1 to perform motion search (Yes in step Sl_2), the inter prediction unit 126 derives the final MV for the current block by searching the peripheral region of the reference picture indicated by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by DMVR.

[0378] FIG. 56 is a conceptual diagram for explaining an example of DMVR for determining an MV.

[0379] First, for example, in merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are specified from the first reference picture (L0) which is the encoded picture in the L0 list. Similarly, according to the candidate MV (L1), reference pixels are specified from the second reference picture (L1) which is the encoded picture in the L1 list. A template is generated by taking the average of these reference pixels.

[0380] Next, using the template, the peripheral regions of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV of the current block. Note that the cost may be calculated using, for example, the difference value between each pixel value of the template and each pixel value of the search region, and the candidate MV value, etc.

[0381] Any process may be used as long as it can search the vicinity of the candidate MV to derive the final MV, even if it is not the process described herein.

[0382] FIG. 57 is a conceptual diagram for explaining another example of the DMVR for determining an MV. In this example shown in FIG. 57, unlike the example of the DMVR shown in FIG. 56, the cost is calculated without generating a template.

[0383] First, the inter prediction unit 126 searches the vicinity of the reference blocks included in the respective reference pictures of the L0 list and the L1 list based on the initial MV, which is the candidate MV obtained from the candidate MV list. For example, as shown in FIG. 57, the initial MV corresponding to the reference block of the L0 list is InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is InitMV_L1. In motion search, the inter prediction unit 126 first sets the search position for the reference picture of the L0 list. The difference vector indicating the set search position, specifically, the difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to the search position is MVd_L0. Then, the inter prediction unit 126 determines the search position in the reference picture of the L1 list. This search position is indicated by the difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. Specifically, the inter prediction unit 126 determines the difference vector as MVd_L1 by mirroring MVd_L0. That is, the inter prediction unit 126 sets the position symmetric to the position indicated by the initial MV as the search position in each of the reference pictures of the L0 list and the L1 list. For each search position, the inter prediction unit 126 calculates the sum of the absolute differences (SAD) of the pixel values within the block at the search position as the cost, and finds the search position where the cost is minimized.

[0384] FIG. 58A is a diagram showing an example of motion search in the DMVR, and FIG. 58B is a flowchart showing an example of the motion search.

[0385] First, in Step1, the inter prediction unit 126 calculates the costs at the search position indicated by the initial MV (also referred to as the starting point) and at eight search positions around it. Then, the inter prediction unit 126 determines whether the cost of a search position other than the starting point is the minimum. Here, when the inter prediction unit 126 determines that the cost of a search position other than the starting point is the minimum, it moves to the search position where the cost is the minimum and performs the processing of Step2. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit 126 skips the processing of Step2 and performs the processing of Step3.

[0386] In Step2, the inter prediction unit 126 performs the same search as in the processing of Step1 with the search position moved according to the processing result of Step1 as a new starting point. Then, the inter prediction unit 126 determines whether the cost of a search position other than the starting point is the minimum. Here, if the cost of a search position other than the starting point is the minimum, the inter prediction unit 126 performs the processing of Step4. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit 126 performs the processing of Step3.

[0387] In Step4, the inter prediction unit 126 treats the search position of the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0388] In Step3, the inter prediction unit 126 determines the pixel position with the minimum cost and decimal precision based on the costs at four points above, below, left, and right of the starting point in Step1 or Step2, and sets the pixel position as the final search position. The pixel position with decimal precision is determined by weighted addition of the vectors at the four points above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)) with the costs at the respective search positions of the four points as weights. Then, the inter prediction unit 126 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0389] [Motion Compensation > BIO / OBMC / LIC] In motion compensation, there are modes for generating a predicted image and correcting the predicted image. Such modes are, for example, BIO, OBMC, and LIC described below.

[0390] FIG. 59 is a flowchart showing an example of generating a predicted image.

[0391] The inter prediction unit 126 generates a predicted image (step Sm_1) and corrects the predicted image according to any of the above-described modes (step Sm_2).

[0392] FIG. 60 is a flowchart showing another example of generating a predicted image.

[0393] The inter prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter prediction unit 126 generates a predicted image using the MV (step Sn_2) and determines whether to perform correction processing (step Sn_3). Here, when the inter prediction unit 126 determines to perform correction processing (Yes in step Sn_3), it generates a final predicted image by correcting the predicted image (step Sn_4). Note that in LIC described below, in step Sn_4, the luminance and color difference may be corrected. On the other hand, when the inter prediction unit 126 determines not to perform correction processing (No in step Sn_3), it outputs the predicted image as the final predicted image without correction (step Sn_5).

[0394] [Motion Compensation > OBMC] Not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks may be used to generate an inter-predicted image. Specifically, a predicted image based on the motion information obtained by motion search (in the reference picture) and a predicted image based on the motion information of adjacent blocks (in the current picture) may be weighted and added to generate an inter-predicted image in units of sub-blocks within the current block. Such inter-prediction (motion compensation) is sometimes called OBMC (overlapped block motion compensation) or the OBMC mode.

[0395] In the OBMC mode, information indicating the size of sub-blocks for OBMC (for example, called the OBMC block size) may be signaled at the sequence level. Further, information indicating whether to apply the OBMC mode (for example, called the OBMC flag) may be signaled at the CU level. Note that the signaling levels of these pieces of information do not have to be limited to the sequence level and the CU level, and may be other levels (for example, the picture level, the slice level, the block level, the CTU level, or the sub-block level).

[0396] The OBMC mode will be described in more detail. FIGS. 61 and 62 are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process by OBMC.

[0397] First, as shown in FIG. 62, a predicted image (Pred) by normal motion compensation is obtained using the MV assigned to the current block. In FIG. 62, the arrow “MV” points to the reference picture and indicates what the current block of the current picture is referring to in order to obtain the predicted image.

[0398] Next, the already-derived MV (MV_L) for the encoded left adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is indicated by the arrow "MV_L" pointing to the reference picture from the current block. Then, the first correction of the predicted image is performed by superimposing the two predicted images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0399] Similarly, the already-derived MV (MV_U) for the encoded upper adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_U). The MV (MV_U) is indicated by the arrow "MV_U" pointing to the reference picture from the current block. Then, the second correction of the predicted image is performed by superimposing the predicted image Pred_U on the predicted image (e.g., Pred and Pred_L) that has undergone the first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block with the boundaries mixed (smoothed) with adjacent blocks.

[0400] Note that the above example is a two-pass correction method using left and upper adjacent blocks, but the correction method may also be a three-pass or more-pass correction method using right and / or lower adjacent blocks.

[0401] Note that the region for superimposition may be only a partial region near the block boundary, rather than the entire pixel region of the block.

[0402] Here, the predictive image correction process of OBMC for obtaining one predictive image Pred by superimposing additional predictive images Pred_L and Pred_U from one reference picture has been described. However, when the predictive image is corrected based on a plurality of reference images, the same process may be applied to each of the plurality of reference pictures. In such a case, by performing the image correction of OBMC based on a plurality of reference pictures, after obtaining the corrected predictive images from each reference picture, the final predictive image is obtained by further superimposing the plurality of obtained corrected predictive images.

[0403] Note that in OBMC, the unit of the current block may be a PU unit or a sub-block unit obtained by further dividing the PU.

[0404] As a method for determining whether to apply OBMC, for example, there is a method using obmc_flag, which is a signal indicating whether to apply OBMC. As a specific example, the encoding device 100 may determine whether the current block belongs to a region with complex motion. When the current block belongs to a region with complex motion, the encoding device 100 sets the value 1 as obmc_flag and applies OBMC for encoding. When the current block does not belong to a region with complex motion, the encoding device 100 sets the value 0 as obmc_flag and performs block encoding without applying OBMC. On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC according to the value.

[0405] [Motion Compensation > BIO] Next, the method for deriving the MV will be described. First, the mode for deriving the MV based on a model assuming uniform linear motion will be described. This mode is sometimes called the BIO (bi - directional optical flow) mode. Also, this bi - directional optical flow may be denoted as BDOF instead of BIO.

[0406] FIG. 63 is a diagram for explaining a model assuming a uniform linear motion. In FIG. 63, (vx, vy) indicates a velocity vector, and τ0 and τ1 respectively indicate the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) indicates the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) indicates the MV corresponding to the reference picture Ref1.

[0407] At this time, under the assumption of the uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are respectively represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) holds.

[0408]

Equation

[0409] Here, I(k) indicates the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from the candidate MV list or the like may be corrected in pixel units.

[0410] Note that the MV may be derived on the decoder 200 side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the MVs of a plurality of adjacent blocks.

[0411] FIG. 64 is a flowchart showing an example of inter prediction according to BIO. Further, FIG. 65 is a diagram showing an example of the functional configuration of an inter prediction unit 126 that performs inter prediction according to the BIO.

[0412] As shown in FIG. 65, the inter prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a prediction image correction unit 126f. Note that the memory 126a may be the frame memory 122.

[0413] The inter prediction unit 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter prediction unit 126 derives a prediction image of the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0414] Next, the interpolation image derivation unit 126b refers to the memory 126a and derives an interpolation image I0 of the current block using the motion vector M0 and the reference picture L0. Further, the interpolation image derivation unit 126b refers to the memory 126a and derives an interpolation image I1 of the current block using the motion vector M1 and the reference picture L1 (step Sy_2). Here, the interpolation image I0 is an image included in the reference picture Ref0 derived for the current block, and the interpolation image I1 is an image included in the reference picture Ref1 derived for the current block. The interpolation image I0 and the interpolation image I1 may each have the same size as the current block. Alternatively, the interpolation image I0 and the interpolation image I1 may each be an image larger than the current block in order to appropriately derive a gradient image described later. Furthermore, the interpolation images I0 and I1 may include prediction images derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1) and a motion compensation filter.

[0415] Further, the gradient image derivation unit 126c derives the gradient images (Ix0, Ix1, Iy0, Iy1) of the current block from the interpolated image I0 and the interpolated image I1 (step Sy_3). Note that the gradient images in the horizontal direction are (Ix0, Ix1), and the gradient images in the vertical direction are (Iy0, Iy1). The gradient image derivation unit 126c may derive the gradient images, for example, by applying a gradient filter to the interpolated image. The gradient images may be those indicating the spatial change amount of pixel values along the horizontal or vertical direction.

[0416] Next, the optical flow derivation unit 126d derives the optical flow (vx, vy), which is the above-described velocity vector, using the interpolated images (I0, I1) and the gradient images (Ix0, Ix1, Iy0, Iy1) in units of a plurality of sub-blocks constituting the current block (step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and may be referred to as a local motion estimation value, a correction motion vector, or a correction weight vector. As an example, the sub-block may be a 4x4 pixel sub-CU. Note that the optical flow may be derived in other units such as pixel units instead of sub-block units.

[0417] Next, the inter prediction unit 126 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives correction values for the values of the pixels included in the current block using the optical flow (vx, vy) (step Sy_5). Then, the predicted image correction unit 126f may correct the predicted image of the current block using the correction values (step Sy_6). Note that the correction values may be derived in units of each pixel, or may be derived in units of a plurality of pixels or sub-blocks.

[0418] Note that the processing flow of BIO is not limited to the processing disclosed in FIG. 64. Only a part of the processing disclosed in FIG. 64 may be performed, different processing may be added or replaced, or the processing may be executed in a different order.

[0419] [Motion Compensation > LIC] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) will be described.

[0420] FIG. 66A is a diagram for explaining an example of a predicted image generation method using luminance correction processing by LIC. Further, FIG. 66B is a flowchart showing an example of the predicted image generation method using the LIC.

[0421] First, the inter prediction unit 126 derives an MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).

[0422] Next, the inter prediction unit 126 extracts information indicating how the luminance value has changed between the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference area (peripheral reference area) and the encoded upper adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV. Then, the inter prediction unit 126 calculates a luminance correction parameter using the information indicating how the luminance value has changed (step Sz_3).

[0423] The inter prediction unit 126 generates a predicted image for the current block by performing a luminance correction process of applying the luminance correction parameter to the reference image in the reference picture specified by the MV (step Sz_4). That is, correction based on the luminance correction parameter is performed on the predicted image, which is the reference image in the reference picture specified by the MV. In this correction, the luminance may be corrected, or the color difference may be corrected. That is, a color difference correction parameter may be calculated using information indicating how the color difference has changed, and a color difference correction process may be performed.

[0424] Note that the shape of the peripheral reference area in FIG. 66A is an example, and other shapes may be used.

[0425] Also, although the process of generating a predicted image from a single reference picture has been described here, the same applies when generating a predicted image from a plurality of reference pictures. For the reference images obtained from each reference picture, a luminance correction process may be performed in the same manner as described above, and then the predicted image may be generated.

[0426] As a method for determining whether to apply LIC, for example, there is a method of using lic_flag, which is a signal indicating whether to apply LIC. As a specific example, in the encoding device 100, it is determined whether the current block belongs to an area where a luminance change has occurred. If it belongs to an area where a luminance change has occurred, the value 1 is set as lic_flag and encoding is performed by applying LIC. If it does not belong to an area where a luminance change has occurred, the value 0 is set as lic_flag and encoding is performed without applying LIC. On the other hand, in the decoding device 200, by decoding the lic_flag described in the stream, decoding may be performed by switching whether to apply LIC according to the value.

[0427] As another method for determining whether to apply LIC, for example, there is also a method of determining according to whether LIC has been applied to surrounding blocks. As a specific example, when the current block is being processed in merge mode, the inter prediction unit 126 determines whether the surrounding encoded blocks selected when deriving the MV in the merge mode have been encoded by applying LIC. The inter prediction unit 126 switches whether to apply LIC according to the result and performs encoding. Note that, also in this example, the same process is applied to the process on the decoding device 200 side.

[0428] Although LIC (luminance correction process) has been described with reference to FIGS. 66A and 66B, the details will be described below.

[0429] First, the inter prediction unit 126 derives an MV for obtaining a reference image corresponding to the current block from a reference picture that is an encoded picture.

[0430] Next, the inter prediction unit 126 extracts information indicating how the luminance values change between the reference picture and the current picture by using the luminance pixel values of the left and upper adjacent coded peripheral reference regions for the current block and the luminance pixel values at the equivalent positions in the reference picture specified by the MV, and calculates the luminance correction parameters. For example, let the luminance pixel value of a certain pixel in the peripheral reference region within the current picture be p0, and the luminance pixel value of the pixel in the peripheral reference region within the reference picture at the equivalent position to this pixel be p1. The inter prediction unit 126 calculates, as the luminance correction parameters, the coefficients A and B that optimize A×p1 + B = p0 for a plurality of pixels in the peripheral reference region.

[0431] Next, the inter prediction unit 126 performs luminance correction processing on the reference image in the reference picture specified by the MV by using the luminance correction parameters, thereby generating a prediction image for the current block. For example, let the luminance pixel value in the reference image be p2, and the luminance pixel value of the prediction image after the luminance correction processing be p3. The inter prediction unit 126 generates the prediction image after the luminance correction processing by calculating A×p2 + B = p3 for each pixel in the reference image.

[0432] Note that a part of the peripheral reference region shown in FIG. 66A may be used. For example, a region including a predetermined number of pixels decimated from each of the upper adjacent pixel and the left adjacent pixel may be used as the peripheral reference region. Also, the peripheral reference region is not limited to the region adjacent to the current block, and may be a region not adjacent to the current block. Further, in the example shown in FIG. 66A, the peripheral reference region in the reference picture is the region specified by the MV of the current picture from the peripheral reference region in the current picture, but may be the region specified by another MV. For example, the other MV may be the MV of the peripheral reference region in the current picture.

[0433] Note that here, the operation in the encoding device 100 has been described, but the operation in the decoding device 200 is the same.

[0434] Note that the LIC may be applied not only to luminance but also to color difference. At this time, correction parameters may be derived individually for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.

[0435] Also, the LIC process may be applied in sub-block units. For example, correction parameters may be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0436] [Prediction control unit] The prediction control unit 128 selects either an intra-predicted image (an image or signal output from the intra-prediction unit 124) or an inter-predicted image (an image or signal output from the inter-prediction unit 126), and outputs the selected predicted image to the subtraction unit 104 and the addition unit 116.

[0437] [Prediction parameter generation unit] The prediction parameter generation unit 130 may output information regarding intra-prediction, inter-prediction, and selection of the predicted image in the prediction control unit 128 as prediction parameters to the entropy encoding unit 110. The entropy encoding unit 110 may generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantized coefficients input from the quantization unit 108. The prediction parameters may be used by the decoding device 200. The decoding device 200 may receive and decode the stream, and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters may include a selected prediction signal (e.g., an MV, a prediction type, or a prediction mode used in the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or indicating the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0438] [Decoding device] Next, a decoding device 200 capable of decoding the stream output from the above-described encoding device 100 will be described. FIG. 67 is a block diagram showing an example of the functional configuration of the decoding device 200 according to the embodiment. The decoding device 200 is a device that decodes a stream, which is an encoded image, in block units.

[0439] As shown in FIG. 67, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a division determination unit 224. Each of the intra prediction unit 216 and the inter prediction unit 218 is configured as a part of the prediction processing unit.

[0440] [Implementation Example of Decoding Device] FIG. 68 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, a plurality of components of the decoding device 200 shown in FIG. 67 are implemented by the processor b1 and the memory b2 shown in FIG. 68.

[0441] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 may be a processor such as a CPU. Also, the processor b1 may be an aggregate of a plurality of electronic circuits. Further, for example, the processor b1 may play the roles of a plurality of components of the decoding device 200 shown in FIG. 67 etc., excluding the components for storing information.

[0442] Memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode a stream is stored. Memory b2 may be an electronic circuit and may be connected to the processor b1. Also, memory b2 may be included in the processor b1. Further, memory b2 may be an aggregate of a plurality of electronic circuits. Also, memory b2 may be a magnetic disk, an optical disk, etc., or may be expressed as a storage or a recording medium, etc. Also, memory b2 may be a non-volatile memory or a volatile memory.

[0443] For example, an image or a stream may be stored in memory b2. Also, a program for the processor b1 to decode a stream may be stored in memory b2.

[0444] Also, for example, memory b2 may serve as a component for storing information among a plurality of components of the decoding device 200 shown in FIG. 67 etc. Specifically, memory b2 may serve as the block memory 210 and the frame memory 214 shown in FIG. 67. More specifically, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.) may be stored in memory b2.

[0445] Note that in the decoding device 200, not all of the plurality of components shown in FIG. 67 etc. need to be implemented, and not all of the plurality of processes described above need to be performed. A part of the plurality of components shown in FIG. 67 etc. may be included in another device, and a part of the plurality of processes described above may be executed by another device.

[0446] First, after explaining the overall processing flow of the decoding device 200, each component included in the decoding device 200 will be described. Among the components included in the decoding device 200, for those that perform the same processing as the components included in the encoding device 100, detailed descriptions will be omitted. For example, the inverse quantization unit 204, inverse transformation unit 206, addition unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 included in the decoding device 200 perform the same processing as the inverse quantization unit 112, inverse transformation unit 114, addition unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120 included in the encoding device 100, respectively.

[0447] [Overall Flow of Decoding Process] FIG. 69 is a flowchart showing an example of the overall decoding process by the decoding device 200.

[0448] First, the division determination unit 224 of the decoding device 200 determines the division pattern of each of a plurality of fixed-size blocks (128×128 pixels) included in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This division pattern is the division pattern selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 for each of the plurality of blocks constituting the division pattern.

[0449] The entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the current block (step Sp_2).

[0450] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore the prediction residual of the current block by performing inverse quantization and inverse transformation on the plurality of quantization coefficients (step Sp_3).

[0451] Next, the prediction processing unit including the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a predicted image of the current block (step Sp_4).

[0452] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the predicted image to the prediction residual (step Sp_5).

[0453] Then, when this reconstructed image is generated, the loop filter unit 212 performs filtering on the reconstructed image (step Sp_6).

[0454] Then, the decoding device 200 determines whether the decoding of the entire picture is completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeatedly executed.

[0455] Note that the processing of these steps Sp_1 to Sp_7 may be sequentially performed by the decoding device 200, or a plurality of some of these processes may be performed in parallel, or the order may be changed.

[0456] [Partition determination unit] FIG. 70 is a diagram showing the relationship between the partition determination unit 224 and other components. The partition determination unit 224 may perform the following processing as an example.

[0457] The splitting determination unit 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoding unit 202. Then, the splitting determination unit 224 may determine a splitting pattern of fixed-size blocks based on the block information and the parameters. Then, the splitting determination unit 224 may output information indicating the determined splitting pattern to the inverse transformation unit 206, the intra prediction unit 216, and the inter prediction unit 218. The inverse transformation unit 206 may perform inverse transformation on the transformation coefficients based on the splitting pattern indicated by the information from the splitting determination unit 224. The intra prediction unit 216 and the inter prediction unit 218 may generate a predicted image based on the splitting pattern indicated by the information from the splitting determination unit 224.

[0458] [Entropy decoding unit] FIG. 71 is a block diagram showing an example of the functional configuration of the entropy decoding unit 202.

[0459] The entropy decoding unit 202 generates quantization coefficients, prediction parameters, parameters related to the splitting pattern, etc. by entropy decoding the stream. For example, CABAC is used for the entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a multivalue conversion unit 202c. The binary arithmetic decoding unit 202a arithmetically decodes the stream into a binary signal using the context value derived by the context control unit 202b. The context control unit 202b derives a context value corresponding to the feature of the syntax element or the surrounding situation, that is, the generation probability of the binary signal, in the same manner as the context control unit 110b of the encoding apparatus 100. The multivalue conversion unit 202c performs multivalue conversion (debinarize) to convert the binary signal output from the binary arithmetic decoding unit 202a into a multivalue signal indicating the above-described quantization coefficients and the like. This multivalue conversion is performed according to the above-described binarization method.

[0460] The entropy decoding unit 202 outputs quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may output prediction parameters included in the stream (see FIG. 1) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can execute the same prediction processes as those performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoder 100 side.

[0461] [Entropy decoding unit] FIG. 72 is a diagram showing the flow of CABAC in the entropy decoding unit 202.

[0462] First, in CABAC in the entropy decoding unit 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding unit 202a and setting of initial context values are performed. Then, the binary arithmetic decoding unit 202a and the multivalue conversion unit 202c execute arithmetic decoding and multivalue conversion on, for example, the encoded data of the CTU. At this time, the context control unit 202b updates the context value every time arithmetic decoding is performed. Then, as post-processing, the context control unit 202b saves the context value. This saved context value is used, for example, as the initial value of the context value for the next CTU.

[0463] [Inverse quantization unit] The inverse quantization unit 204 inverse quantizes the quantization coefficients of the current block, which is the input from the entropy decoding unit 202. Specifically, for each quantization coefficient of the current block, the inverse quantization unit 204 inverse quantizes the quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0464] FIG. 73 is a block diagram showing an example of the functional configuration of the inverse quantization unit 204.

[0465] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204a, a predictive quantization parameter generation unit 204b, a quantization parameter storage unit 204d, and an inverse quantization processing unit 204e.

[0466] FIG. 74 is a flowchart showing an example of inverse quantization by the inverse quantization unit 204.

[0467] As an example, the inverse quantization unit 204 may perform inverse quantization processing for each CU based on the flow shown in FIG. 74. Specifically, the quantization parameter generation unit 204a determines whether to perform inverse quantization (step Sv_11). Here, if it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation unit 204a acquires the differential quantization parameter of the current block from the entropy decoding unit 202 (step Sv_12).

[0468] Next, the predictive quantization parameter generation unit 204b acquires the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). The predictive quantization parameter generation unit 204b generates the predictive quantization parameter of the current block based on the acquired quantization parameter (step Sv_14).

[0469] Then, the quantization parameter generation unit 204a adds the differential quantization parameter of the current block acquired from the entropy decoding unit 202 and the predictive quantization parameter of the current block generated by the predictive quantization parameter generation unit 204b (step Sv_15). By this addition, the quantization parameter of the current block is generated. Also, the quantization parameter generation unit 204a stores the quantization parameter of the current block in the quantization parameter storage unit 204d (step Sv_16).

[0470] Next, the inverse quantization processing unit 204e inverse quantizes the quantization coefficient of the current block into a conversion coefficient using the quantization parameter generated in step Sv_15 (step Sv_17).

[0471] Note that the differential quantization parameter may be decoded at the bit sequence level, picture level, slice level, block level, or CTU level. Also, the initial value of the quantization parameter may be decoded at the sequence level, picture level, slice level, block level, or CTU level. At this time, the quantization parameter may be generated using the initial value of the quantization parameter and the differential quantization parameter.

[0472] Note that the inverse quantization unit 204 may include a plurality of inverse quantizers, and the quantization coefficients may be inverse quantized using an inverse quantization method selected from a plurality of inverse quantization methods.

[0473] [Inverse conversion unit] The inverse conversion unit 206 restores the prediction residual by inverse-converting the conversion coefficients that are the input from the inverse quantization unit 204.

[0474] For example, when the information decoded from the stream indicates that EMT or AMT is to be applied (for example, the AMT flag is true), the inverse conversion unit 206 inverse-converts the conversion coefficients of the current block based on the information indicating the decoded conversion type.

[0475] Also, for example, when the information decoded from the stream indicates that NSST is to be applied, the inverse conversion unit 206 applies an inverse inverse-conversion to the conversion coefficients.

[0476] FIG. 75 is a flowchart showing an example of the processing by the inverse conversion unit 206.

[0477] For example, the inverse transform unit 206 determines whether information indicating that orthogonal transformation is not to be performed exists in the stream (step St_11). Here, if it is determined that the information does not exist (No in step St_11), the inverse transform unit 206 acquires information indicating the transform type, which has been decoded by the entropy decoding unit 202 (step St_12). Next, the inverse transform unit 206 determines the transform type used for the orthogonal transform of the encoding apparatus 100 based on the information (step St_13). Then, the inverse transform unit 206 performs inverse orthogonal transformation using the determined transform type (step St_14).

[0478] FIG. 76 is a flowchart showing another example of the processing by the inverse transform unit 206.

[0479] For example, the inverse transform unit 206 determines whether the transform size is equal to or less than a predetermined value (step Su_11). Here, if it is determined that the transform size is equal to or less than the predetermined value (Yes in step Su_11), the inverse transform unit 206 acquires from the entropy decoding unit 202 information indicating which one of one or more transform types included in the first transform type group has been used by the encoding apparatus 100 (step Su_12). Note that such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.

[0480] The inverse transform unit 206 determines the transform type used for the orthogonal transform in the encoding apparatus 100 based on the information (step Su_13). Then, the inverse transform unit 206 performs inverse orthogonal transformation on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, if the inverse transform unit 206 determines in step Su_11 that the transform size is not equal to or less than the predetermined value (No in step Su_11), the inverse transform unit 206 performs inverse orthogonal transformation on the transform coefficients of the current block using the second transform type group (step Su_15).

[0481] Note that the inverse orthogonal transformation by the inverse transformation unit 206 may be performed according to the flow shown in FIG. 75 or FIG. 76 for each TU as an example. Alternatively, the inverse orthogonal transformation may be performed using a predefined transformation type without decoding the information indicating the transformation type used for the orthogonal transformation. Specifically, the transformation type is, for example, DST7 or DCT8, and in the inverse orthogonal transformation, an inverse transformation basis function corresponding to the transformation type is used.

[0482] [Addition unit] The addition unit 208 reconstructs the current block by adding the prediction residual, which is the input from the inverse transformation unit 206, and the prediction image, which is the input from the prediction control unit 220. That is, a reconstructed image of the current block is generated. Then, the addition unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.

[0483] [Block memory] The block memory 210 is a storage unit for storing blocks within the current picture, which are blocks referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the addition unit 208.

[0484] [Loop filter unit] The loop filter unit 212 applies a loop filter to the reconstructed image generated by the addition unit 208 and outputs the filtered reconstructed image to the frame memory 214 and a display device or the like.

[0485] When the information indicating the on / off of the ALF read from the stream indicates that the ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.

[0486] FIG. 77 is a block diagram showing an example of the functional configuration of the loop filter unit 212. Note that the loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding device 100.

[0487] As shown in FIG. 77 for example, the loop filter unit 212 includes a deblocking filter processing unit 212a, a SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 212b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. Further, the ALF processing unit 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. Note that the loop filter unit 212 does not necessarily include all the processing units disclosed in FIG. 77, and may include only some of the processing units. Also, the loop filter unit 212 may be configured to perform the above-described respective processes in an order different from the processing order disclosed in FIG. 77.

[0488] [Frame memory] The frame memory 214 is a storage unit for storing reference pictures used for inter prediction, and may also be called a frame buffer. Specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter unit 212.

[0489] [Prediction unit (intra prediction unit, inter prediction unit, prediction control unit)] FIG. 78 is a flowchart showing an example of the processing performed in the prediction unit of the decoding apparatus 200. As an example, the prediction unit includes all or some of the components of an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220. The prediction processing unit includes, for example, the intra prediction unit 216 and the inter prediction unit 218.

[0490] The prediction unit generates a predicted image of the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a predicted block. Note that the prediction signal includes, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction unit generates a predicted image of the current block using a reconstructed image that has already been obtained by performing generation of a predicted image for other blocks, restoration of a prediction residual, and addition of the predicted image. The prediction unit of the decoding apparatus 200 generates the same predicted image as the predicted image generated by the prediction unit of the encoding apparatus 100. That is, the methods for generating the predicted images used in those prediction units are common or corresponding to each other.

[0491] The reconstructed image may be, for example, an image of a reference picture, or an image of a decoded block (i.e., the other blocks described above) in the current picture that is the picture including the current block. The decoded block in the current picture is, for example, an adjacent block of the current block.

[0492] FIG. 79 is a flowchart showing another example of the processing performed by the prediction unit of the decoding apparatus 200.

[0493] The prediction unit determines a method or mode for generating a predicted image (step Sr_1). For example, this method or mode may be determined based on, for example, prediction parameters.

[0494] When the prediction unit determines the first method as the mode for generating the predicted image, it generates the predicted image according to the first method (step Sr_2a). Also, when the prediction unit determines the second method as the mode for generating the predicted image, it generates the predicted image according to the second method (step Sr_2b). Also, when the prediction unit determines the third method as the mode for generating the predicted image, it generates the predicted image according to the third method (step Sr_2c).

[0495] The first method, the second method, and the third method are different methods for generating a predicted image, and may be, for example, an inter prediction method, an intra prediction method, and other prediction methods, respectively. In these prediction methods, the reconstructed image described above may be used.

[0496] FIGS. 80A and 80B are flowcharts showing other examples of the processing performed by the prediction unit of the decoding apparatus 200.

[0497] As an example, the prediction unit may perform prediction processing according to the flow shown in FIGS. 80A and 80B. Note that the intra block copy shown in FIGS. 80A and 80B is one mode belonging to inter prediction, and is a mode in which a block included in the current picture is referred to as a reference picture or a reference block. That is, in the intra block copy, a picture different from the current picture is not referred to. Also, the PCM mode shown in FIG. 80A is one mode belonging to intra prediction, and is a mode in which no conversion and quantization are performed.

[0498] [Intra Prediction Unit] The intra prediction unit 216 generates a predicted image (that is, an intra predicted image) of the current block by performing intra prediction with reference to a block in the current picture stored in the block memory 210 based on the intra prediction mode decoded from the stream. Specifically, the intra prediction unit 216 generates an intra predicted image by performing intra prediction with reference to the pixel values (for example, luminance values and color difference values) of the blocks adjacent to the current block, and outputs the intra predicted image to the prediction control unit 220.

[0499] Note that when an intra prediction mode in which a luminance block is referred to in the intra prediction of a color difference block is selected, the intra prediction unit 216 may predict the color difference component of the current block based on the luminance component of the current block.

[0500] Also, when the information decoded from the stream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0501] FIG. 81 is a diagram showing an example of the processing by the intra prediction unit 216 of the decoding apparatus 200.

[0502] The intra prediction unit 216 first determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, if it is determined that the MPM flag indicating 1 exists (Yes in step Sw_11), the intra prediction unit 216 acquires, from the entropy decoding unit 202, information indicating the intra prediction mode selected in the encoding apparatus 100 among the MPMs (step Sw_12). Note that the information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Next, the intra prediction unit 216 determines the MPM (step Sw_13). The MPM consists of, for example, six intra prediction modes. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information acquired in step Sw_12 from among the plurality of intra prediction modes included in the MPM (step Sw_14).

[0503] On the other hand, when the intra prediction unit 216 determines in step Sw_11 that the MPM flag indicating 1 does not exist in the stream (No in step Sw_11), it acquires information indicating the intra prediction mode selected in the encoding apparatus 100 (step Sw_15). That is, the intra prediction unit 216 acquires, from the entropy decoding unit 202, information indicating the intra prediction mode selected in the encoding apparatus 100 from among one or more intra prediction modes not included in the MPM. Note that the information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information acquired in step Sw_15 from among the one or more intra prediction modes not included in the MPM (step Sw_17).

[0504] The intra prediction unit 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).

[0505] [Inter prediction unit] The inter prediction unit 218 predicts the current block with reference to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks within the current block. Note that the sub-blocks are included in the block and are units smaller than the block. The size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may be switched in units such as slices, bricks, or pictures.

[0506] For example, the inter prediction unit 218 generates an inter prediction image of the current block or sub-block by performing motion compensation using motion information (e.g., MV) decoded from a stream (e.g., prediction parameters output from the entropy decoding unit 202), and outputs the inter prediction image to the prediction control unit 220.

[0507] When the information decoded from the stream indicates that the OBMC mode is to be applied, the inter prediction unit 218 generates an inter prediction image using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.

[0508] Also, when the information decoded from the stream indicates that the FRUC mode is to be applied, the inter prediction unit 218 derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the stream. Then, the inter prediction unit 218 performs motion compensation (prediction) using the derived motion information.

[0509] Further, when the BIO mode is applied, the inter prediction unit 218 derives an MV based on a model assuming a uniform linear motion. Further, when the information decoded from the stream indicates that the affine mode is to be applied, the inter prediction unit 218 derives an MV in sub-block units based on the MVs of a plurality of adjacent blocks.

[0510] [Flow of MV Derivation] FIG. 82 is a flowchart showing an example of MV derivation in the decoding apparatus 200.

[0511] The inter prediction unit 218 determines, for example, whether to decode motion information (e.g., MV). For example, the inter prediction unit 218 may determine according to the prediction mode included in the stream, or may determine based on other information included in the stream. Here, when the inter prediction unit 218 determines to decode the motion information, it derives the MV of the current block in the mode of decoding the motion information. On the other hand, when the inter prediction unit 218 determines not to decode the motion information, it derives the MV in the mode of not decoding the motion information.

[0512] Here, the modes of MV derivation include a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, etc., which will be described later. Among these modes, the modes of decoding motion information include a normal inter mode, a normal merge mode, and an affine mode (specifically, an affine inter mode and an affine merge mode), etc. Note that the motion information may include not only the MV but also prediction MV selection information, which will be described later. Further, the mode of not decoding the motion information includes a FRUC mode, etc. The inter prediction unit 218 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0513] FIG. 83 is a flowchart showing another example of MV derivation in the decoding apparatus 200.

[0514] The inter prediction unit 218 determines, for example, whether to decode the differential MV. For example, the inter prediction unit 218 may determine according to the prediction mode included in the stream, or may determine based on other information included in the stream. Here, when the inter prediction unit 218 determines to decode the differential MV, it may derive the MV of the current block in the mode of decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.

[0515] On the other hand, when the inter prediction unit 218 determines not to decode the differential MV, it derives the MV in the mode of not decoding the differential MV. In this case, the encoded differential MV is not included in the stream.

[0516] Here, as described above, the modes of deriving the MV include a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, etc. to be described later. Among these modes, the modes of encoding the differential MV include a normal inter mode and an affine mode (specifically, an affine inter mode), etc. Also, the modes of not encoding the differential MV include a FRUC mode, a normal merge mode, and an affine mode (specifically, an affine merge mode), etc. The inter prediction unit 218 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.

[0517] [MV Derivation > Normal Inter Mode] For example, when the information read from the stream indicates that the normal inter mode is to be applied, the inter prediction unit 218 derives the MV in the normal merge mode based on the information read from the stream, and performs motion compensation (prediction) using the MV.

[0518] FIG. 84 is a flowchart showing an example of inter prediction in the normal inter mode in the decoding apparatus 200.

[0519] The inter prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks around the current block temporally or spatially (step Sg_11). That is, the inter prediction unit 218 creates a candidate MV list.

[0520] Next, the inter prediction unit 218 extracts each of N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_11 as a prediction motion vector candidate (also referred to as a prediction MV candidate) according to a predetermined priority order (step Sg_12). The priority order is determined in advance for each of the N prediction MV candidates.

[0521] Next, the inter prediction unit 218 decodes prediction MV selection information from the input stream, and uses the decoded prediction MV selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction MV of the current block (step Sg_13).

[0522] Next, the inter prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the differential value, which is the decoded differential MV, to the selected prediction MV (step Sg_14).

[0523] Finally, the inter prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are executed for each block. For example, when the processes of steps Sg_11 to Sg_15 are executed for each of all the blocks included in a slice, the inter prediction using the normal inter mode for that slice ends. Also, when the processes of steps Sg_11 to Sg_15 are executed for each of all the blocks included in a picture, the inter prediction using the normal inter mode for that picture ends. Note that if the processes of steps Sg_11 to Sg_15 are not executed for all the blocks included in a slice but are executed for some of the blocks, the inter prediction using the normal inter mode for that slice may end. Similarly, if the processes of steps Sg_11 to Sg_15 are executed for some of the blocks included in a picture, the inter prediction using the normal inter mode for that picture may end.

[0524] [MV Derivation > Normal Merge Mode] For example, when the information decoded from the stream indicates the application of the normal merge mode, the inter prediction unit 218 derives an MV in the normal merge mode and performs motion compensation (prediction) using the MV.

[0525] FIG. 85 is a flowchart showing an example of inter prediction in the normal merge mode in the decoding apparatus 200.

[0526] First, the inter prediction unit 218 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sh_11). That is, the inter prediction unit 218 creates a candidate MV list.

[0527] Next, the inter prediction unit 218 derives the MV of the current block by selecting one candidate MV from among the plurality of candidate MVs acquired in step Sh_11 (step Sh_12). Specifically, the inter prediction unit 218 acquires, for example, MV selection information included as a prediction parameter in the stream, and selects the candidate MV identified by the MV selection information as the MV of the current block.

[0528] Finally, the inter prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processes of steps Sh_11 to Sh_13 are executed, for example, for each block. For example, when the processes of steps Sh_11 to Sh_13 are executed for each of all the blocks included in a slice, the inter prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_11 to Sh_13 are executed for each of all the blocks included in a picture, the inter prediction using the normal merge mode for that picture ends. Note that, if the processes of steps Sh_11 to Sh_13 are not executed for all the blocks included in a slice but are executed for some of the blocks, the inter prediction using the normal merge mode for that slice may end. Similarly, if the processes of steps Sh_11 to Sh_13 are executed for some of the blocks included in a picture, the inter prediction using the normal merge mode for that picture may end.

[0529] [MV Derivation > FRUC Mode] For example, when the information decoded from the stream indicates the application of the FRUC mode, the inter prediction unit 218 derives the MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoder 100 side but is derived on the decoder 200 side. For example, the decoder 200 may derive the motion information by performing motion search. In this case, the decoder 200 performs the motion search without using the pixel values of the current block.

[0530] FIG. 86 is a flowchart showing an example of inter prediction in the FRUC mode in the decoder 200.

[0531] First, the inter prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block, and generates a list showing those MVs as candidate MVs (that is, a candidate MV list, which may be common to the candidate MV list in the normal merge mode) (step Si_11). Next, the inter prediction unit 218 selects the best candidate MV from among the plurality of candidate MVs registered in the candidate MV list (step Si_12). For example, the inter prediction unit 218 calculates the evaluation value of each candidate MV included in the candidate MV list, and selects one candidate MV as the best candidate MV based on the evaluation value. Then, the inter prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Also, for example, in the peripheral region of the position in the reference picture corresponding to the selected best candidate MV, the MV for the current block may be derived by performing pattern matching. That is, pattern matching and search using the evaluation value are performed on the region around the best candidate MV in the reference picture, and if there is an MV with a better evaluation value, the best candidate MV may be updated to that MV and used as the final MV for the current block. It is not necessary to perform the update to an MV with a better evaluation value.

[0532] Finally, the inter prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processes of steps Si_11 to Si_15 are executed, for example, for each block. For example, when the processes of steps Si_11 to Si_15 are executed for each of all the blocks included in a slice, the inter prediction using the FRUC mode for that slice is completed. Also, when the processes of steps Si_11 to Si_15 are executed for each of all the blocks included in a picture, the inter prediction using the FRUC mode for that picture is completed. The processing may be performed in the same manner as the above-described block unit even in units of sub-blocks.

[0533] [MV Derivation > Affine Merge Mode] For example, when the information decoded from the stream indicates the application of the affine merge mode, the inter prediction unit 218 derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.

[0534] FIG. 87 is a flowchart showing an example of inter prediction by the affine merge mode in the decoding apparatus 200.

[0535] In the affine merge mode, first, the inter prediction unit 218 derives each MV of the control points of the current block (step Sk_11). As shown in FIG. 46A, the control points are the upper left and upper right points of the current block, or as shown in FIG. 46B, the upper left, upper right, and lower left points of the current block.

[0536] For example, when using the MV derivation method shown in FIGS. 47A to 47C, the inter prediction unit 218 inspects these blocks in the order of the decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left) as shown in FIG. 47A, and identifies the first valid block decoded in the affine mode.

[0537] The inter prediction unit 218 derives the MV of the control point using the first valid block decoded in the specified affine mode. For example, if block A is specified and block A has two control points, as shown in FIG. 47B, the inter prediction unit 218 projects the motion vectors v3 and v4 at the upper left and upper right corners of the decoded block including block A onto the current block, thereby calculating the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block. Thereby, the MV of each control point is derived.

[0538] Note that, as shown in FIG. 49A, when block A is specified and block A has two control points, the MVs of three control points may be calculated. As shown in FIG. 49B, when block A is specified and block A has three control points, the MVs of two control points may be calculated.

[0539] Also, when the MV selection information is included in the stream as a prediction parameter, the inter prediction unit 218 may derive the MV of each control point of the current block using the MV selection information.

[0540] Next, the inter prediction unit 218 performs motion compensation for each of the plurality of sub-blocks included in the current block. That is, for each of the plurality of sub-blocks, the inter prediction unit 218 calculates the MV of the sub-block as an affine MV using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_12). Then, the inter prediction unit 218 performs motion compensation for the sub-block using those affine MVs and the decoded reference picture (step Sk_13). When the processes of steps Sk_12 and Sk_13 are executed for each of all the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block ends. That is, motion compensation is performed for the current block, and a predicted picture of the current block is generated.

[0541] Note that in step Sk_11, the above candidate MV list may be generated. The candidate MV list may be, for example, a list including candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47A to 47C, the MV derivation methods shown in FIGS. 48A and 48B, the MV derivation methods shown in FIGS. 49A and 49B, and other MV derivation methods.

[0542] Note that the candidate MV list may include candidate MVs for a mode that performs prediction in units of sub-blocks other than the affine mode.

[0543] Note that, as a candidate MV list, for example, a candidate MV list including a candidate MV in an affine merge mode having two control points and a candidate MV in an affine merge mode having three control points may be generated. Or, a candidate MV list including a candidate MV in an affine merge mode having two control points and a candidate MV list including a candidate MV in an affine merge mode having three control points may be generated respectively. Or, a candidate MV list including candidate MVs in one of the modes of an affine merge mode having two control points and an affine merge mode having three control points may be generated.

[0544] [MV Derivation > Affine Inter Mode] For example, when the information decoded from the stream indicates the application of the affine inter mode, the inter prediction unit 218 derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.

[0545] FIG. 88 is a flowchart showing an example of inter prediction by the affine inter mode in the decoding apparatus 200.

[0546] In the affine inter mode, first, the inter prediction unit 218 derives respective predicted MVs (v0, v1) or (v0, v1, v2) of two or three control points of the current block (step Sj_11). The control points are, for example, the points at the upper left corner, upper right corner, or lower left corner of the current block as shown in FIG. 46A or FIG. 46B.

[0547] The inter prediction unit 218 obtains prediction MV selection information included as a prediction parameter in the stream, and derives the prediction MV for each control point of the current block using the MV identified by the prediction MV selection information. For example, when using the MV derivation method shown in FIGS. 48A and 48B, the inter prediction unit 218 selects the MV of the block identified by the prediction MV selection information from among the decoded blocks in the vicinity of each control point of the current block shown in FIGS. 48A or 48B, thereby deriving the prediction MV (v0, v1) or (v0, v1, v2) of the control point of the current block.

[0548] Next, the inter prediction unit 218 obtains each differential MV included as a prediction parameter in the stream, for example, and adds the prediction MV of each control point of the current block and the differential MV corresponding to the prediction MV (step Sj_12). Thereby, the MV of each control point of the current block is derived.

[0549] Next, the inter prediction unit 218 performs motion compensation for each of a plurality of sub-blocks included in the current block. That is, the inter prediction unit 218 calculates the MV of each of the plurality of sub-blocks as an affine MV using two motion vectors v0 and v1 and the above-described formula (1A), or using three motion vectors v0, v1, and v2 and the above-described formula (1B) (step Sj_13). Then, the inter prediction unit 218 performs motion compensation on the sub-block using those affine MVs and the decoded reference picture (step Sj_14). When the processes of steps Sj_13 and Sj_14 are executed for each of all the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0550] Note that in step Sj_11, similar to step Sk_11, the above-described candidate MV list may be generated.

[0551] [MV Derivation > Triangle Mode] For example, when the information decoded from the stream indicates the application of the triangle mode, the inter prediction unit 218 derives the MV in the triangle mode and performs motion compensation (prediction) using the MV.

[0552] FIG. 89 is a flowchart showing an example of inter prediction by the triangle mode in the decoding apparatus 200.

[0553] In the triangle mode, first, the inter prediction unit 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter prediction unit 218 may obtain, from the stream, as prediction parameters, partition information which is information regarding the division into each partition. Then, the inter prediction unit 218 may divide the current block into the first partition and the second partition according to the partition information.

[0554] Next, the inter prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks around the current block temporally or spatially (step Sx_12). That is, the inter prediction unit 218 creates a candidate MV list.

[0555] Then, the inter prediction unit 218 selects, as a first MV and a second MV, the candidate MV for the first partition and the candidate MV for the second partition, respectively, from among the plurality of candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter prediction unit 218 may obtain, from the stream, as prediction parameters, MV selection information for identifying the selected candidate MVs. Then, the inter prediction unit 218 may select the first MV and the second MV according to the MV selection information.

[0556] Next, the inter prediction unit 218 generates a first predicted image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter prediction unit 218 generates a second predicted image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).

[0557] Finally, the inter prediction unit 218 generates a predicted image of the current block by weighted addition of the first predicted image and the second predicted image (step Sx_16).

[0558] [Motion Search > DMVR] For example, when the information decoded from the stream indicates the application of DMVR, the inter prediction unit 218 performs motion search with DMVR.

[0559] FIG. 90 is a flowchart showing an example of motion search by DMVR in the decoding apparatus 200.

[0560] The inter prediction unit 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter prediction unit 218 derives the final MV for the current block by searching the peripheral region of the reference picture indicated by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is determined by DMVR.

[0561] FIG. 91 is a flowchart showing a detailed example of motion search by DMVR in the decoding apparatus 200.

[0562] First, in Step 1 shown in FIG. 58A, the inter prediction unit 218 calculates the costs at the search position indicated by the initial MV (also referred to as the starting point) and at eight search positions around it. Then, the inter prediction unit 218 determines whether the cost of a search position other than the starting point is the minimum. Here, when the inter prediction unit 218 determines that the cost of a search position other than the starting point is the minimum, it moves to the search position where the cost is the minimum and performs the processing of Step 2 shown in FIG. 58A. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit 218 skips the processing of Step 2 shown in FIG. 58A and performs the processing of Step 3.

[0563] In Step 2 shown in FIG. 58A, the inter prediction unit 218 performs the same search as in the processing of Step 1, using the search position moved according to the processing result of Step 1 as a new starting point. Then, the inter prediction unit 218 determines whether the cost of a search position other than the starting point is the minimum. Here, if the cost of a search position other than the starting point is the minimum, the inter prediction unit 218 performs the processing of Step 4. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit 218 performs the processing of Step 3.

[0564] In Step 4, the inter prediction unit 218 treats the search position of the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0565] In Step 3 shown in FIG. 58A, the inter prediction unit 218 determines the pixel position with the minimum cost and fractional precision based on the costs at four points above, below, left, and right of the starting point of Step 1 or Step 2, and sets the pixel position as the final search position. The pixel position with fractional precision is determined by weighted addition of the vectors of the four points above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)), with the costs at the respective search positions of the four points as weights. Then, the inter prediction unit 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0566] [Motion Compensation > BIO / OBMC / LIC] For example, when the information decoded from the stream indicates the application of prediction image correction, the inter prediction unit 218 corrects the prediction image according to the correction mode when generating the prediction image. The mode is, for example, the above-mentioned BIO, OBMC, and LIC, etc.

[0567] FIG. 92 is a flowchart showing an example of generating a prediction image in the decoding apparatus 200.

[0568] The inter prediction unit 218 generates a prediction image (step Sm_11), and corrects the prediction image according to any of the above modes (step Sm_12).

[0569] FIG. 93 is a flowchart showing another example of generating a prediction image in the decoding apparatus 200.

[0570] The inter prediction unit 218 derives the MV of the current block (step Sn_11). Next, the inter prediction unit 218 generates a prediction image using the MV (step Sn_12), and determines whether to perform correction processing (step Sn_13). For example, the inter prediction unit 218 acquires the prediction parameters included in the stream, and determines whether to perform correction processing based on the prediction parameters. This prediction parameter is, for example, a flag indicating whether to apply each of the above modes. Here, when the inter prediction unit 218 determines to perform correction processing (Yes in step Sn_13), it generates the final prediction image by correcting the prediction image (step Sn_14). Note that in LIC, in step Sn_14, the luminance and color difference of the prediction image may be corrected. On the other hand, when the inter prediction unit 218 determines not to perform correction processing (No in step Sn_13), it outputs the prediction image as the final prediction image without correction (step Sn_15).

[0571] [Motion Compensation > OBMC] For example, when the information decoded from the stream indicates the application of OBMC, the inter prediction unit 218 corrects the predicted image according to OBMC when generating the predicted image.

[0572] FIG. 94 is a flowchart showing an example of correcting a predicted image by OBMC in the decoding apparatus 200. Note that the flowchart of FIG. 94 shows the flow of correcting the predicted image using the current picture and reference pictures shown in FIG. 62.

[0573] First, as shown in FIG. 62, the inter prediction unit 218 obtains a predicted image (Pred) by normal motion compensation using the MV assigned to the current block.

[0574] Next, the inter prediction unit 218 applies (reuses) the MV (MV_L) already derived for the decoded left adjacent block to the current block to obtain a predicted image (Pred_L). Then, the inter prediction unit 218 performs the first correction of the predicted image by overlapping the two predicted images Pred and Pred_L. This has the effect of mixing the boundaries between adjacent blocks.

[0575] Similarly, the inter prediction unit 218 applies (reuses) the MV (MV_U) already derived for the decoded upper adjacent block to the current block to obtain a predicted image (Pred_U). Then, the inter prediction unit 218 performs the second correction of the predicted image by overlapping the predicted image Pred_U with the predicted image (for example, Pred and Pred_L) that has undergone the first correction. This has the effect of mixing the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block in which the boundaries with adjacent blocks are mixed (smoothed).

[0576] [Motion Compensation > BIO] For example, when the information decoded from the stream indicates the application of BIO, the inter prediction unit 218 corrects the predicted image according to BIO when generating the predicted image.

[0577] FIG. 95 is a flowchart showing an example of correction of a predicted image by BIO in the decoding apparatus 200.

[0578] As shown in FIG. 63, the inter prediction unit 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter prediction unit 218 derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0579] Next, the inter prediction unit 218 derives an interpolated image I0 of the current block using the motion vector M0 and the reference picture L0. Also, the inter prediction unit 218 derives an interpolated image I1 of the current block using the motion vector M1 and the reference picture L1 (step Sy_12). Here, the interpolated image I0 is an image included in the reference picture Ref0 derived for the current block, and the interpolated image I1 is an image included in the reference picture Ref1 derived for the current block. The interpolated image I0 and the interpolated image I1 may each have the same size as the current block. Alternatively, the interpolated image I0 and the interpolated image I1 may each be an image larger than the current block in order to appropriately derive a gradient image described later. Further, the interpolated images I0 and I1 may include predicted images derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1) and a motion compensation filter.

[0580] Further, the inter prediction unit 218 derives the gradient images (Ix0, Ix1, Iy0, Iy1) of the current block from the interpolation images I0 and I1 (step Sy_13). Note that the gradient images in the horizontal direction are (Ix0, Ix1), and the gradient images in the vertical direction are (Iy0, Iy1). The inter prediction unit 218 may derive the gradient images, for example, by applying a gradient filter to the interpolation images. The gradient images may be those indicating the spatial change amount of pixel values along the horizontal or vertical direction.

[0581] Next, the inter prediction unit 218 derives the optical flow (vx, vy), which is the above-described velocity vector, using the interpolation images (I0, I1) and the gradient images (Ix0, Ix1, Iy0, Iy1) in units of a plurality of sub-blocks constituting the current block (step Sy_14). As an example, the sub-blocks may be 4x4 pixel sub-CUs.

[0582] Next, the inter prediction unit 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter prediction unit 218 derives a correction value for the value of the pixel included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter prediction unit 218 may correct the predicted image of the current block using the correction value (step Sy_16). Note that the correction values may be derived for each pixel, or may be derived for a plurality of pixel units or sub-block units.

[0583] Note that the processing flow of BIO is not limited to the processing disclosed in FIG. 95. Only a part of the processing disclosed in FIG. 95 may be performed, different processing may be added or replaced, or the processing may be executed in a different order.

[0584] [Motion Compensation > LIC] For example, when the information decoded from the stream indicates the application of LIC, the inter prediction unit 218 corrects the predicted image according to LIC when generating the predicted image.

[0585] FIG. 96 is a flowchart showing an example of correction of a predicted image by LIC in the decoding apparatus 200.

[0586] First, the inter prediction unit 218 acquires a reference image corresponding to the current block from the decoded reference picture using the MV (step Sz_11).

[0587] Next, the inter prediction unit 218 extracts information indicating how the luminance value has changed between the reference picture and the current picture for the current block (step Sz_12). This extraction is performed based on the luminance pixel values of the decoded left adjacent reference area (peripheral reference area) and the decoded upper adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV, as shown in FIG. 66A. Then, the inter prediction unit 218 calculates a luminance correction parameter using the information indicating how the luminance value has changed (step Sz_13).

[0588] The inter prediction unit 218 generates a predicted image for the current block by performing a luminance correction process of applying the luminance correction parameter to the reference image in the reference picture specified by the MV (step Sz_14). That is, correction based on the luminance correction parameter is performed on the predicted image, which is the reference image in the reference picture specified by the MV. In this correction, the luminance may be corrected or the color difference may be corrected.

[0589] [Prediction control unit] The prediction control unit 220 selects either the intra prediction image or the inter prediction image, and outputs the selected prediction image to the addition unit 208. Overall, the configuration, function, and processing of the prediction control unit 220, the intra prediction unit 216, and the inter prediction unit 218 on the decoding apparatus 200 side may correspond to the configuration, function, and processing of the prediction control unit 128, the intra prediction unit 124, and the inter prediction unit 126 on the encoding apparatus 100 side.

[0590] [Regarding multi-layer and single-layer] For the encoding of moving images, there is scalable encoding. In scalable encoding, the encoded bitstream is encoded so as to have a scalability function. The encoded bitstream having the scalability function has a stream structure composed of multiple layers. On the other hand, the encoded bitstream having no scalability function has a stream structure composed of a single layer. As the scalability function, there are a temporal layer scalability function, that is, a case where the Temporal Sub-Layer shown in FIGS. 4 and 5 is used as multiple layers, and a spatial / SN / multi-viewpoint scalability function, that is, a case where a Layer is used as multiple layers. In the following first aspect and second aspect, when the encoded bitstream has a scalability function, it will be described as having a spatial / SN / multi-viewpoint scalability function. The spatial / SN / multi-viewpoint scalability function is realized by a picture (PU) in an access unit (AU). Note that an access unit is a collection of data corresponding to one picture, and is a section obtained by grouping several NAL units in order to access the information in the encoded bitstream in picture units.

[0591] A stream structure composed of multiple layers, that is, two or more layers, is called a multi-layer structure. Hereinafter, the multi-layer structure used in the following first aspect and second aspect will be described with reference to FIG. 97.

[0592] FIG. 97 is a diagram for explaining the concept of a stream structure composed of multiple layers according to the first aspect and the like. Here, an example in the case of a stream structure composed of three layers is shown as the stream structure of the encoded bitstream having the spatial / SN / multi-viewpoint scalability function. Layer0 is also called the base layer, and Layer1 to Layer2 are also called the enhancement layers.

[0593] Specifically, in the example of FIG. 97, each of the four access units (AUs) denoted by AU0 to AU3 is a group of three pictures (PUs) to which the same display timing (POC: Picture Order Count) is assigned. To each of these three pictures (PUs), one picture (PU) is assigned for each of the three layers of Layer0 to Layer3. That is, the access unit (AU) of this aspect is a group of a plurality of PUs to which the same display timing (POC) is assigned, and is a group of a plurality of PUs to which one PU is assigned for each of the plurality of layers.

[0594] Also, the PU at the tip of the arrow shown in FIG. 97 represents a reference destination PU that is referred to when the PU at the start point of the arrow is encoded or decoded.

[0595] For example, in Layer0, the PU of POC_3 refers only to the PU of POC_2, the PU of POC_2 refers only to the PU of POC_1, and the PU of POC_1 refers only to the PU of POC_0. That is, in Layer0, since each PU refers only to the PU of the same Layer0, decoding is possible if there is only the stream data of Layer0.

[0596] On the other hand, for example, in Layer1, the PU of POC_3 refers to the PU of POC_3 in Layer0 in addition to the PU of POC_2 in the same layer, and the PU of POC_2 refers to the PU of POC_2 in Layer0 in addition to the PU of POC_1 in the same layer. Similarly, the PU of POC_1 in Layer1 refers to the PU of POC_1 in Layer0 in addition to the PU of POC_0 in the same layer. That is, since each PU in Layer1 refers to the PU in Layer0 in addition to the PU in the same Layer1, both the stream data of Layer0 and Layer1 are required to decode each PU in Layer1. Similarly, all the stream data from Layer0 to Layer2 are required to decode each PU in Layer2.

[0597] In addition, the set of layers that require decoding when decoding and displaying a specific layer is called OLS (Output Layer Set). For example, the OLS required to decode each PU in Layer1 is Layer0 and Layer1, and the OLS required to decode each PU in Layer2 is Layer0 to Layer2.

[0598] Here, for example, assume that Layer0 shown in FIG. 97 is composed of PUs with low resolution, Layer1 is composed of PUs with medium resolution, and Layer2 is composed of PUs with high resolution.

[0599] When it is desired to display an encoded bitstream having a stream structure composed of such three layers on a decoding device 200 with low processing performance or when it is desired to transmit data with a narrow transmission bandwidth, the encoding device 100 may send only the stream data of Layer0 to the decoding device 200. As a result, although the decoding device 200 will decode with a rough image commensurate with the processing performance or the transmission bandwidth, the encoded bitstream can be decoded.

[0600] On the other hand, assume that an encoded bitstream having a stream structure composed of such three layers can be displayed on a decoding device 200 with sufficiently high processing performance or data can be transmitted with a wide transmission bandwidth. In this case, the encoding device 100 may send all the stream data from Layer0 to Layer2 to the decoding device 200. As a result, the decoding device 200 can decode the encoded bitstream with a fine image.

[0601] Note that a stream structure composed of one layer is also called a single-layer structure. Using the example shown in FIG. 97, the single-layer structure is composed only of the PUs belonging to Layer0 and has a structure in which one AU has only one PU.

[0602] [First Aspect] This aspect describes the Sequence Parameter Set (SPS) applied to an encoded video sequence in Versatile Video Coding (VVC). Here, as described above, the SPS is a type of header information that includes parameters used for the encoded video sequence. The SPS includes a set (syntax) of encoding parameters that a decoder 200 refers to in order to decode the encoded video sequence.

[0603] FIG. 98 is a diagram showing syntax elements as part of the set of encoding parameters included in the SPS.

[0604] In VVC, the SPS applied to an encoded video sequence can refer to a Video Parameter Set (VPS). Here, the VPS is a type of header information that includes various information about layers included in the encoded video sequence or the encoded bitstream.

[0605] In this aspect, when the encoding device 100 generates an encoded bitstream including a multi-layer structure in the encoded bitstream, for example, the encoding device 100 generates the encoded bitstream including the SPS that refers to the VPS. On the other hand, when the encoding device 100 generates an encoded bitstream without including a multi-layer structure in the encoded bitstream, for example, the encoding device 100 generates the encoded bitstream including the SPS that does not refer to the VPS. Although the VPS also existed in standards prior to VVC, it was not effectively utilized during decoding. In this aspect, when a multi-layer structure is included in the encoded bitstream, the VPS includes information used to decode each layer of the multi-layer.

[0606] More specifically, as the SPS that refers to the VPS, the ID of the SPS described with a value of 0 or 1 is used. That is, when the encoding device 100 generates an encoded bitstream including a multi-layer structure therein, as the SPS that refers to the VPS, the encoded bitstream is generated including the ID of the SPS described with a value of 1. On the other hand, when the encoding device 100 generates an encoded bitstream without including a multi-layer structure therein, as the SPS that does not refer to the VPS, the encoded bitstream is generated including the ID of the SPS described with a value of 1.

[0607] Also, as the ID of the SPS described with a value of 0 or 1, sps_video_parameter_set_id among the syntax elements shown in FIG. 98 can be used.

[0608] Note that when a multi-layer structure is included in the encoded bitstream, the encoded video sequence includes at least one NAL (Network Abstraction Layer) unit with a nuh_layer_id greater than 0. For this reason, the value of the ID of the SPS, that is, the syntax element sps_video_parameter_set_id, is greater than 0. Here, nuh_layer_id is information specifying the ID of one layer among the multi-layers. The NAL unit is like a packet and includes data for one picture. And an AU is composed of a plurality of NAL units.

[0609] Also, in this aspect, the decoding device 200 decodes the encoded bitstream and checks whether the SPS refers to the VPS when decoding the encoded bitstream. As described above, when the SPS refers to the VPS, a multi-layer structure is included in the encoded bitstream, and when the SPS does not refer to the VPS, a multi-layer structure is not included in the encoded bitstream. More specifically, when the SPS refers to the VPS, the ID of the SPS has a value of 1 described therein, and when the SPS does not refer to the VPS, the ID of the SPS has a value of 0 described therein.

[0610] Here, assume that the sps_video_parameter_set_id among the syntax elements shown in FIG. 98 is used as the ID of the SPS where a value of 0 or 1 is described. In this case, the decoding device 200 can know when the VPS is used and when it is not used by checking the value of the sps_video_parameter_set_id.

[0611] For example, when the ID of the SPS, that is, the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 0, the decoding device 200 does not refer to the VPS. Therefore, the decoding device 200 determines that the coded bitstream does not include a multi-layer structure, and does not extract and decode layers from the coded bitstream. In other words, when the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 0, the VPS is not referred to, and layers cannot be extracted from the coded bitstream.

[0612] More specifically, assume that the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 0. In this case, the value of the variable TargetLayerIdList can only include 0, and when applying the sub-bitstream extraction process, only layer 0, which is the base layer, can be extracted. Note that TargetLayerIdList is a variable indicating a list of extended layers to be decoded, and when the value is 0, it indicates that there is no extended layer to be decoded.

[0613] Therefore, when the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 0, since there is no extended layer to be decoded, the decoding device 200 is prohibited from using reference pictures that span layers. That is, when the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 0, the decoding device 200 cannot use reference pictures that span layers in the reference picture list.

[0614] On the one hand, when the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 1 (when it is greater than 0), since there is an extended layer to be decoded, the decoder 200 is permitted to use reference pictures that span layers. Also, when the ID of the SPS, that is, the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 1 (when it is greater than 0), the decoder 200 refers to the VPS. In this way, by confirming that the SPS refers to a valid VPS, the decoder 200 can use the syntax elements notified by the valid VPS in the decoding process.

[0615] Here, for example, assume that the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 is 1 and the encoded bitstream includes a three-layer multi-layer structure. In this case, since the value of the variable TargetLayerIdList can include any of 0 to 2, when applying the sub-bitstream extraction process, not only layer 0 which is the base layer but also layers 1 and 2 which are extended layers can be extracted and decoded.

[0616] In this way, when decoding the encoded bitstream, the decoder 200 can determine the value of the variable TargetLayerIdList when the SPS refers to the VPS. On the other hand, when decoding the encoded bitstream, the decoder 200 determines the value of the variable TargetLayerIdList to be 0 when the SPS does not refer to the VPS. That is, the decoder 200 can determine which pictures are output from the decoding picture buffer by checking the value of the syntax element sps_video_parameter_set_id shown in FIG. 98 and determining the value of the variable TargetLayerIdList.

[0617] Next, an example of the process by which the decoder 200 according to this aspect decodes the encoded bitstream will be described.

[0618] FIG. 99 is a flowchart showing an example of a method for decoding an encoded bitstream performed by the decoding apparatus 200 according to the first aspect.

[0619] First, the decoding apparatus 200 (the entropy decoding unit 202 thereof) analyzes a parameter set available in the encoded bitstream (S101). Here, examples of the parameter set include DPS (Decoding Parameter Set), VPS, SPS, PPS (Picture Parameter Set), and APS (Adaptation Parameter Set). Any encoded bitstream includes at least one SPS.

[0620] Next, the decoding apparatus 200 analyzes the encoded video sequence and determines the SPS to be used (S102). Here, the entropy decoding unit 202 of the decoding apparatus 200 determines to use at least the syntax element sps_video_parameter_set_id shown in FIG. 98 as the SPS.

[0621] Next, the decoding apparatus 200 checks whether the value of the syntax element sps_video_parameter_set_id is 0 (S103).

[0622] In step S103, if the value of the syntax element sps_video_parameter_set_id is 0 (yes in S103), the decoding apparatus 200 determines the value of the variable TargetLayerIdList to be 0 (S104). More specifically, when the value of the syntax element sps_video_parameter_set_id is 0, since the encoded bitstream does not include a multi-layer structure, only layer 0, which is the base layer, can be extracted. Therefore, the decoding apparatus 200 determines and sets the value of the variable TargetLayerIdList to 0 so that only layer 0 is included as the extended layer to be decoded.

[0623] On the one hand, in step S103, if the value of the syntax element sps_video_parameter_set_id is not 0 (no in S103), the decoder 200 acquires information from the VPS referred to in the SPS (S105). More specifically, when the value of the syntax element sps_video_parameter_set_id is 1, as the VPS to be referred to, information regarding the layer is acquired from vps_video_parameter_set_id equal to sps_video_parameter_set_id.

[0624] Next, the decoder 200 determines the value of the variable TargetLayerIdList using the information acquired in step S105 (S106). More specifically, the decoder 200 can determine which layers among the extended layers are to be decoded by determining and setting the value of the variable TargetLayerIdList using the information regarding the layer acquired in step S105. Note that the value of the variable TargetLayerIdList may be determined not only using the information regarding the layer acquired in step S105 but also using external inputs such as user input. Also, the value of the variable TargetLayerIdList may be determined using external inputs such as user input without using the information regarding the layer.

[0625] Next, the decoder 200 determines the value of the variable HighestTid (S107). Here, the variable HighestTid is a variable indicating the highest layer among one or more layers to be decoded. For example, if the value of the variable HighestTid is 0, it indicates that the base layer is the highest layer to be decoded. The value of the variable HighestTid may be determined using external inputs such as user input.

[0626] Next, the decoding device 200 extracts a sub-bitstream using the determined value of the variable TargetLayerIdList and the value of the variable HighestTid (S108). More specifically, the decoding device 200 provides the value of the variable TargetLayerIdList determined in step S104 or S106 and the value of the variable HighestTid determined in step S107 to the sub-bitstream extraction process. The decoding device 200 executes the sub-bitstream extraction process and extracts the sub-bitstream.

[0627] Finally, the decoding device 200 decodes the encoded bitstream by decoding the extracted sub-bitstream (S109).

[0628] [Effect of the First Aspect] According to the first aspect of the present disclosure, in the case where the encoding device 100 and the encoding method generate an encoded bitstream that does not include a multi-layer structure, since the VPS is not referenced in the SPS, it may not be necessary to transmit the VPS. That is, it may be possible to avoid transmitting the VPS of the encoded bitstream including a single-layer structure. Therefore, since the encoding device 100 and the encoding method do not have to include the VPS in the encoded bitstream, there is a possibility that the processing efficiency can be improved together with the encoding efficiency.

[0629] Similarly, according to the first aspect of the present disclosure, the decoding device 200 and the decoding method can confirm whether the encoded bitstream to be decoded includes a multi-layer structure only by checking whether the SPS references the VPS. Therefore, there is a possibility that the processing efficiency of the decoding device 200 when decoding the encoded bitstream can be improved.

[0630] [Second Aspect] Hereinafter, as a second aspect, an example in which the encoding device 100 prohibits including a VPS having an ID with a value of 0 in the encoded bitstream will be described.

[0631] FIG. 100 is a diagram showing an example of syntax elements included in the VPS used in this embodiment.

[0632] The VPS used in this embodiment is a VPS referred to by the SPS, and the vps_video_parameter_set_id among the syntax elements shown in FIG. 100 is used. The syntax element vps_video_parameter_set_id includes a unique identification number within the syntax, and a value of 0 or a non-zero value can be described.

[0633] Also, in this embodiment, whether the SPS refers to the VPS or not can be used to determine whether the encoded stream includes a multi-layer structure.

[0634] However, the SPS, that is, the syntax element sps_video_parameter_set_id, cannot refer to a VPS with a value of 0, that is, a syntax element vps_video_parameter_set_id with a value of 0. This is because a vps_video_parameter_set_id with a value of 0 indicates that the encoded stream does not include a multi-layer structure. That is, when the value of the syntax element sps_video_parameter_set_id is 1, the VPS, that is, the syntax element vps_video_parameter_set_id, is referred to in order to indicate that the encoded bitstream includes a multi-layer structure. However, if the value of the syntax element vps_video_parameter_set_id referred to by the SPS is 0, the encoded bitstream will not include a multi-layer structure. For this reason, on the decoding side, even if the SPS is analyzed, it is impossible to determine whether the encoded bitstream includes a multi-layer structure or not, resulting in confusion.

[0635] Therefore, in this aspect, it is assumed that the syntax element vps_video_parameter_set_id does not have a value of 0. In other words, in this aspect, it is prohibited to set 0 for the value of the syntax element vps_video_parameter_set_id which is a VPS included in the encoded bitstream.

[0636] [Effect of the Second Aspect] According to the second aspect of the present disclosure, it is prohibited to set 0 for the value of the syntax element vps_video_parameter_set_id which is a VPS included in the encoded bitstream. As a result, it is possible to avoid transmitting a VPS that cannot be referenced by any SPS, and thus there is a possibility of avoiding confusion on the decoding side.

[0637] More specifically, according to the second aspect of the present disclosure, the encoding apparatus 100 and the encoding method can suppress contradictions such as indicating that although the SPS references the VPS, the referenced VPS is an encoded bitstream that does not include a multi-layer structure. Therefore, the encoding apparatus can suppress causing confusion on the decoding side due to the content indicated by the VPS and the SPS being different.

[0638] Similarly, according to the second aspect of the present disclosure, the decoding apparatus 200 and the decoding method can suppress contradictions such as indicating that although the SPS references the VPS, the referenced VPS is an encoded bitstream that does not include a multi-layer structure. Therefore, the decoding apparatus 200 may be able to suppress the occurrence of confusion due to the content indicated by the video parameter set and the sequence parameter set being different.

[0639] [Typical Examples of Configuration and Processing] Typical examples of the configuration and processing of the encoding apparatus 100 and the decoding apparatus 200 shown above are shown below.

[0640] FIG. 101 is a flowchart showing the operations performed by the encoding apparatus 100 according to the embodiment. For example, the encoding apparatus 100 includes a circuit and a memory connected to the circuit. The circuit and the memory included in the encoding apparatus 100 may correspond to the processor a1 and the memory a2 shown in FIG. 8. The circuit of the encoding apparatus 100 performs the following in operation.

[0641] For example, the circuit of the encoding apparatus 100 checks whether to include a multi-layer structure in the encoded bitstream (S311). If it is to include a multi-layer structure in the encoded bitstream in step S311 (yes in S311), the circuit of the encoding apparatus 100 generates an encoded bitstream including a sequence parameter set that refers to the video parameter set (S312). On the other hand, if it is to generate the encoded bitstream without including a multi-layer structure in step S311 (no in S311), the circuit of the encoding apparatus 100 generates an encoded bitstream including a sequence parameter set that does not refer to the video parameter set (S313).

[0642] Thereby, when the encoding apparatus 100 generates an encoded bitstream that does not include a multi-layer structure, the video parameter set may not be referred to in the sequence parameter set, so there may be a case where the video parameter set does not need to be transmitted. Therefore, since the encoding apparatus 100 may not transmit the video parameter set, the video parameter set does not need to be included in the encoded bitstream, and there is a possibility that the processing efficiency can be improved together with the encoding efficiency.

[0643] Note that the entropy encoding unit 110 of the encoding apparatus 100 may perform the above-described operations as the circuit of the encoding apparatus 100.

[0644] FIG. 102 is a flowchart showing the operations performed by the decoding apparatus 200 according to the embodiment. For example, the decoding apparatus 200 includes a circuit and a memory connected to the circuit. The circuit and the memory included in the decoding apparatus 200 may correspond to the processor b1 and the memory b2 shown in FIG. 68. The circuit of the decoding apparatus 200 performs the following in operation.

[0645] For example, the circuit of the decoding apparatus 200 checks whether the sequence parameter set refers to the video parameter set (S411). In step S411, if the sequence parameter set refers to the video parameter set (yes in S411), the circuit of the decoding apparatus 200 decodes the encoded bitstream including the multi-layer structure (S412). On the other hand, in step S411, if the sequence parameter set does not refer to the video parameter set (no in S411), the circuit of the decoding apparatus 200 decodes the encoded bitstream not including the multi-layer structure (S413).

[0646] Thereby, the decoding apparatus 200 can check whether the encoded bitstream to be decoded includes a multi-layer structure only by checking whether the sequence parameter set refers to the video parameter set. Therefore, there is a possibility that the decoding apparatus can improve the processing efficiency when decoding the encoded bitstream.

[0647] Note that the entropy decoding unit 202 of the decoding apparatus 200 may perform the above-described operations as the circuit of the decoding apparatus 200.

[0648] [Other Examples] The encoding apparatus 100 and the decoding apparatus 200 in each of the above-described examples may be used as an image encoding apparatus and an image decoding apparatus, respectively, or may be used as a moving image encoding apparatus and a moving image decoding apparatus.

[0649] Alternatively, the encoding device 100 and the decoding device 200 may each be used as an entropy encoding device and an entropy decoding device. That is, the encoding device 100 and the decoding device 200 may each correspond only to the entropy encoding unit 110 and the entropy decoding unit 202. And the other components may be included in other devices.

[0650] Also, the encoding device 100 may include an input unit and an output unit. For example, one or more pictures are input to the input unit of the encoding device 100, and an encoded bit stream is output from the output unit of the encoding device 100. The decoding device 200 may also include an input unit and an output unit. For example, an encoded bit stream is input to the input unit of the decoding device 200, and one or more pictures are output from the output unit of the decoding device 200. The encoded bit stream may include quantized coefficients to which variable length coding is applied and control information.

[0651] Also, at least a part of each of the above-described examples may be used as an encoding method, a decoding method, an entropy encoding method, an entropy decoding method, or other methods.

[0652] Also, each component may be configured by dedicated hardware or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0653] Specifically, each of the encoding device 100 and the decoding device 200 may include a processing circuit and a storage device electrically connected to the processing circuit and accessible from the processing circuit. For example, the processing circuit corresponds to processor a1 or b1, and the storage device corresponds to memory a2 or b2.

[0654] The processing circuit includes at least one of dedicated hardware and a program execution unit, and executes processing using a storage device. Further, when the processing circuit includes a program execution unit, the storage device stores a software program executed by the program execution unit.

[0655] Here, the software that realizes the above-described encoding device 100 or decoding device 200 is the following program.

[0656] For example, this program causes a computer to execute an encoding method of generating an encoded bitstream, and when generating the encoded bitstream including a multi-layer structure therein, generating the encoded bitstream including a sequence parameter set that refers to a video parameter set, and when generating the encoded bitstream without including a multi-layer structure therein, generating the encoded bitstream including a sequence parameter set that does not refer to the video parameter set.

[0657] Further, for example, this program causes a computer to execute a decoding method of decoding an encoded bitstream, and when decoding the encoded bitstream, checking whether a sequence parameter set refers to a video parameter set, and when the sequence parameter set refers to the video parameter set, the encoded bitstream includes a multi-layer structure, and when the sequence parameter set does not refer to the video parameter set, the encoded bitstream does not include a multi-layer structure.

[0658] Also, as described above, each component may be a circuit. These circuits may constitute one circuit as a whole, or may be separate circuits respectively. Further, each component may be realized by a general-purpose processor or a dedicated processor.

[0659] Also, another component may execute the processes executed by a specific component. Also, the order in which the processes are executed may be changed, or a plurality of processes may be executed in parallel. Further, the encoding / decoding device may include the encoding device 100 and the decoding device 200.

[0660] Also, ordinal numbers such as the first and second used in the description may be appropriately changed. Also, ordinal numbers may be newly given to or removed from components and the like.

[0661] As described above, the aspects of the encoding device 100 and the decoding device 200 have been described based on a plurality of examples. However, the aspects of the encoding device 100 and the decoding device 200 are not limited to these examples. As long as the gist of the present disclosure is not deviated from, forms obtained by applying various modifications conceived by those skilled in the art to each example or by combining components in different examples may also be included in the scope of the aspects of the encoding device 100 and the decoding device 200.

[0662] One or more aspects disclosed herein may be implemented in combination with at least a part of other aspects in the present disclosure. Also, a part of the processes described in the flowchart of one or more aspects disclosed herein, a part of the configuration of the device, a part of the syntax, etc. may be implemented in combination with other aspects.

[0663] [Implementation and Application] In each of the above embodiments, each of the functional or operative blocks can usually be realized by an MPU (micro processing unit), a memory, and the like. Also, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads...

Claims

1. A circuit, and a memory connected to the circuit, wherein the circuit, in operation, generates a CVS (Coded Video Sequence), generates the SPS including a VPS identifier indicating an identifier of a VPS (Video Parameter Set) referred to by the SPS (Sequence Parameter Set), when the CVS has one or more layers, the VPS identifier is set to a value greater than 0, the SPS having the VPS identifier with a value of zero indicates (i) that the SPS does not refer to the VPS and (ii) that the CVS includes only one layer, none of the VPSs has an identifier of the VPS with a value of zero, an encoding apparatus.

2. A circuit, and a memory connected to the circuit, wherein the circuit, in operation, obtains from a bitstream the CVS (Coded Video Sequence) and the SPS including a VPS identifier indicating an identifier of a VPS (Video Parameter Set) referred to by the SPS (Sequence Parameter Set), decodes the CVS, when the CVS has one or more layers, the VPS identifier is set to a value greater than 0, the SPS having the VPS identifier with a value of zero indicates (i) that the SPS does not refer to the VPS and (ii) that the CVS includes only one layer, none of the VPSs has an identifier of the VPS with a value of zero, a decoding apparatus.

3. Generate a CVS (Coded Video Sequence), generate the SPS including a VPS identifier indicating an identifier of a VPS (Video Parameter Set) referred to by the SPS (Sequence Parameter Set), When the CVS has one or more layers, the VPS identifier is set to a value greater than 0. The SPS having the VPS identifier with a value of zero indicates that (i) the SPS does not refer to the VPS and (ii) the CVS includes only one layer. None of the VPSs have an identifier of the VPS with a value of zero. Encoding method. **Claim 4**: Obtain from a bitstream the SPS that includes a VPS identifier indicating an identifier of a VPS (Video Parameter Set) referred to by a CVS (Coded Video Sequence) and an SPS (Sequence Parameter Set), decode the CVS, when the CVS has one or more layers, the VPS identifier is set to a value greater than 0, the SPS having the VPS identifier of 0 indicates that (i) the SPS does not refer to the VPS and (ii) the CVS includes only one layer, none of the VPSs have an identifier of the VPS with a value of 0. Decoding method. **Claim 5** A bitstream generation device that generates a bitstream, comprising a memory, and a circuit connected to the memory, wherein the circuit generates a CVS (Coded Video Sequence), generates the SPS that includes a VPS identifier indicating an identifier of a VPS (Video Parameter Set) referred to by an SPS (Sequence Parameter Set), when the CVS has one or more layers, the VPS identifier is set to a value greater than 0, the SPS having the VPS identifier with a value of zero indicates that (i) the SPS does not refer to the VPS and (ii) the CVS includes only one layer. None of the VPSs have an identifier of the VPS with a value of zero. Bitstream generation device.