Encoding device, decoding device, and bit stream generation device

By introducing a picture-level and sub-picture-level slice index allocation mechanism in the video encoding device, the need for improved encoding efficiency and image quality in the existing technology is addressed, the processing volume and circuit size are reduced, the processing speed and the appropriateness of element selection are improved, and the encoding and decoding processes are optimized.

CN120915955APending Publication Date: 2025-11-07PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511235364.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-12-11
Filing Date
2020-12-08
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

There is a need to improve existing video coding technologies in terms of coding efficiency, image quality, processing volume, circuit scale, and processing speed. In particular, there is a lack of effective means to select appropriate elements or actions such as filters, blocks, sizes, motion vectors, and reference images or reference blocks.

Method used

By introducing a picture-level and sub-picture-level slice index allocation mechanism in the encoding and decoding devices, the sequential consistency and continuity of the indexes in the bitstream are ensured, reducing processing complexity, and the encoding and decoding process is optimized through the synergy of circuits and memory.

Benefits of technology

It achieves improvements in encoding efficiency, image quality, reduced processing volume, smaller circuit size, and faster processing speed. At the same time, it appropriately selects elements such as filters, blocks, sizes, motion vectors, and reference images to improve the overall performance of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120915955A_ABST
    Figure CN120915955A_ABST
Patent Text Reader

Abstract

The invention relates to an encoding device, a decoding device, and a bitstream generation device. A picture includes a plurality of sub-pictures, the picture includes a plurality of slices, each of the plurality of slices is included in one of the plurality of sub-pictures, and a picture-level slice index assigned at a picture level and a sub-picture-level slice index assigned at a sub-picture level are assigned to each of the plurality of slices. A plurality of sub-picture-level slice indexes are encoded into a plurality of slice headers respectively corresponding to a plurality of slices, each of the plurality of slices is encoded into a bitstream, and a picture-level slice index of a slice to be processed included in a sub-picture to be processed is assigned to the bitstream. The NAL unit is calculated by adding (1) the value of a sub-picture level slice index of a slice to be processed, which is an integer value that increases one by one from 0, and (2) the total number of slices included in a sub-picture that is encoded earlier than the sub-picture to be processed, the sub-picture level slice index being stored in a slice header and being an integer value that increases one by one from 0, and the bit stream including a plurality of slices as a plurality of NAL units. The order of the plurality of NAL units is determined based on the order of the sub-picture level slice indices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of the patent application with application number 202080082276.2, filed on December 8, 2020, and with the title “Encoding apparatus and decoding apparatus”. TECHNICAL FIELD

[0002] The present application relates to an encoding apparatus and a decoding apparatus. BACKGROUND

[0003] Video coding technology has progressed from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to handle the ever-increasing amount of digital video data in various uses, there is always a need to provide improvements and optimizations in video coding technology. The present application relates to further progress, improvements, and optimizations in video coding.

[0004] Furthermore, Non-Patent Literature 1 relates to an example of an existing standard related to the above-described video coding technology.

[0005] Prior Art Documents

[0006] Non-Patent Literature

[0007] Non-Patent Literature 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) SUMMARY

[0008] Problems to be Solved by the Invention

[0009] With regard to the above-described coding method, for improvement in coding efficiency, improvement in picture quality, reduction in processing amount, reduction in circuit size, or appropriate selection of elements or actions such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, or the like, it is desirable to propose a new method.

[0010] The present application provides a structure or a method that can contribute to one or more of, for example, improvement in coding efficiency, improvement in picture quality, reduction in processing amount, reduction in circuit size, improvement in processing speed, and appropriate selection of elements or actions. Furthermore, the present application can include a structure or a method that can contribute to benefits other than the above.

[0011] Means for Solving the Problems

[0012] For example, an encoding apparatus of an aspect of the present application includes a circuit and a memory connected to the circuit, a picture includes a plurality of sub-pictures, the picture includes a plurality of slices, each of the plurality of slices is included in one of the plurality of sub-pictures, the circuit, in operation, allocates a picture-level slice index given at a picture level and a sub-picture-level slice index given at a sub-picture level to each of the plurality of slices, encodes the plurality of sub-picture-level slice indexes into a plurality of slice headers respectively corresponding to the plurality of slices, and encodes each of the plurality of slices into a bitstream, the picture-level slice index given to a target slice included in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice to (2) a total number of slices included in a sub-picture that is encoded before the target sub-picture, the sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is incremented by 1 from 0, and the bitstream includes the plurality of slices as a plurality of NAL units, and an order of the plurality of NAL units is determined based on an order of sub-picture-level slice indexes.

[0013] For example, an encoding apparatus of an aspect of the present application includes a circuit and a memory connected to the circuit, a picture includes a plurality of sub-pictures, the picture includes a plurality of slices, each of the plurality of slices is included in one of the plurality of sub-pictures, the circuit, in operation, allocates a picture-level slice index given at a picture level and a sub-picture-level slice index given at a sub-picture level to each of the plurality of slices, encodes the plurality of sub-picture-level slice indexes into a plurality of slice headers respectively corresponding to the plurality of slices, and encodes each of the plurality of slices into a bitstream, the picture-level slice index given to a target slice included in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice to (2) a total number of slices included in a sub-picture that is encoded before the target sub-picture, the sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is incremented by 1 from 0, and the bitstream includes the plurality of slices as a plurality of NAL units, and an order of the plurality of NAL units is determined based on an order of sub-picture-level slice indexes.

[0014] For example, a bitstream generation apparatus of an aspect of the present application includes a circuit and a memory connected to the circuit, a picture includes a plurality of sub-pictures, the picture includes a plurality of slices, each of the plurality of slices is included in one of the plurality of sub-pictures, the circuit, in operation, assigns a picture-level slice index given at a picture level and a sub-picture-level slice index given at a sub-picture level to each of the plurality of slices, includes a plurality of the sub-picture-level slice indexes in a plurality of slice headers corresponding to the plurality of slices, respectively, generates a bitstream including the plurality of slice headers and the plurality of slices, the picture-level slice index given to a target slice included in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice to (2) a total number of slices included in a sub-picture that is encoded before the target sub-picture, the sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is increased by 1 from 0, the bitstream includes the plurality of slices as a plurality of NAL units, and an order of the plurality of NAL units is determined based on an order of the sub-picture-level slice indexes.

[0015] In video encoding technology, a new method is desired to be proposed for improvement of encoding efficiency, improvement of picture quality, reduction of circuit size, and the like.

[0016] The structure or method of each embodiment of the present application or a part thereof, respectively, can achieve at least any one of improvement of encoding efficiency, improvement of picture quality, reduction of processing amount of encoding / decoding, reduction of circuit size, or improvement of processing speed of encoding / decoding, and the like. Alternatively, the structure or method of each embodiment of the present application or a part thereof, respectively, can make appropriate selection of constituent elements / actions of filters, blocks, sizes, motion vectors, reference pictures, reference blocks, and the like in encoding and decoding. In addition, the present application also includes disclosure of a structure or method that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in processing amount, and the like.

[0017] Further advantages and effects of an aspect of the present application are clarified from the description and the drawings. The advantages and / or effects are respectively obtained by several embodiments and features described in the description and the drawings, but all of them are not necessarily required to obtain one or more of the advantages and / or effects.

[0018] In addition, these general or specific aspects can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can also be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0019] Effects of the Invention

[0020] The structure or method of an aspect of the present application can contribute to one or more of, for example, improvement of coding efficiency, improvement of picture quality, reduction of processing amount, reduction of circuit size, improvement of processing speed, and appropriate selection of elements or actions. In addition, the structure or method of an aspect of the present application can contribute to benefits other than the above. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a diagram showing an example of a structure of a transmission system of an embodiment.

[0022] Figure 2 is a diagram showing an example of a hierarchical structure of data in a stream.

[0023] Figure 3 is a diagram showing an example of a structure of a slice.

[0024] Figure 4 is a diagram showing an example of a structure of a tile.

[0025] Figure 5 is a diagram showing an example of an encoding structure at the time of scalable encoding.

[0026] Figure 6 is a diagram showing an example of an encoding structure at the time of scalable encoding.

[0027] Figure 7 is a block diagram showing an example of a functional structure of an encoding apparatus of an embodiment.

[0028] Figure 8 is a block diagram showing an example of a configuration of an encoding apparatus.

[0029] Figure 9 is a flowchart showing an example of an overall encoding process performed by an encoding apparatus.

[0030] Figure 10 is a diagram showing an example of block partitioning.

[0031] Figure 11 is a diagram showing an example of a functional structure of a partitioning section.

[0032] Figure 12 is a diagram showing an example of a partitioning pattern.

[0033] Figure 13A is a diagram showing an example of a syntax tree of a partitioning pattern.

[0034] Figure 13B is a diagram showing another example of a syntax tree of a partitioning pattern.

[0035] Figure 14is a table indicating a transform basis function corresponding to each transform type.

[0036] Figure 15 is a diagram indicating an example of SVT.

[0037] Figure 16 is a flowchart indicating an example of processing by the transform section.

[0038] Figure 17 is a flowchart indicating another example of processing by the transform section.

[0039] Figure 18 is a block diagram indicating an example of a functional structure of the quantization section.

[0040] Figure 19 is a flowchart indicating an example of quantization by the quantization section.

[0041] Figure 20 is a block diagram indicating an example of a functional structure of the entropy encoding section.

[0042] Figure 21 is a diagram indicating a flow of CABAC in the entropy encoding section.

[0043] Figure 22 is a block diagram indicating an example of a functional structure of the loop filtering section.

[0044] Figure 23A is a diagram indicating an example of a shape of a filter used in ALF (adaptive loop filter).

[0045] Figure 23B is a diagram indicating another example of a shape of a filter used in ALF.

[0046] Figure 23C is a diagram indicating another example of a shape of a filter used in ALF.

[0047] Figure 23D is a diagram indicating an example of CCALF for Cb using Y samples (1st component) and CCALF for Cr (a plurality of components different from the 1st component).

[0048] Figure 23E is a diagram indicating a diamond-shaped filter.

[0049] Figure 23F is a diagram indicating an example of JC-CCALF.

[0050] Figure 23G is a diagram indicating an example of weight_index candidates of JC-CCALF.

[0051] Figure 24 is a block diagram showing an example of a detailed structure of a loop filter that functions as a DBF.

[0052] Figure 25 is a diagram showing an example of deblocking filtering having a filter characteristic symmetrical with respect to a block boundary.

[0053] Figure 26 is a diagram for explaining an example of a block boundary on which deblocking filtering is performed.

[0054] Figure 27 is a diagram showing an example of a Bs value.

[0055] Figure 28 is a flowchart showing an example of a process performed by a prediction section of an encoding apparatus.

[0056] Figure 29 is a flowchart showing another example of a process performed by a prediction section of an encoding apparatus.

[0057] Figure 30 is a flowchart showing another example of a process performed by a prediction section of an encoding apparatus.

[0058] Figure 31 is a diagram showing an example of 67 intra prediction modes in intra prediction.

[0059] Figure 32 is a flowchart showing an example of a process performed by an intra prediction section.

[0060] Figure 33 is a diagram showing an example of each reference picture.

[0061] Figure 34 is a conceptual diagram showing an example of a reference picture list.

[0062] Figure 35 is a flowchart showing a flow of a basic process of inter prediction.

[0063] Figure 36 is a flowchart showing an example of MV derivation.

[0064] Figure 37 is a flowchart showing another example of MV derivation.

[0065] Figure 38A is a diagram showing an example of classification of each mode of MV derivation.

[0066] Figure 38B is a diagram showing an example of classification of each mode of MV derivation.

[0067] Figure 39 is a flowchart showing an example of inter prediction based on a normal inter mode.

[0068] Figure 40 is a flowchart showing an example of inter prediction based on the normal merge mode.

[0069] Figure 41 is a diagram for explaining an example of MV derivation processing based on the normal merge mode.

[0070] Figure 42 is a diagram for explaining an example of MV derivation processing based on the HMVP mode.

[0071] Figure 43 is a flowchart showing an example of FRUC (frame rate up conversion).

[0072] Figure 44 is a diagram for explaining an example of pattern matching (bi-directional matching) between two blocks along a motion trajectory.

[0073] Figure 45 is a diagram for explaining an example of pattern matching (template matching) between a template within a current picture and a block within a reference picture.

[0074] Figure 46A is a diagram for explaining an example of derivation of MVs in sub-block units in an affine mode using two control points.

[0075] Figure 46B is a diagram for explaining an example of derivation of MVs in sub-block units in an affine mode using three control points.

[0076] Figure 47A is a conceptual diagram for explaining an example of MV derivation of control points in an affine mode.

[0077] Figure 47B is a conceptual diagram for explaining an example of MV derivation of control points in an affine mode.

[0078] Figure 47C is a conceptual diagram for explaining an example of MV derivation of control points in an affine mode.

[0079] Figure 48A is a diagram for explaining an affine mode with two control points.

[0080] Figure 48B is a diagram for explaining an affine mode with three control points.

[0081] Figure 49A is a conceptual diagram for explaining an example of a control point MV derivation method in a case where the number of control points is different between an encoded block and a current block.

[0082] Figure 49B is a conceptual diagram for explaining another example of a method of deriving an MV of a control point in a case where the number of control points is different between an already coded block and a current block.

[0083] Figure 50 is a flowchart showing an example of a process of an affine merge mode.

[0084] Figure 51 is a flowchart showing an example of a process of an affine inter mode.

[0085] Figure 52A is a diagram for explaining generation of a prediction image of 2 triangles.

[0086] Figure 52B is a conceptual diagram showing an example of a first part of a first partition and a first sample set and a second sample set.

[0087] Figure 52C is a conceptual diagram showing a first part of a first partition.

[0088] Figure 53 is a flowchart showing an example of a triangle mode.

[0089] Figure 54 is a diagram showing an example of an ATMVP mode in which an MV is derived in a subblock unit.

[0090] Figure 55 is a diagram showing a relationship between a merge mode and DMVR (dynamic motion vector refreshing).

[0091] Figure 56 is a conceptual diagram for explaining an example of DMVR.

[0092] Figure 57 is a conceptual diagram for explaining another example of DMVR for deciding an MV.

[0093] Figure 58A is a diagram showing an example of a motion search in DMVR.

[0094] Figure 58B is a flowchart showing an example of a motion search in DMVR.

[0095] Figure 59 is a flowchart showing an example of generation of a prediction image.

[0096] Figure 60 is a flowchart showing another example of generation of a prediction image.

[0097] Figure 61is a flowchart for explaining an example of an OBMC (overlapped block motion compensation)-based prediction image correction process.

[0098] Figure 62 is a conceptual diagram for explaining an example of an OBMC-based prediction image correction process.

[0099] Figure 63 is a diagram for explaining a model assuming constant velocity straight line motion.

[0100] Figure 64 is a flowchart showing an example of inter prediction according to BIO.

[0101] Figure 65 is a diagram showing an example of functional structure of an inter prediction section that performs inter prediction according to BIO.

[0102] Figure 66A is a diagram for explaining an example of a prediction image generation method using LIC (local illumination compensation)-based luminance correction processing.

[0103] Figure 66B is a flowchart showing an example of a prediction image generation method using LIC-based luminance correction processing.

[0104] Figure 67 is a block diagram showing a functional structure of a decoding device of the embodiment.

[0105] Figure 68 is a block diagram showing an example of installation of a decoding device.

[0106] Figure 69 is a flowchart showing an example of overall decoding processing by the decoding device.

[0107] Figure 70 is a diagram showing a relationship of a partitioning decision section with other constituent elements.

[0108] Figure 71 is a block diagram showing an example of functional structure of an entropy decoding section.

[0109] Figure 72 is a diagram showing a flow of CABAC in the entropy decoding section.

[0110] Figure 73 is a block diagram showing an example of functional structure of an inverse quantization section.

[0111] Figure 74 is a flowchart showing an example of inverse quantization by the inverse quantization section.

[0112] Figure 75 is a flowchart showing an example of the processing by the inverse transform section.

[0113] Figure 76 is a flowchart showing another example of the processing by the inverse transform section.

[0114] Figure 77 is a block diagram showing an example of the functional structure of the loop filter section.

[0115] Figure 78 is a flowchart showing an example of the processing by the prediction section of the decoding device.

[0116] Figure 79 is a flowchart showing another example of the processing by the prediction section of the decoding device.

[0117] Figure 80A is a flowchart showing a part of another example of the processing by the prediction section of the decoding device.

[0118] Figure 80B is a flowchart showing the remaining part of another example of the processing by the prediction section of the decoding device.

[0119] Figure 81 is a diagram showing an example of the processing by the intra prediction section of the decoding device.

[0120] Figure 82 is a flowchart showing an example of the MV derivation in the decoding device.

[0121] Figure 83 is a flowchart showing another example of the MV derivation in the decoding device.

[0122] Figure 84 is a flowchart showing an example of the inter prediction based on the normal inter mode in the decoding device.

[0123] Figure 85 is a flowchart showing an example of the inter prediction based on the normal merge mode in the decoding device.

[0124] Figure 86 is a flowchart showing an example of the inter prediction based on the FRUC mode in the decoding device.

[0125] Figure 87 is a flowchart showing an example of the inter prediction based on the affine merge mode in the decoding device.

[0126] Figure 88 is a flowchart showing an example of the inter prediction based on the affine inter mode in the decoding device.

[0127] Figure 89 is a flowchart representing an example of triangle mode based inter prediction in a decoding apparatus.

[0128] Figure 90 is a flowchart representing an example of DMVR based motion search in a decoding apparatus.

[0129] Figure 91 is a flowchart representing a detailed example of DMVR based motion search in a decoding apparatus.

[0130] Figure 92 is a flowchart representing an example of prediction picture generation in a decoding apparatus.

[0131] Figure 93 is a flowchart representing another example of prediction picture generation in a decoding apparatus.

[0132] Figure 94 is a flowchart representing an example of OBMC based prediction picture modification in a decoding apparatus.

[0133] Figure 95 is a flowchart representing an example of BIO based prediction picture modification in a decoding apparatus.

[0134] Figure 96 is a flowchart representing an example of LIC based prediction picture modification in a decoding apparatus.

[0135] Figure 97 is a conceptual diagram representing a plurality of picture level slice indices with discontinuity within a subpicture.

[0136] Figure 98 is a conceptual diagram representing a plurality of picture level slice indices without discontinuity within a subpicture.

[0137] Figure 99 is a conceptual diagram representing a pseudo code for maintaining a picture level slice index of an initial slice in each subpicture.

[0138] Figure 100 is a conceptual diagram representing a pseudo code for calculating a picture level slice index of a processing target slice using a picture level slice index of an initial slice.

[0139] Figure 101 is a flowchart representing an encoding process of a subpicture level slice index at the time of encoding of a slice.

[0140] Figure 102 is a flowchart representing a calculation process of a picture level slice index at the time of decoding of a slice.

[0141] Figure 103is a conceptual diagram indicating a pseudo code for holding the number of slices in each sub-picture.

[0142] Figure 104 is a conceptual diagram indicating a pseudo code for calculating a picture-level slice index of a processing target slice using the number of slices.

[0143] Figure 105 is a flowchart indicating the action of an encoding apparatus of an embodiment.

[0144] Figure 106 is a flowchart indicating the action of a decoding apparatus of an embodiment.

[0145] Figure 107 is a whole configuration diagram indicating a content supply system that implements a content distribution service.

[0146] Figure 108 is a diagram indicating an example of a display screen of a web page.

[0147] Figure 109 is a diagram indicating an example of a display screen of a web page.

[0148] Figure 110 is a diagram indicating an example of a smart phone.

[0149] Figure 111 is a block diagram indicating an example of the structure of a smart phone. DETAILED DESCRIPTION

[0150] [INTRODUCTION]

[0151] In encoding and decoding of a moving image, pictures constituting the moving image are divided into various units of regions such as CTU, tile, slice, and sub-picture, whereby image processing for encoding and decoding of the moving image is efficiently performed.

[0152] As for the various units of regions, for example, a CTU corresponds to a square region of a fixed size. A tile is a rectangular region determined by one or more rows within a picture, one or more columns within the picture, or both. A slice corresponds to a NAL unit that is one data packet. Further, a slice can correspond to one or more tiles, or can also correspond to one or more CTUs in series in one tile.

[0153] A sub-picture is a rectangular region, and corresponds to one or more slices. Further, a sub-picture can correspond to one or more tiles, or can correspond to one or more CTU rows in one tile.

[0154] Each region obtained by dividing a picture in a region unit like this is sometimes identified by an index. For example, as an index for identifying a slice, there are two kinds of slice index, a picture-level slice index and a sub-picture-level slice index.

[0155] With the picture-level slice index, a slice can be identified from among a plurality of slices included in the entire picture. Further, with the sub-picture-level slice index, a slice can be identified from among one or more slices included in a sub-picture.

[0156] However, with the two kinds of slice index, the picture-level slice index and the sub-picture-level slice index, it is possible to complicate the processing.

[0157] Therefore, for example, an encoding apparatus of one aspect of the present application has a circuit and a memory connected to the circuit, the circuit, in operation, assigns a plurality of picture-level slice indexes to a plurality of slices included in a picture, respectively, the plurality of picture-level slice indexes being continuously and uninterruptedly in the entire picture and being continuously and uninterruptedly in a plurality of sub-pictures included in the picture, respectively, assigns a plurality of sub-picture-level slice indexes to the plurality of slices, respectively, the plurality of sub-picture-level slice indexes being continuously and uninterruptedly in each of the plurality of sub-pictures and having the same order as the order of the plurality of picture-level slice indexes in each of the plurality of sub-pictures, encodes the plurality of sub-picture-level slice indexes into a plurality of slice headers corresponding to the plurality of slices, respectively, and encodes each of the plurality of slices into a bitstream.

[0158] Thereby, one or more picture-level slice indexes and one or more sub-picture-level slice indexes are assigned to one or more slices included in each sub-picture in the same order, and it is possible to suppress complication of the processing.

[0159] Further, for example, the plurality of picture-level slice indexes are continuously and uninterruptedly from 0 in the entire picture, and the plurality of sub-picture-level slice indexes are continuously and uninterruptedly from 0 in each of the plurality of sub-pictures.

[0160] Thereby, it is possible to suppress the picture-level slice index and the sub-picture-level slice index from becoming too large.

[0161] Further, for example, the bitstream includes, for each of the plurality of sub-pictures, one or more slices included in the sub-picture as one or more NAL units, and for each of the plurality of sub-pictures, the order of one or more sub-picture-level slice indexes assigned to the one or more slices among the plurality of sub-picture-level slice indexes coincides with the order of the one or more slices in the bitstream.

[0162] Thus, it is possible to appropriately assign one or more picture-level slice indexes and one or more sub-picture-level slice indexes to each sub-picture in the order of one or more slices in the bitstream.

[0163] Further, for example, for each of the plurality of slices, a slice header corresponding to the slice includes a slice_address parameter as an encoding parameter, the slice_address parameter indicating a sub-picture-level slice index of the slice, for each of the plurality of sub-pictures, the one or more sub-picture-level slice indexes are indicated by one or more slice_address parameters included in one or more slice headers corresponding to the one or more slices, the order of the one or more sub-picture-level slice indexes indicated by the one or more slice_address parameters coincides with the order of the one or more slices in the bitstream.

[0164] Thus, each of the one or more sub-picture-level slice indexes having an order coinciding with the order of one or more slices in the bitstream is possible to be appropriately indicated by a slice_address parameter.

[0165] Further, for example, the bitstream includes the plurality of slices as a plurality of NAL units, the order of the plurality of picture-level slice indexes assigned to the plurality of slices coincides with the order of the plurality of slices in the bitstream.

[0166] Thus, it is possible to appropriately assign a plurality of picture-level slice indexes and a plurality of sub-picture-level slice indexes to the entirety of a picture in the order of a plurality of slices in the bitstream.

[0167] Further, for example, the plurality of slices include a processing target slice, the plurality of sub-pictures include a processing target sub-picture, the processing target sub-picture includes one or more slices including the processing target slice, and the circuitry calculates a picture-level slice index of the processing target slice by adding a sub-picture-level slice index of the processing target slice to a picture-level slice index of a slice included in the processing target sub-picture, which is a first slice in the bitstream.

[0168] Thus, it is possible to appropriately calculate a picture-level slice index of a processing target slice from a picture-level slice index of a slice included in a processing target sub-picture, which is a first slice in the bitstream.

[0169] Further, for example, the plurality of slices include a processing target slice, the plurality of sub-pictures include a processing target sub-picture, the processing target sub-picture includes one or more slices including the processing target slice, and the circuit calculates a picture-level slice index of the processing target slice by adding a sub-picture-level slice index of the processing target slice to a number of one or more slices included in one or more sub-pictures before the processing target sub-picture in the plurality of sub-pictures.

[0170] Thus, it is possible to appropriately calculate a picture-level slice index of a processing target slice in accordance with a number of one or more slices included in one or more sub-pictures before the processing target sub-picture.

[0171] Further, for example, the circuit allocates a first picture-level slice index of the plurality of picture-level slice indexes to a first slice of the plurality of slices, allocates a first sub-picture-level slice index of the plurality of sub-picture-level slice indexes to the first slice, encodes the first sub-picture-level slice index to a slice header of the first slice, and encodes the first slice, and then allocates a second picture-level slice index of the plurality of picture-level slice indexes to a second slice of the plurality of slices, allocates a second sub-picture-level slice index of the plurality of sub-picture-level slice indexes to the second slice, encodes the second sub-picture-level slice index to a slice header of the second slice, and encodes the second slice.

[0172] Thus, since the plurality of slices are processed in order, it is possible to suppress the amount of memory usage.

[0173] Further, for example, a decoding device of an aspect of the present application includes a circuit and a memory connected to the circuit, and in operation, the circuit decodes a plurality of sub-picture-level slice indexes continuously and uninterruptedly included in each of a plurality of sub-pictures included in a picture from a plurality of slice headers respectively corresponding to a plurality of slices included in the picture, allocates the plurality of sub-picture-level slice indexes to the plurality of slices, respectively, allocates a plurality of picture-level slice indexes continuously and uninterruptedly in the entire picture and continuously and uninterruptedly in each of the plurality of sub-pictures and having the same order as an order of the plurality of sub-picture-level slice indexes in each of the plurality of sub-pictures to the plurality of slices, respectively, and decodes each of the plurality of slices from a bitstream.

[0174] Thus, one or more picture-level slice indexes and one or more sub-picture-level slice indexes are allocated to one or more slices included in each sub-picture in the same order, and it is possible to suppress complication of processing.

[0175] Further, for example, the plurality of picture-level slice indexes are continuous from 0 without interruption in the entirety of the picture, and the plurality of sub-picture-level slice indexes are continuous from 0 without interruption in each of the plurality of sub-pictures.

[0176] Thus, it is possible to suppress the picture-level slice indexes and the sub-picture-level slice indexes from becoming too large.

[0177] Further, for example, the bitstream contains, for each of the plurality of sub-pictures, one or more slices of the plurality of slices that are contained in the sub-picture as one or more NAL units, and the order of one or more sub-picture-level slice indexes of the plurality of sub-picture-level slice indexes that are assigned to the one or more slices coincides with the order of the one or more slices in the bitstream.

[0178] Thus, it is possible to appropriately assign one or more picture-level slice indexes and one or more sub-picture-level slice indexes to each sub-picture in the order of one or more slices in the bitstream.

[0179] Further, for example, for each of the plurality of slices, a slice header corresponding to the slice contains a slice_address parameter as an encoding parameter, the slice_address parameter indicating a sub-picture-level slice index of the slice, and for each of the plurality of sub-pictures, the one or more sub-picture-level slice indexes are indicated by one or more slice_address parameters contained in one or more slice headers corresponding to the one or more slices, the order of the one or more sub-picture-level slice indexes indicated by the one or more slice_address parameters coinciding with the order of the one or more slices in the bitstream.

[0180] Thus, each of the one or more sub-picture-level slice indexes having an order that coincides with the order of one or more slices in the bitstream is possible to be appropriately indicated by a slice_address parameter.

[0181] Further, for example, the bitstream contains the plurality of slices as a plurality of NAL units, and the order of the plurality of picture-level slice indexes assigned to the plurality of slices coincides with the order of the plurality of slices in the bitstream.

[0182] Thus, it is possible to appropriately assign a plurality of picture-level slice indexes and a plurality of sub-picture-level slice indexes to the entirety of a picture in the order of a plurality of slices in the bitstream.

[0183] Further, for example, the plurality of slices include a processing target slice, the plurality of sub-pictures include a processing target sub-picture, the processing target sub-picture includes one or more slices including the processing target slice, and the circuit calculates the picture-level slice index of the processing target slice by adding the sub-picture-level slice index of the processing target slice to the number of one or more slices included in the picture-level slice index of the first slice in the bitstream among the one or more slices included in the processing target sub-picture.

[0184] Thus, it is possible to appropriately calculate the picture-level slice index of the processing target slice from the picture-level slice index of the first slice among the one or more slices included in the processing target sub-picture.

[0185] Further, for example, the plurality of slices include a processing target slice, the plurality of sub-pictures include a processing target sub-picture, the processing target sub-picture includes one or more slices including the processing target slice, and the circuit calculates the picture-level slice index of the processing target slice by adding the sub-picture-level slice index of the processing target slice to the number of one or more slices included in the picture-level slice index of the first slice in the bitstream among the one or more slices included in the processing target sub-picture.

[0186] Thus, it is possible to appropriately calculate the picture-level slice index of the processing target slice from the number of one or more slices included in the one or more sub-pictures before the processing target sub-picture.

[0187] Further, for example, the circuit decodes a first sub-picture-level slice index in the plurality of sub-picture-level slice indexes from a slice header of a first slice in the plurality of slices, assigns the first sub-picture-level slice index to the first slice, assigns a first picture-level slice index in the plurality of picture-level slice indexes to the first slice, and decodes the first slice, and then decodes a second sub-picture-level slice index in the plurality of sub-picture-level slice indexes from a slice header of a second slice in the plurality of slices, assigns the second sub-picture-level slice index to the second slice, assigns a second picture-level slice index in the plurality of picture-level slice indexes to the second slice, and decodes the second slice.

[0188] Thus, since the plurality of slices are processed sequentially, it is possible to suppress the amount of memory usage.

[0189] Further, for example, an encoding method of an aspect of the present application assigns a plurality of picture-level slice indexes to a plurality of slices included in a picture, respectively, the plurality of picture-level slice indexes being continuously and uninterruptedly in an entirety of the picture and being continuously and uninterruptedly in each of a plurality of sub-pictures included in the picture, assigns a plurality of sub-picture-level slice indexes to the plurality of slices, respectively, the plurality of sub-picture-level slice indexes being continuously and uninterruptedly in each of the plurality of sub-pictures and having the same order as an order of the plurality of picture-level slice indexes in each of the plurality of sub-pictures, encodes the plurality of sub-picture-level slice indexes into a plurality of slice headers corresponding to the plurality of slices, respectively, and encodes the plurality of slices into a bitstream, respectively.

[0190] Thus, one or more picture-level slice indexes and one or more sub-picture-level slice indexes are assigned to one or more slices included in each sub-picture in the same order, and it is possible to suppress complication of processing.

[0191] Further, for example, a decoding method of an aspect of the present application decodes a plurality of sub-picture-level slice indexes that are continuously and uninterruptedly in each of a plurality of sub-pictures included in a picture from a plurality of slice headers corresponding to a plurality of slices included in the picture, respectively, assigns the plurality of sub-picture-level slice indexes to the plurality of slices, respectively, assigns a plurality of picture-level slice indexes that are continuously and uninterruptedly in an entirety of the picture and are continuously and uninterruptedly in each of the plurality of sub-pictures and have the same order as an order of the plurality of sub-picture-level slice indexes in each of the plurality of sub-pictures to the plurality of slices, respectively, and decodes each of the plurality of slices from a bitstream.

[0192] Thus, one or more picture-level slice indexes and one or more sub-picture-level slice indexes are assigned to one or more slices included in each sub-picture in the same order, and it is possible to suppress complication of processing.

[0193] Further, for example, an encoding apparatus of an aspect of the present application includes an input unit, a division unit, an intra prediction unit, an inter prediction unit, a loop filter unit, a transform unit, a quantization unit, an entropy coding unit, and an output unit.

[0194] The input unit inputs a current picture. The division unit divides the current picture into a plurality of blocks.

[0195] The intra prediction unit generates a prediction signal of a current block included in the current picture using a reference image included in the current picture. The inter prediction unit generates a prediction signal of a current block included in the current picture using a reference image included in a reference picture different from the current picture. The loop filter unit applies a filter to a reconstructed block of the current block included in the current picture.

[0196] The transform section transforms a prediction error of an original signal of a current block included in a current picture and a prediction signal generated by the intra prediction section or the inter prediction section, and generates a transform coefficient. The quantization section quantizes the transform coefficient, and generates a quantized coefficient. The entropy encoding section applies variable length encoding to the quantized coefficient, and generates an encoded bitstream. And, the encoded bitstream including the quantized coefficient to which the variable length encoding is applied and control information is output from the output section.

[0197] Further, for example, the entropy encoding section, in operation, allocates a plurality of picture-level slice indexes to a plurality of slices included in a picture respectively, the plurality of picture-level slice indexes being continuously uninterrupted in the entirety of the picture and continuously uninterrupted in each of a plurality of sub-pictures included in the picture, allocates a plurality of sub-picture-level slice indexes to the plurality of slices respectively, the plurality of sub-picture-level slice indexes being continuously uninterrupted in each of the plurality of sub-pictures and having the same order as an order of the plurality of picture-level slice indexes in each of the plurality of sub-pictures, encodes the plurality of sub-picture-level slice indexes into a plurality of slice headers corresponding to the plurality of slices respectively, and encodes each of the plurality of slices into a bitstream.

[0198] Further, for example, a decoding apparatus of an aspect of the present application includes an input section, an entropy decoding section, an inverse quantization section, an inverse transform section, an intra prediction section, an inter prediction section, a loop filtering section, and an output section.

[0199] An encoded bitstream is input to the input section. The entropy decoding section applies variable length decoding to the encoded bitstream, and derives a quantized coefficient. The inverse quantization section inversely quantizes the quantized coefficient, and derives a transform coefficient. The inverse transform section inversely transforms the transform coefficient, and derives a prediction error.

[0200] The intra prediction section generates a prediction signal of a current block included in a current picture using a reference image included in the current picture. The inter prediction section generates a prediction signal of a current block included in a current picture using a reference image included in a reference picture different from the current picture.

[0201] The loop filtering section applies a filter to a reconstructed block of a current block included in the current picture. Then, the current picture is output from the output section.

[0202] Further, for example, the entropy decoding section, in operation, decodes, from a plurality of slice headers respectively corresponding to a plurality of slices included in a picture, a plurality of subpicture-level slice indexes continuously and uninterruptedly included in each of a plurality of subpictures included in the picture, respectively allocates the plurality of subpicture-level slice indexes to the plurality of slices, respectively allocates a plurality of picture-level slice indexes continuously and uninterruptedly in the entire picture and continuously and uninterruptedly in each of the plurality of subpictures and having the same order as the order of the plurality of subpicture-level slice indexes in each of the plurality of subpictures to the plurality of slices, and decodes each of the plurality of slices from a bitstream.

[0203] Also, these inclusive or specific modes can be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a CD-ROM that is computer-readable, and can be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0204] [Definitions of Terms]

[0205] As an example, each term can be defined in the following manner.

[0206] (1) Image

[0207] is a unit of data composed of a set of pixels, composed of a picture or a block smaller than a picture, and includes still images in addition to moving images.

[0208] (2) Picture

[0209] is a processing unit of an image composed of a set of pixels, sometimes referred to as a frame or a field.

[0210] (3) Block

[0211] is a processing unit including a set of a certain number of pixels, and the name is not limited as listed in the following examples. Also, the shape is not limited, and for example, of course includes a rectangle composed of M x N pixels, a square composed of M x M pixels, and also includes a triangle, a circle, and other shapes.

[0212] (Examples of Blocks)

[0213] • Slice / Tile / Brick

[0214] • CTU / Superblock / Basic Partition Unit

[0215] • VPDU / Hardware Processing Partition Unit

[0216] • CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transform Block Unit (TU) / Unit

[0217] • Sub-block

[0218] (4) Pixel / sample

[0219] is a point that constitutes the smallest unit of an image, and includes not only an integer-position pixel but also a fractional-position pixel generated based on the integer-position pixel.

[0220] (5) Pixel value / sample value

[0221] is an inherent value that a pixel has, and of course includes a luminance value, a color difference value, a gray scale of RGB, and further includes a depth value or a 0, 1 binary value.

[0222] (6) Flag

[0223] In addition to 1 bit, there is a case where a plurality of bits are included, for example, it can also be a parameter or an index of 2 bits or more. In addition, not only a binary value using a binary number but also a multi-value using another number of digits can be used.

[0224] (7) Signal

[0225] is a signal that is subjected to symbolization or encoding in order to transmit information, and in addition to a digital signal that is discretized, an analog signal that takes a continuous value is also included.

[0226] (8) Stream / bit stream

[0227] refers to a data string of digital data or a stream of digital data. The stream / bit stream can be composed of a plurality of streams in addition to 1 stream, and can be divided into a plurality of hierarchies. In addition, in addition to a case where it is transmitted on a single transmission path through serial communication, a case where it is transmitted through packet communication in a plurality of transmission paths is also included.

[0228] (9) Difference / differential

[0229] In the case of a scalar, as long as an operation including a difference is included in addition to a simple difference (x-y), an absolute value of a difference (|x-y|), a squared difference (x^2-y^2), a square root of a difference (sqrt(x-y)), a weighted difference (ax-by: a, b are constants), and an offset difference (x-y+a: a is an offset) are included.

[0230] (10) Sum

[0231] In the case of a scalar, as long as an operation including a sum is included in addition to a simple sum (x+y), an absolute value of a sum (|x+y|), a squared sum (x^2+y^2), a square root of a sum (sqrt(x+y)), a weighted sum (ax+by: a, b are constants), and an offset sum (x+y+a: a is an offset) are included.

[0232] (11) based on

[0233] Also included is a case where an element other than the object based on is added. In addition, a case where a result is obtained via an intermediate result is included in addition to a case where a direct result is obtained.

[0234] (12) used, using

[0235] Also included is a case where an element other than the object used is added. In addition, a case where a result is obtained via an intermediate result is included in addition to a case where a direct result is obtained.

[0236] (13) prohibit, forbid

[0237] Also referred to as not allowed. In addition, not prohibited or allowed does not necessarily mean an obligation.

[0238] (14) limit, restriction, restrict, restricted

[0239] Also referred to as not allowed. In addition, not prohibited or allowed does not necessarily mean an obligation. Also, as long as a part is prohibited in terms of quantity or quality, a case where it is completely prohibited is included.

[0240] (15) chroma

[0241] An adjective denoted by the symbols Cb and Cr, which specifies a sample arrangement or a single sample indicating one of two colour difference signals associated with a primary colour. Also, the term chroma can be replaced with the term chrominance.

[0242] (16) luma

[0243] An adjective denoted by the symbol or subscript Y or L, which specifies a sample arrangement or a single sample indicating a monochrome signal associated with a primary colour. The term luma can be replaced with the term luminance.

[0244] [Related to the description recorded]

[0245] In the drawings, the same reference numbers denote the same or similar constituent elements. In addition, the size and relative positions of the constituent elements in the drawings are not necessarily drawn to scale.

[0246] Hereinafter, the embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below each show an inclusive or specific example. The numerical values, shapes, materials, component configurations, arrangement positions and connection modes of components, steps, relationships and orders of steps, and the like shown in the embodiments below are one example, and are not intended to limit the scope of the claims.

[0247] Hereinafter, the embodiments of the encoding apparatus and the decoding apparatus will be described. The embodiments are examples of the encoding apparatus and the decoding apparatus capable of applying the processes and / or structures described in each aspect of the present application. The processes and / or structures can be implemented in the encoding apparatus and the decoding apparatus different from the embodiments. For example, with respect to the processes and / or structures applied to the embodiments, for example, any one of the following can be performed.

[0248] (1) Any one of the plurality of components of the encoding apparatus or the decoding apparatus of the embodiments described in each aspect of the present application can be replaced with another component described in any one of the aspects of the present application, or a combination thereof;

[0249] (2) In the encoding apparatus or the decoding apparatus of the embodiments, arbitrary changes such as addition, replacement, deletion, and the like of the functions or processes performed by part of the plurality of components of the encoding apparatus or the decoding apparatus can be made. For example, any one of the functions or processes can be replaced with another function or process described in any one of the aspects of the present application, or a combination thereof;

[0250] (3) In the method implemented by the encoding apparatus or the decoding apparatus of the embodiments, arbitrary changes such as addition, replacement, deletion, and the like can be made with respect to part of the plurality of processes included in the method. For example, any one of the processes in the method can be replaced with another process described in any one of the aspects of the present application, or a combination thereof;

[0251] (4) Part of the plurality of components of the encoding apparatus or the decoding apparatus of the embodiments can be combined with a component described in any one of the aspects of the present application, a component having part of the functions described in any one of the aspects of the present application, or a component implementing part of the processes implemented by the components described in any one of the aspects of the present application;

[0252] (5) A configuration element that has a part of a function of the encoding apparatus or the decoding apparatus of the embodiment, or a configuration element that performs a part of a process of the encoding apparatus or the decoding apparatus of the embodiment, is combined with or replaced by a configuration element described in any of the aspects of the present application, a configuration element that has a part of a function described in any of the aspects of the present application, or a configuration element that performs a part of a process described in any of the aspects of the present application.

[0253] (6) In a method performed by the encoding apparatus or the decoding apparatus of the embodiment, a part of a plurality of processes included in the method is replaced by a process described in any of the aspects of the present application or the same process, or a combination thereof.

[0254] (7) A part of a plurality of processes included in a method performed by the encoding apparatus or the decoding apparatus of the embodiment can also be combined with a process described in any of the aspects of the present application.

[0255] (8) The aspects of the present application are not limited to the encoding apparatus or the decoding apparatus of the embodiment. For example, the processes and / or structures can also be implemented in an apparatus used for a purpose different from the motion picture encoding or the motion picture decoding disclosed in the embodiment.

[0256] [system configuration]

[0257] Figure 1 is a schematic diagram showing an example of a configuration of a transmission system of the embodiment.

[0258] The transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, as shown in Figure 1 , an encoding apparatus 100, a network Nw, and a decoding apparatus 200.

[0259] An image is input to the encoding apparatus 100. The encoding apparatus 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, an encoded image and control information for decoding the encoded image. The image is compressed by the encoding.

[0260] Further, the original image before encoding input to the encoding apparatus 100 is also referred to as an original image, an original signal, or an original sample. In addition, the image can be a moving image or a still image. In addition, the image is a higher concept of a sequence, a picture, a block, and the like, and is not limited by a region in space and time unless otherwise specified. In addition, the image is constituted by an arrangement of pixels or pixel values, and a signal or a pixel value representing the image is also referred to as a sample. In addition, the stream can be referred to as a bit stream, an encoded bit stream, a compressed bit stream, or an encoded signal. Furthermore, the encoding apparatus can also be referred to as an image encoding apparatus or a moving image encoding apparatus, and the encoding method of the encoding apparatus 100 can also be referred to as an encoding method, an image encoding method, or a moving image encoding method.

[0261] The network Nw transmits the stream generated by the encoding apparatus 100 to the decoding apparatus 200. The network Nw can be the Internet, a wide area network (WAN), a small-scale network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network, and can be a unidirectional communication network that transmits a broadcast wave such as terrestrial digital broadcasting or satellite broadcasting. In addition, the network Nw can also be replaced by a storage medium in which a stream of a DVD (Digital Versatile Disc), a BD (Blu-Ray Disc (registered trademark)), or the like is recorded.

[0262] The decoding apparatus 200 generates, for example, a decoded image that is a non-compressed image by decoding the stream transmitted by the network Nw. For example, the decoding apparatus decodes the stream in accordance with a decoding method corresponding to the encoding method of the encoding apparatus 100.

[0263] In addition, the decoding apparatus can also be referred to as an image decoding apparatus or a moving image decoding apparatus, and the decoding method of the decoding apparatus 200 can also be referred to as a decoding method, an image decoding method, or a moving image decoding method.

[0264] [Data structure]

[0265] Figure 2 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. The video sequence includes, for example, as shown in (a) of FIG. 13, a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and a plurality of pictures. Figure 2 ​

[0266] The VPS includes encoding parameters common to a plurality of layers in a moving image composed of the plurality of layers, and encoding parameters associated with each layer or the plurality of layers included in the moving image.

[0267] The SPS includes parameters used for a sequence, that is, encoding parameters referred to by the decoding device 200 in order to decode the sequence. For example, the encoding parameters can also indicate the width or height of a picture. In addition, there can be a plurality of SPSs.

[0268] The PPS includes parameters used for a picture, that is, encoding parameters referred to by the decoding device 200 in order to decode each picture within a sequence. For example, the encoding parameters can include a reference value of the quantization width used in decoding of a picture, and a flag indicating the application of weighted prediction. In addition, there can be a plurality of PPSs. Furthermore, the SPS and the PPS are sometimes referred to as a parameter set.

[0269] As shown in (b) of FIG. 1, a picture can include a picture header and one or more slices. The picture header includes encoding parameters referred to by the decoding device 200 in order to decode the one or more slices. Figure 2 As shown in (c) of FIG. 1, a slice includes a slice header and one or more tiles. The slice header includes encoding parameters referred to by the decoding device 200 in order to decode the one or more tiles.

[0270] Figure 2 As shown in (d) of FIG. 1, a tile includes one or more CTUs (Coding Tree Units).

[0271] As shown in (e) of FIG. 1, such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes encoding parameters referred to by the decoding device 200 in order to decode the one or more CUs. Figure 2 A CU can also be divided into a plurality of small CUs. In addition, as shown in (f) of FIG. 1, such a CU includes a CU header and one or more PUs (Prediction Units).

[0272]

[0273] Figure 2

[0274] Figure 2 ​​​​​The CU includes a CU header, prediction information, and residual coefficient information as shown in (f). The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual described later. Furthermore, the CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but can include a plurality of TUs smaller than the CU, for example, in SBT described later. In addition, the CU can also process each VPDU (Virtual Pipeline Decoding Unit) constituting the CU. The VPDU is, for example, a fixed unit that can be processed in one stage when pipeline processing is performed in hardware.

[0275] Furthermore, a stream can also not have Figure 2 The order of these layers can be exchanged, and any one of the layers can be replaced with another layer. In addition, a picture that is the object of processing by the apparatus 100 or the apparatus 200 or the like at the current time point is referred to as a current picture. If the processing is encoding, the current picture is synonymous with an encoding target picture, and if the processing is decoding, the current picture is synonymous with a decoding target picture. In addition, a block such as a CU or the like that is the object of processing by the apparatus 100 or the apparatus 200 or the like at the current time point is referred to as a current block. If the processing is encoding, the current block is synonymous with an encoding target block, and if the processing is decoding, the current block is synonymous with a decoding target block.

[0276] [Structure of picture Slice / Tile]

[0277] In order to decode a picture in parallel, a picture is sometimes constituted by a slice unit or a tile unit.

[0278] A slice is a basic encoding unit constituting a picture. A picture is constituted by one or more slices, for example. In addition, a slice is constituted by one or more continuous CTUs.

[0279] Figure 3is a diagram showing an example of a structure of a slice. For example, a picture contains 11 x 8 CTUs, and is divided into 4 slices (slices 1 to 4). Slice 1 is composed of, for example, 16 CTUs, slice 2 is composed of, for example, 21 CTUs, slice 3 is composed of, for example, 29 CTUs, and slice 4 is composed of, for example, 22 CTUs. Here, each CTU within the picture belongs to a certain slice. The shape of the slice becomes a shape in which the picture is divided in the horizontal direction. The boundary of the slice need not be the picture end, and can be somewhere in the boundary of the CTU within the picture. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, a raster scan order. In addition, the slice contains a slice header and encoded data. In the slice header, the CTU address of the beginning of the slice, the slice type, and the like, which are characteristics of the slice, can be described.

[0280] A tile is a unit of a rectangular region constituting a picture. Each tile can be assigned a number called Tileld in a raster scan order.

[0281] Figure 4 is a diagram showing an example of a structure of a tile. For example, a picture contains 11 x 8 CTUs, and is divided into 4 rectangular region tiles (tiles 1 to 4). In the case of using tiles, the processing order of the CTUs is changed compared to the case of not using tiles. In the case of not using tiles, a plurality of CTUs within the picture are processed, for example, in a raster scan order. In the case of using tiles, in each of a plurality of tiles, at least one CTU is processed, for example, in a raster scan order. For example, as shown in Figure 4 , the processing order of a plurality of CTUs contained in tile 1 is an order from the left end of the 1st column of tile 1 toward the right end of the 1st column of tile 1, and then from the left end of the 2nd column of tile 1 toward the right end of the 2nd column of tile 1.

[0282] In addition, sometimes one tile contains one or more slices, and sometimes one slice contains one or more tiles.

[0283] Further, a picture can also be constituted by a tile set unit. A tile set can contain one or more tile groups, and can contain one or more tiles. A picture can be constituted only by one of a tile set, a tile group, and a tile. For example, an order in which a plurality of tiles are scanned in a raster order for each tile set is set as a basic encoding order of the tile. A set of one or more tiles in which the basic encoding order is continuous within each tile set is set as a tile group. Such a picture can also be constituted by the division section 102 (refer to Figure 7 ) described later.

[0284] [scalable encoding]

[0285] Figure 5 and Figure 6is an example of a structure of a scalable stream.

[0286] As shown in Figure 5 , the encoding apparatus 100 can generate a temporally / spatially scalable stream by encoding a plurality of pictures separately into a certain layer of a plurality of layers. For example, the encoding apparatus 100 realizes scalability in which an enhancement layer exists in a higher level than a base layer by encoding the pictures per layer. Such encoding of each picture is called scalable encoding. Thus, the decoding apparatus 200 can switch the quality of an image displayed by decoding the stream. That is, the decoding apparatus 200 decides which layer to decode according to an internal factor such as its own performance and an external factor such as the state of a communication band. As a result, the decoding apparatus 200 can freely switch the same content to a low-resolution content and a high-resolution content to decode. For example, a user of the stream is moving, uses a smartphone to listen to a moving image of the stream halfway, and after returning home, uses an Internet TV or the like to listen to the latter part of the moving image. In addition, the decoding apparatus 200 having the same or different performance is assembled in each of the smartphone and the device described above. In this case, if the device decodes a higher layer in the stream, the user can listen to a high-quality moving image after returning home. Thus, the encoding apparatus 100 does not need to generate a plurality of streams having the same content but different qualities, and can reduce the processing load.

[0287] Further, the enhancement layer can also include meta information based on statistical information of the image or the like. The decoding apparatus 200 can generate a high-quality moving image by super-resolution of the picture of the base layer based on the meta information. The super-resolution can be one of improving the SN ratio in the same resolution and expanding the resolution. The meta information includes information for determining a filter coefficient of linearity or non-linearity used in the super-resolution processing, or information for determining a parameter value of a filter processing, machine learning, or minimum 2 multiplication used in the super-resolution processing, and the like.

[0288] Alternatively, the picture can be divided into tiles or the like according to the meaning of each object or the like in the picture. In this case, the decoding apparatus 200 can decode only a part of the area in the picture by selecting a tile as an object of decoding. Further, the attribute of the object (person, car, ball, or the like) and the position in the picture (coordinate position in the same picture or the like) can be saved as meta information. In this case, the decoding apparatus 200 can determine the position of the desired object based on the meta information, and decide a tile including the object. For example, as shown in Figure 6 , the meta information can be saved using a data saving structure different from pixel data such as SEI in HEVC. The meta information indicates, for example, the position, size, or color of the main object, or the like.

[0289] In addition, the meta information can be stored in units of a plurality of pictures such as a stream, a sequence, or a random access unit. Thus, the decoding apparatus 200 can acquire a time at which a specific person appears in a moving image, and by using the time and information of a picture unit, can determine a picture in which the target exists and a position of the target in the picture.

[0290] [Encoding apparatus]

[0291] Next, an encoding apparatus 100 according to the embodiment will be described. Figure 7 is a block diagram showing an example of a functional structure of the encoding apparatus 100 according to the embodiment. The encoding apparatus 100 encodes an image in a block unit.

[0292] As shown in Figure 7 , the encoding apparatus 100 is an apparatus that encodes an image in a block unit, and includes a division section 102, a subtraction section 104, a transform section 106, a quantization section 108, an entropy encoding section 110, an inverse quantization section 112, an inverse transform section 114, an addition section 116, a block memory 118, a loop filter section 120, a frame memory 122, an intra prediction section 124, an inter prediction section 126, a prediction control section 128, and a prediction parameter generation section 130. Further, the intra prediction section 124 and the inter prediction section 126 each constitute a part of a prediction processing section.

[0293] [Mounting example of encoding apparatus]

[0294] Figure 8 is a block diagram showing a mounting example of the encoding apparatus 100. The encoding apparatus 100 includes a processor al and a memory a2. For example, Figure 7 , a plurality of constituent elements of the encoding apparatus 100 are mounted by Figure 8 the processor al and the memory a2 shown in

[0295] The processor al is a circuit that performs information processing, and is a circuit that can access the memory a2. For example, the processor al is a dedicated or general electronic circuit that encodes an image. The processor al can also be a processor such as a CPU. In addition, the processor al can be a collection of a plurality of electronic circuits. In addition, for example, the processor al can function as a plurality of constituent elements of the encoding apparatus 100 shown in Figure 7 , except for a plurality of constituent elements for storing information.

[0296] The memory a2 is a dedicated or general-purpose memory that stores information used by the processor a1 to encode the image. The memory a2 can be an electronic circuit or can be connected to the processor a1. Alternatively, the memory a2 can be included in the processor a1. Further, the memory a2 can be a collection of a plurality of electronic circuits. Further, the memory a2 can be a magnetic disk or an optical disk or the like, or can be a storage or a recording medium or the like. Further, the memory a2 can be a nonvolatile memory or a volatile memory.

[0297] For example, the memory a2 can store an encoded image or can store a stream corresponding to the encoded image. Alternatively, a program used by the processor a1 to encode the image can be stored in the memory a2.

[0298] Further, for example, the memory a2 can function as Figure 7 the memory a2 can function as Figure 7 the block memory 118 and the frame memory 122 illustrated in FIG. 1. More specifically, the reconstructed image (specifically, a reconstructed block and a reconstructed picture, or the like) can be stored in the memory a2.

[0299] Further, in the encoding apparatus 100, all of the plurality of components illustrated in FIG. 1 can not be installed, and all of the plurality of processes described above can not be performed. Figure 7 Further, in the encoding apparatus 100, all of the plurality of components illustrated in FIG. 1 can not be installed, and all of the plurality of processes described above can not be performed. Figure 7 Further, in the encoding apparatus 100, all of the plurality of components illustrated in FIG. 1 can not be installed, and all of the plurality of processes described above can not be performed.

[0300] Hereinafter, after the flow of the overall process of the encoding apparatus 100 is described, each component included in the encoding apparatus 100 is described.

[0301] [Overall flow of encoding process]

[0302] Figure 9 is a flowchart showing an example of the overall encoding process performed by the encoding apparatus 100.

[0303] First, the division section 102 of the encoding apparatus 100 divides a picture included in an original image into a plurality of blocks having a fixed size (128 x 128 pixels) (step Sa_1). Then, the division section 102 selects a division pattern for the block having the fixed size (step Sa_2). That is, the division section 102 further divides the block having the fixed size into a plurality of blocks constituting the selected division pattern. Then, the encoding apparatus 100 performs the processes of steps Sa_3 to Sa_9 for each of the plurality of blocks.

[0304] The prediction processing section composed of the intra prediction section 124 and the inter prediction section 126, and the prediction control section 128 generate a prediction image of the current block (step Sa_3). Further, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.

[0305] Next, the subtraction section 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual is also referred to as a prediction error.

[0306] Next, the transform section 106 and the quantization section 108 generate a plurality of quantized coefficients by performing transform and quantization on the prediction residual (step Sa_5).

[0307] Next, the entropy coding section 110 generates a stream by performing coding (specifically, entropy coding) on the plurality of quantized coefficients and a prediction parameter related to the generation of the prediction image (step Sa_6).

[0308] Next, the inverse quantization section 112 and the inverse transform section 114 reproduce the prediction residual by performing inverse quantization and inverse transform on the plurality of quantized coefficients (step Sa_7).

[0309] Next, the addition section 116 reconstructs the current block by adding the reproduced prediction residual to the prediction image (step Sa_8). Thereby, a reconstructed image is generated. Further, the reconstructed image is also referred to as a reconstructed block, and specifically, the reconstructed image generated by the encoding apparatus 100 is also referred to as a local decoded block or a local decoded image.

[0310] When the reconstructed image is generated, the loop filter section 120 filters the reconstructed image as necessary (step Sa_9).

[0311] Then, the encoding apparatus 100 determines whether or not the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (NO in step Sa_10), the processing from step Sa_2 is repeated.

[0312] In addition, in the example described above, the encoding apparatus 100 selects one partitioning style for the blocks of a fixed size, and performs the encoding of each block in accordance with the partitioning style, but it is also possible to perform the encoding of each block in accordance with each of a plurality of partitioning styles. In this case, the encoding apparatus 100 can evaluate the cost for each of the plurality of partitioning styles, and for example, can select a stream obtained by performing the encoding in accordance with the partitioning style of the smallest cost as the final output stream.

[0313] Further, the processing of these steps Sa_1 to Sa_10 can be performed sequentially by the encoding apparatus 100, a part of the plurality of processing among these processing can be performed in parallel, or the order can be changed.

[0314] The encoding process of such an encoding apparatus 100 is hybrid encoding using prediction encoding and transform encoding. Furthermore, the prediction encoding is performed by an encoding loop constituted by the subtraction section 104, the transform section 106, the quantization section 108, the inverse quantization section 112, the inverse transform section 114, the addition section 116, the loop filter section 120, the block memory 118, the frame memory 122, the intra prediction section 124, the inter prediction section 126, and the prediction control section 128. That is, the prediction processing section constituted by the intra prediction section 124 and the inter prediction section 126 constitutes a part of the encoding loop.

[0315] [DIVISION SECTION]

[0316] The division section 102 divides each picture included in the original image into a plurality of blocks and outputs each block to the subtraction section 104. For example, the division section 102 first divides a picture into blocks of a fixed size (e.g., 128 x 128 pixels). Such a block of a fixed size is referred to as a coding tree unit (CTU). Furthermore, the division section 102 divides each block of a fixed size into blocks of a variable size (e.g., 64 x 64 pixels or less) based on, for example, a recursive quadtree and / or binary tree block division. That is, the division section 102 selects a division pattern. Such a block of a variable size is referred to as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). Note that in various installation examples, the CU, the PU, and the TU are not necessarily distinguished, and a part or all of the blocks within a picture can be processed as a CU, a PU, or a TU.

[0317] Figure 10 FIG. 1 is a diagram illustrating an example of block division according to an embodiment. In FIG. 1, a solid line indicates a block boundary based on quadtree block division, and a dashed line indicates a block boundary based on binary tree block division. Figure 10

[0318] Here, the block 10 is a square block of 128 x 128 pixels. The block 10 is first divided into four square blocks of 64 x 64 pixels (quadtree block division).

[0319] The square block of 64 x 64 pixels on the left upper side is further vertically divided into two rectangular blocks each of which is constituted by 32 x 64 pixels, and the rectangular block of 32 x 64 pixels on the left side is further vertically divided into two rectangular blocks each of which is constituted by 16 x 64 pixels (binary tree block division). As a result, the square block of 64 x 64 pixels on the left upper side is divided into two rectangular blocks 11 and 12 each of which is constituted by 16 x 64 pixels, and a rectangular block 13 of 32 x 64 pixels.

[0320] The square block of 64 x 64 pixels on the right upper side is horizontally divided into two rectangular blocks 14 and 15 each of which is constituted by 64 x 32 pixels (binary tree block division).​

[0321] The 64x64-pixel square block on the lower left is divided into four 32x32-pixel square blocks (quad-tree block division). The upper left and lower right of the four 32x32-pixel square blocks are further divided. The 32x32-pixel square block on the upper left is vertically divided into two 16x32-pixel rectangular blocks, and the 16x32-pixel rectangular block on the right is further horizontally divided into two 16x16-pixel square blocks (binary-tree block division). The 32x32-pixel square block on the lower right is horizontally divided into two 32x16-pixel rectangular blocks (binary-tree block division). As a result, the 64x64-pixel square block on the lower left is divided into a 16x32-pixel rectangular block 16, two 16x16-pixel square blocks 17 and 18 each, two 32x32-pixel square blocks 19 and 20 each, and two 32x16-pixel rectangular blocks 21 and 22 each.

[0322] The 64x64-pixel block 23 on the lower right is not divided.

[0323] As above, in the example of FIG. 1, the block 10 is divided into 13 variable-size blocks 11 to 23 based on the recursive quad-tree and binary-tree block division. Such division is referred to as QTBT (quad-tree plus binary tree) division. Figure 10 In addition, in the example of FIG. 1, one block is divided into four or two blocks (quad-tree or binary-tree block division), but the division is not limited to these. For example, one block can be divided into three blocks (ternary-tree division). Division including such ternary-tree division is referred to as MBT (multi type tree) division.

[0324] Figure 10 FIG. 1 is a diagram showing an example of a functional structure of the division section 102. As shown in FIG. 1, the division section 102 can include a block division determination section 102a. As an example, the block division determination section 102a can perform the following processing.

[0325] Figure 11 The block division determination section 102a collects block information, for example, from the block memory 118 or the frame memory 122, and determines the division pattern described above based on the block information. The division section 102 divides the original image in accordance with the division pattern, and outputs one or more blocks obtained by the division to the subtraction section 104. Figure 11

[0326] The block division determination section 102a collects block information, for example, from the block memory 118 or the frame memory 122, and determines the division pattern described above based on the block information. The division section 102 divides the original image in accordance with the division pattern, and outputs one or more blocks obtained by the division to the subtraction section 104.

[0327] ​​Further, the block division decision section 102a outputs, for example, a parameter indicating the above-described division pattern to the transform section 106, the inverse transform section 114, the intra prediction section 124, the inter prediction section 126, and the entropy coding section 110. The transform section 106 can transform the prediction residual based on the parameter, the intra prediction section 124 and the inter prediction section 126 can generate the prediction image based on the parameter. In addition, the entropy coding section 110 can also entropy code the parameter.

[0328] As an example, the parameter related to the division pattern can also be written in the stream as follows.

[0329] Figure 12 is a diagram indicating an example of the division pattern. The division pattern has, for example, quadtree (QT) that divides the block into 2 in the horizontal direction and the vertical direction, ternary tree (HT or VT) that divides the block in the same direction at a ratio of 1 to 2 to 1, binary tree (HB or VB) that divides the block in the same direction at a ratio of 1 to 1, and no split (NS).

[0330] In addition, in the case of the quadtree and the no split, the division pattern does not have the block division direction, and in the case of the binary tree and the ternary tree, the division pattern has the division direction information.

[0331] Figure 13A and Figure 13B is a diagram indicating an example of the syntax tree of the division pattern. In Figure 13A , first, information indicating whether or not to perform the division (S: Split flag) is present first, and then information indicating whether or not to perform the quadtree division (QT: QT flag) is present. Then, information indicating whether to perform the ternary tree division or the binary tree division (TT: TT flag or BT: BT flag) is present, and finally information indicating the division direction (Ver: Vertical flag or Hor: Horizontal flag) is present. In addition, the division can be further repeatedly applied to each of one or more blocks obtained by the division based on the division pattern by the same processing. That is, as an example, the determination of whether or not to perform the division, whether or not to perform the quadtree division, whether the division method is the horizontal direction or the vertical direction, and whether to perform the ternary tree division or the binary tree division can also be implemented recursively, and the results of the implemented determinations can be encoded in the stream in the order disclosed by the syntax tree shown in Figure 13A .

[0332] In addition, in the syntax tree shown in Figure 13A , the information is arranged in the order of S, QT, TT, and Ver, but the information can also be arranged in the order of S, QT, Ver, and BT. That is, in Figure 13BIn the example of FIG. 9, first, there is information indicating whether or not division is performed (S: Split flag), and then, there is information indicating whether or not quad-division is performed (QT: QT flag). Next, there is information indicating the division direction (Ver: Vertical flag or Hor: Horizontal flag), and finally, there is information indicating whether or not bi-division is performed or tri-division is performed (BT: BT flag or TT: TT flag).

[0333] In addition, the division pattern described herein is an example, and a division pattern other than the division pattern described can be used, or only a part of the division pattern described can be used.

[0334] [Subtracting section]

[0335] The subtracting section 104 subtracts the prediction image (prediction image input from the prediction control section 128) from the original image in units of blocks input from and divided by the dividing section 102. That is, the subtracting section 104 calculates the prediction residual of the current block. Also, the subtracting section 104 outputs the calculated prediction residual to the transforming section 106.

[0336] The original image is an input signal of the encoding apparatus 100, and is, for example, a signal (e.g., a luma signal and two chroma signals) indicating an image of each picture constituting a moving image.

[0337] [Transforming section]

[0338] The transforming section 106 transforms the prediction residual in the spatial domain into a transform coefficient in the frequency domain, and outputs the transform coefficient to the quantizing section 108. Specifically, the transforming section 106, for example, performs a predetermined discrete cosine transform (DCT) or a discrete sine transform (DST) on the prediction residual in the spatial domain.

[0339] In addition, the transforming section 106 can also adaptively select a transform type from among a plurality of transform types, and transform the prediction residual into a transform coefficient using a transform basis function corresponding to the selected transform type. Such a transform is referred to as EMT (explicit multiple core transform) or AMT (adaptive multiple transform). Furthermore, the transform basis function is sometimes simply referred to as a basis.

[0340] The plurality of transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Furthermore, these transform types can be respectively denoted as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14is a table indicating a transform basis function corresponding to each transform type. In Figure 14 In the above, N indicates the number of input pixels. The selection of the transform type from among these multiple transform types can be dependent on, for example, the kind of prediction (intra prediction and inter prediction, etc.) or the intra prediction mode.

[0341] Information indicating whether to apply such EMT or AMT (for example, referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. In addition, the signaling of these pieces of information need not be limited to the CU level, and can be at another level (for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level).

[0342] Further, the transform unit 106 can also perform a re-transformation on the transform coefficients (i.e., the transform result). Such a re-transformation is referred to as an AST (adaptive secondary transform) or an NSST (non-separable secondary transform). For example, the transform unit 106 performs a re-transformation on each sub-block (for example, a 4 x 4 pixel sub-block) included in a block of transform coefficients corresponding to the intra prediction residual. Information indicating whether to apply the NSST and information related to the transform matrix used in the NSST are generally signaled at the CU level. In addition, the signaling of these pieces of information need not be limited to the CU level, and can be at another level (for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level).

[0343] The transform unit 106 can apply both a separable transform and a non-separable transform. The separable transform refers to a manner in which a plurality of transforms are performed in each direction separately corresponding to the number of dimensions of the input, and the non-separable transform refers to a manner in which 2 or more dimensions are considered as one dimension and a transform is performed on the whole.

[0344] For example, as an example of the non-separable transform, there is a manner in which, in the case where the input is a block of 4 x 4 pixels, the block is considered as one arrangement having 16 elements, and a transform process is performed on the arrangement with a 16 x 16 transform matrix.

[0345] Further, in a further example of the non-separable transform, a transform in which Givens rotation is performed a plurality of times on the arrangement after the input block of 4 x 4 pixels is considered as one arrangement having 16 elements (Hypercube Givens Transform) can also be performed.

[0346] In the transform in the transform section 106, the transform type of the transform basis function to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is an SVT (Spatially Varying Transform).

[0347] Figure 15 is a diagram showing an example of the SVT.

[0348] In the SVT, as shown in Figure 15 , the CU is bisected in the horizontal direction or the vertical direction, and the transform into the frequency domain is performed only on the region of either side. The transform type can be set for each region, and, for example, DST7 and DCT8 are used. For example, for the region of position 0 in the 2 regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used. Alternatively, for the region of position 1 in the 2 regions, DST7 is used. Similarly, for the region of position 0 in the 2 regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 are used. Alternatively, for the region of position 1 in the 2 regions, DST7 is used. In such Figure 15 , in the example shown, the transform is performed only on one of the 2 regions within the CU, and the other is not transformed, but the transform can be performed on both of the 2 regions. In addition, the division method is not only bisected, but can also be quartered. Furthermore, it is also possible to be more flexible, and encode the information showing the division method, and signalize the same as the CU division, and the like. In addition, the SVT is sometimes referred to as SBT (Sub-block Transform).

[0349] The aforementioned AMT and EMT can also be referred to as MTS (Multiple Transform Selection). In the case where MTS is applied, a transform type such as DST7 or DCT8 can be selected, and information indicating the selected transform type can be encoded as index information for each CU. On the other hand, as a process of selecting a transform type used in orthogonal transform on a CU basis without encoding the index information, there is a process referred to as IMTS (Implicit MTS). In the case where IMTS is applied, for example, if the shape of a CU is rectangular, DST7 is used on the short side of the rectangle, and DCT2 is used on the long side, and orthogonal transform is performed on each. Also, for example, in the case where the shape of a CU is square, if MTS is effective within a sequence, orthogonal transform is performed using DCT2, and if MTS is not effective, orthogonal transform is performed using DST7. DCT2 and DST7 are examples, and other transform types can be used, and the combination of the transform types used can be set to different combinations. IMTS can be available only in an intra prediction block, or can be available together in an intra prediction block and an inter prediction block.

[0350] In the above, as a selection process of selectively switching a transform type used in orthogonal transform, three processes of MTS, SBT, and IMTS are described, but all of the three selection processes can be effective, or only a part of the selection processes can be selectively made effective. As to whether each selection process is effective, it can be recognized by flag information and the like within a header such as SPS. For example, if all of the three selection processes are effective, one is selected from the three selection processes in a CU unit to perform orthogonal transform. Also, as long as a selection process of selectively switching a transform type can achieve at least one of the following four functions [1] to [4], a different selection process from the above three selection processes can be used, or the above three selection processes can be replaced by other processes, respectively. Function [1] is a function of performing orthogonal transform on the entire range within a CU, and encoding information indicating a transform type used in the transform. Function [2] is a function of performing orthogonal transform on the entire range of a CU, and determining a transform type based on a prescribed rule without encoding information indicating the transform type. Function [3] is a function of performing orthogonal transform on a region of a part of a CU, and encoding information indicating a transform type used in the transform. Function [4] is a function of performing orthogonal transform on a region of a part of a CU, and determining a transform type based on a prescribed rule without encoding information indicating a transform type used in the transform, and the like.

[0351] Also, whether each of MTS, IMTS, and SBT is applied or not can be determined on a per-process unit basis. For example, whether to apply or not can be determined in a sequence unit, a picture unit, a tile unit, a slice unit, a CTU unit, or a CU unit.

[0352] In addition, the tool that selectively switches the transform type in the present application can also be called a method of adaptively selecting a basis used in the transform processing, a selection process, or a process of selecting a basis. In addition, the tool that selectively switches the transform type can also be called a mode of adaptively selecting the transform type.

[0353] Figure 16 is a flowchart showing an example of the processing by the transform unit 106.

[0354] For example, the transform unit 106 determines whether or not to perform the orthogonal transform (step St_1). Here, when it is determined to perform the orthogonal transform (Yes in step St_1), the transform unit 106 selects a transform type used for the orthogonal transform from among a plurality of transform types (step St_2). Next, the transform unit 106 performs the orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). Then, the transform unit 106 outputs information indicating the selected transform type to the entropy encoding unit 110, and thereby encodes the information (step St_4). On the other hand, when it is determined not to perform the orthogonal transform (No in step St_1), the transform unit 106 outputs information indicating that the orthogonal transform is not performed to the entropy encoding unit 110, and thereby encodes the information (step St_5). Further, the determination of whether or not to perform the orthogonal transform in step St_1 can be determined, for example, on the basis of the size of the transform block, the prediction mode applied to the CU, and the like. In addition, the information indicating the transform type used for the orthogonal transform is not encoded, and the orthogonal transform can be performed using a transform type specified in advance.

[0355] Figure 17 is a flowchart showing another example of the processing by the transform unit 106. In addition, Figure 17 The example shown in Figure 16 is an example of the orthogonal transform in the case where the method of selectively switching the transform type used for the orthogonal transform is applied, as with the example shown in

[0356] As an example, the first transform type group can include DCT2, DST7, and DCT8. In addition, as an example, the second transform type group can include DCT2. In addition, the transform types included in the first transform type group and the second transform type group can be partially repeated or can be all different transform types.

[0357] Specifically, the transform unit 106 determines whether the transform size is equal to or smaller than a prescribed value (step Su_1). Here, when it is determined that the transform size is equal to or smaller than the prescribed value (Yes in step Su_1), the transform unit 106 performs orthogonal transform on the prediction residual of the current block using a transform type included in the first transform type group (step Su_2). Next, the transform unit 106 outputs information indicating which one of one or more transform types included in the first transform type group is used to the entropy encoding unit 110, and encodes the information (step Su_3). On the other hand, when it is determined that the transform size is not equal to or smaller than the prescribed value (No in step Su_1), the transform unit 106 performs orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4).

[0358] In step Su_3, the information indicating the transform type used for the orthogonal transform can be information indicating a combination of a transform type applied to the vertical direction and a transform type applied to the horizontal direction of the current block. In addition, the first transform type group can include only one transform type, and the information indicating the transform type used for the orthogonal transform can not be encoded. The second transform type group can include a plurality of transform types, and the information indicating the transform type used in the orthogonal transform among one or more transform types included in the second transform type group can be encoded.

[0359] In addition, the transform type can be determined based only on the transform size. Furthermore, if the transform type used for the orthogonal transform is determined based on the transform size, the determination is not limited to whether the transform size is equal to or smaller than the prescribed value.

[0360] [Quantization unit]

[0361] The quantization unit 108 quantizes the transform coefficient output from the transform unit 106. Specifically, the quantization unit 108 scans a plurality of transform coefficients of the current block in a prescribed scan order, and quantizes the transform coefficients based on a quantization parameter (QP) corresponding to the scanned transform coefficients. Furthermore, the quantization unit 108 outputs the quantized plurality of transform coefficients (hereinafter referred to as quantized coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0362] The prescribed scan order is an order for quantization / inverse quantization of the transform coefficients. For example, the prescribed scan order is defined by an ascending order of frequency (an order from low frequency to high frequency) or a descending order of frequency (an order from high frequency to low frequency).

[0363] The quantization parameter (QP) refers to a parameter defining a quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the error of the quantized coefficients (quantization error) increases.

[0364] In addition, in quantization, a quantization matrix is sometimes used. For example, a plurality of quantization matrices are sometimes used in correspondence with a frequency transform size such as 4 x 4 and 8 x 8, a prediction mode such as intra prediction and inter prediction, a pixel component such as luminance and chrominance, and the like. In addition, quantization refers to digitization of a value sampled at a predetermined interval in correspondence with a predetermined level, and in this technical field, rounding, rounding, or scaling is sometimes used.

[0365] As a method of using a quantization matrix, there is a method of using a quantization matrix directly set on the encoding device 100 side and a method of using a default quantization matrix (default matrix). On the encoding device 100 side, by directly setting a quantization matrix, it is possible to set a quantization matrix corresponding to the characteristics of an image. However, in this case, there is a disadvantage that the amount of encoding increases due to encoding of the quantization matrix. In addition, instead of directly using a default quantization matrix or an encoded quantization matrix, it is also possible to generate a quantization matrix used in quantization of the current block based on a default quantization matrix or an encoded quantization matrix.

[0366] On the other hand, there is also a method of quantizing in such a manner that the coefficients of high frequency components and the coefficients of low frequency components are all the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix in which all the coefficients are the same value (flat matrix).

[0367] The quantization matrix can be encoded, for example, at the sequence level, the picture level, the slice level, the tile level, or the CTU level.

[0368] The quantization section 108 scales, for example, the quantization width or the like calculated based on the quantization parameter or the like by the value of the quantization matrix for each transform coefficient in the case of using the quantization matrix. The quantization processing performed without using the quantization matrix can also be processing of quantizing the transform coefficient based on the quantization width calculated based on the quantization parameter or the like. Furthermore, in the quantization processing performed without using the quantization matrix, it is also possible to multiply the quantization width by a predetermined value common to all the transform coefficients within the block.

[0369] Figure 18 is a block diagram that shows an example of the functional structure of the quantization section 108.

[0370] The quantization section 108, for example, has a difference quantization parameter generation section 108a, a prediction quantization parameter generation section 108b, a quantization parameter generation section 108c, a quantization parameter storage section 108d, and a quantization processing section 108e.

[0371] Figure 19 is a flowchart that shows an example of quantization performed by the quantization section 108.

[0372] As an example, the quantization section 108 can generate a quantization matrix based on the quantization parameter calculated by the quantization parameter generation section 108c.Figure 19 The flowchart shown performs quantization per CU. Specifically, the quantization parameter generation section 108c determines whether or not to perform quantization (step Sv l). Here, when it is determined to perform quantization (Yes in step Sv l), the quantization parameter generation section 108c generates a quantization parameter of the current block (step Sv_2), and saves the quantization parameter to the quantization parameter storage section 108d (step Sv_3).

[0373] Next, the quantization processing section 108e quantizes the transform coefficients of the current block using the quantization parameter generated in step Sv_2 (step Sv_4). Then, the prediction quantization parameter generation section 108b acquires a quantization parameter of a processing unit other than the current block from the quantization parameter storage section 108d (step Sv_5). The prediction quantization parameter generation section 108b generates a prediction quantization parameter of the current block based on the acquired quantization parameter (step Sv_6). The difference quantization parameter generation section 108a calculates a difference between the quantization parameter of the current block generated by the quantization parameter generation section 108c and the prediction quantization parameter of the current block generated by the prediction quantization parameter generation section 108b (step Sv_7). By this calculation of the difference, a difference quantization parameter is generated. The difference quantization parameter generation section 108a outputs the difference quantization parameter to the entropy encoding section 110, and thereby encodes the difference quantization parameter (step Sv_8).

[0374] Further, the difference quantization parameter can also be encoded at a sequence level, a picture level, a slice level, a tile level, or a CTU level. Further, an initial value of the quantization parameter can be encoded at a sequence level, a picture level, a slice level, a tile level, or a CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the difference quantization parameter.

[0375] Further, the quantization section 108 can have a plurality of quantizers, and dependent quantization in which a quantization method selected from a plurality of quantization methods is used to quantize the transform coefficients can also be applied.

[0376] [Entropy encoding section]

[0377] Figure 20 is a block diagram showing an example of a functional structure of the entropy encoding section 110.

[0378] The entropy coding section 110 performs entropy coding with respect to the quantization coefficients input from the quantization section 108 and the prediction parameters input from the prediction parameter generation section 130, thereby generating a stream. In this entropy coding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy coding section 110 has, for example, a binarization section 110a, a context control section 110b, and a binary arithmetic coding section 110c. The binarization section 110a performs binarization of converting a multi-valued signal such as the quantization coefficients and the prediction parameters into a binary signal. The manner of binarization is, for example, Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, or the like. The context control section 110b derives a context value, that is, a probability of occurrence of a binary signal, corresponding to a feature of a syntax element or a situation of the surroundings. In the method of deriving the context value, there are, for example, bypass, syntax element reference, upper / left neighboring block reference, level information reference, and others. The binary arithmetic coding section 110c performs arithmetic coding of the binarized signal using the derived context value.

[0379] Figure 21 is a diagram showing a flow of CABAC in the entropy coding section 110.

[0380] First, in CABAC in the entropy coding section 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding section 110c and setting of an initial context value are performed. Then, the binarization section 110a and the binary arithmetic coding section 110c, for example, sequentially perform binarization and arithmetic coding with respect to a plurality of quantization coefficients of a CTU, respectively. At this time, the context control section 110b performs updating of the context value each time arithmetic coding is performed. Then, the context control section 110b causes the context value to be backed up as a post-process. This backed-up context value is used, for example, as an initial value of the context value with respect to the next CTU.

[0381] [Inverse quantization section]

[0382] The inverse quantization section 112 performs inverse quantization of the quantization coefficients input from the quantization section 108. Specifically, the inverse quantization section 112 performs inverse quantization of the quantization coefficients of the current block in a prescribed scan order. Also, the inverse quantization section 112 outputs the inverse-quantized transform coefficients of the current block to the inverse transform section 114.

[0383] [Inverse transform section]

[0384] The inverse transform unit 114 restores the prediction residual by performing inverse transform on the transform coefficient input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual of the current block by performing inverse transform on the transform coefficient corresponding to the transform performed by the transform unit 106. Further, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.

[0385] In addition, the restored prediction residual generally does not coincide with the prediction error calculated by the subtraction unit 104 because information is lost by quantization. That is, the restored prediction residual generally includes quantization error.

[0386] [Addition unit]

[0387] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 to the prediction image input from the prediction control unit 128. As a result, a reconstructed image is generated. Further, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.

[0388] [Block memory]

[0389] The block memory 118 is, for example, a storage unit for storing a block referred to in intra prediction and a block within the current picture. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.

[0390] [Frame memory]

[0391] The frame memory 122 is, for example, a storage unit for storing a reference picture used in inter prediction, and is also referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.

[0392] [Loop filter unit]

[0393] The loop filter unit 120 performs loop filtering on the reconstructed image output from the addition unit 116, and outputs the reconstructed image subjected to the filtering to the frame memory 122. Loop filtering refers to filtering used within an encoding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).

[0394] Figure 22 is a block diagram showing an example of the functional structure of the loop filter unit 120.

[0395] For example Figure 22As shown, the loop filter 120 includes a deblocking filter processing section 120a, an SAO processing section 120b, and an ALF processing section 120c. The deblocking filter processing section 120a applies the deblocking filter processing described above to the reconstructed image. The SAO processing section 120b applies the SAO processing described above to the reconstructed image after the deblocking filter processing. In addition, the ALF processing section 120c applies the ALF processing described above to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is processing for improving the quality of the image by reducing ringing (a phenomenon in which the pixel values around the edge are deformed in a wavy manner) and correcting the deviation of the pixel values. In this SAO processing, for example, there are an edge offset processing and a band offset processing. In addition, the loop filter 120 can not include Figure 22 All of the processing sections disclosed can include only a part of the processing sections. In addition, the loop filter 120 can be configured to perform the above-described processing in a different order from the processing order disclosed in the first embodiment. Figure 22

[0396] [Loop filter > Adaptive loop filter]

[0397] In the ALF, a least square error filter for removing coding distortion is used, and for example, one filter selected from a plurality of filters based on the direction and activity of the gradient in each 2x2 pixel sub-block in the current block is used.

[0398] Specifically, first, the sub-blocks (for example, 2x2 pixel sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed, for example, based on the direction and activity of the gradient. In a specific example, using the direction value D (for example, 0-2 or 0-4) of the gradient and the activity value A (for example, 0-4) of the gradient, a classification value C (for example, C=5D+A) is calculated. And, based on the classification value C, the sub-blocks are classified into a plurality of classes.

[0399] The direction value D of the gradient is derived, for example, by comparing the gradients of a plurality of directions (for example, horizontal, vertical, and two diagonal directions). In addition, the activity value A of the gradient is derived, for example, by adding the gradients of a plurality of directions and quantizing the addition result.

[0400] Based on the result of such classification, a filter for the sub-block is determined from among a plurality of filters.

[0401] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figure 23A-23C is a diagram showing a plurality of examples of the shape of the filter used in the ALF. Figure 23A shows a 5x5 diamond shape filter,​Figure 23B represents a 7x7 diamond-shaped filter, Figure 23C represents a 9x9 diamond-shaped filter. Information representing the shape of a filter is usually signaled at a picture level. In addition, the signaling of information representing the shape of a filter need not be limited to the picture level, but can be another level (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0402] The turning on / off of ALF can also be determined, for example, at a picture level or a CU level. For example, as for luma, whether to employ ALF can be determined at a CU level, and as for chroma, whether to employ ALF can be determined at a picture level. Information representing the turning on / off of ALF is usually signaled at a picture level or a CU level. In addition, the signaling of information representing the turning on / off of ALF need not be limited to the picture level or the CU level, but can be another level (e.g., sequence level, slice level, tile level, or CTU level).

[0403] In addition, as described above, one filter is selected from a plurality of filters to perform ALF processing on a sub-block. For each of the plurality of filters (e.g., up to 15 or 25 filters), a coefficient set composed of a plurality of coefficients used by the filter is usually signaled at a picture level. In addition, the signaling of the coefficient set need not be limited to the picture level, but can be another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0404] [Loop filter] Cross Component Adaptive Loop Filter (or CC-ALF)

[0405] Figure 23D is a diagram representing an example in which a Y sample (1st component) is used for CCALF of Cb and CCALF of Cr (a plurality of components different from the 1st component). Figure 23E is a diagram representing a diamond-shaped filter.

[0406] One example of CC-ALF is a linear diamond-shaped filter (7x7) applied to a sub-block. Figure 23D , Figure 23E) applied to the luma channel of each chroma component. For example, the filter coefficients are sent in APS, scaled by a factor of 2Λ10, and rounded for fixed-point representation. The application of the filter is controlled to be variable block size and signaled by a context-coded flag received per block of samples. The block size and CC-ALF enable flag are received at slice level per chroma component. The syntax and semantics of CC-ALF are provided in Appendix. In this document, block sizes of 16x16, 32x32, 64x64, 128x128 are supported (in chroma samples).

[0407] [Loop filter > Joint Chroma Cross Component Adaptive Loop Filter]

[0408] Figure 23F is an example of a picture representing JC-CCALF. Figure 23G is an example of a picture representing JC-CCALF weight_index candidates.

[0409] One example of JC-CCALF uses only one CCALF filter, generates one CCALF filter output as the chroma adjustment signal for only one color component, and applies appropriately weighted versions of the same chroma adjustment signal to the other color components. In this way, the complexity of the existing CCALF is roughly halved.

[0410] The weight value is coded as a sign flag and a weight index. The weight index, denoted as weight_index, is coded as 3 bits, specifying the size of the JC-CCALF weight JcCcWeight. It cannot be equal to 0. The size of JcCcWeight is determined as follows.

[0411] • For weight_index less than or equal to 4, JcCcWeight is equal to weight_index » 2.

[0412] • For other cases, JcCcWeight is equal to 4 / (weight_index - 4).

[0413] The block-level on / off control of the ALF filtering for Cb and Cr is separate. This is the same as for CCALF, and two separate sets of block-level on / off control flags are coded. Here, unlike for CCALF, the on / off control block sizes for Cb and Cr are the same, so only one block size variable is coded.

[0414] [Loop filter>Deblocking filter]

[0415] In the deblocking filter processing, the loop filter 120 reduces the distortion generated at the block boundary by performing the filter processing on the block boundary of the reconstructed image.

[0416] Figure 24 is a block diagram showing an example of the detailed structure of the deblocking filter processing section 120a.

[0417] The deblocking filter processing section 120a has, for example, a boundary determination section 1201, a filter determination section 1203, a filter processing section 1205, a processing determination section 1208, a filter characteristic decision section 1207, and switches 1202, 1204, and 1206.

[0418] The boundary determination section 1201 determines whether or not there is a pixel (i.e., an object pixel) on which the deblocking filter processing is performed in the vicinity of the block boundary. Then, the boundary determination section 1201 outputs the result of the determination to the switch 1202 and the processing determination section 1208.

[0419] In a case where it is determined by the boundary determination section 1201 that the object pixel exists in the vicinity of the block boundary, the switch 1202 outputs the image before the filter processing to the switch 1204. On the contrary, in a case where it is determined by the boundary determination section 1201 that the object pixel does not exist in the vicinity of the block boundary, the switch 1202 outputs the image before the filter processing to the switch 1206. In addition, the image before the filter processing is an image constituted of the object pixel and at least one peripheral pixel located in the periphery of the object pixel.

[0420] The filter determination section 1203 determines whether or not to perform the deblocking filter processing on the object pixel on the basis of the pixel value of at least one peripheral pixel located in the periphery of the object pixel. Then, the filter determination section 1203 outputs the result of the determination to the switch 1204 and the processing determination section 1208.

[0421] In a case where it is determined by the filter determination section 1203 that the deblocking filter processing is performed on the object pixel, the switch 1204 outputs the image before the filter processing acquired via the switch 1202 to the filter processing section 1205. On the contrary, in a case where it is determined by the filter determination section 1203 that the deblocking filter processing is not performed on the object pixel, the switch 1204 outputs the image before the filter processing acquired via the switch 1202 to the switch 1206.

[0422] In a case where the image before the filter processing is acquired via the switches 1202 and 1204, the filter processing section 1205 performs the deblocking filter processing on the object pixel with the filter characteristic decided by the filter characteristic decision section 1207. Then, the filter processing section 1205 outputs the pixel after the filter processing to the switch 1206.

[0423] According to control of the processing determination section 1208, the switch 1206 selectively outputs the pixel which is not deblocking filter-processed and the pixel which is deblocking filter-processed by the filter processing section 1205.

[0424] The processing determination section 1208 controls the switch 1206 based on respective determination results of the boundary determination section 1201 and the filter determination section 1203. That is, the processing determination section 1208 outputs the pixel which is deblocking filter-processed from the switch 1206 in a case where it is determined by the boundary determination section 1201 that the object pixel exists in the vicinity of the block boundary and it is determined by the filter determination section 1203 that the deblocking filter processing is performed on the object pixel. In addition, the processing determination section 1208 outputs the pixel which is not deblocking / filter-processed from the switch 1206 in a case other than the above-described case. By repeating the output of such a pixel, the image which is filter-processed is output from the switch 1206. In addition, Figure 24 The structure illustrated is an example of the structure in the deblocking filter processing section 120a, and the deblocking filter processing section 120a can have other structures.

[0425] Figure 25 is a diagram which shows an example of deblocking filter having a filter characteristic which is symmetrical with respect to the block boundary.

[0426] In the deblocking filter processing, for example, using the pixel value and the quantization parameter, either one of two deblocking filters, that is, a strong filter and a weak filter, which have different characteristics is selected. In the strong filter, as shown in Figure 25 In a case where the pixels p0 to p2 and the pixels q0 to q2 exist across the block boundary, as shown, the pixel value of each of the pixels q0 to q2 is changed to the pixel value q'0 to q'2 by performing the operation shown in the following expression.

[0427] q'0 = (p1 + 2 x p0 + 2 x q0 + 2 x q1 + q2 + 4) / 8

[0428] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0429] q'2 = (p0 + q0 + q1 + 3 x q2 + 2 x q3 + 4) / 8

[0430] Further, in the above-described expressions, p0 to p2 and q0 to q2 are the pixel value of each of the pixels p0 to p2 and the pixels q0 to q2. In addition, q3 is the pixel value of the pixel q3 which is adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above-described expressions, the coefficient which is multiplied by the pixel value of each pixel used in the deblocking filter processing is a filter coefficient.

[0431] Furthermore, in the deblocking filtering process, limiting can also be performed in a way that prevents the calculated pixel value from changing even if it exceeds a threshold. In this limiting process, a threshold determined by the quantization parameters is used to limit the calculated pixel value based on the above formula to "the pixel value before calculation ± 2 × the threshold". This prevents excessive smoothing.

[0432] Figure 26 This is a diagram illustrating an example of block boundaries used in deblocking filtering. Figure 27 This is a graph representing an example of BS value.

[0433] The block boundaries for deblocking filtering are, for example, Figure 26 The boundaries of the CU, PU, ​​or TU of the 8×8 pixel block are shown. Deblocking filtering is performed, for example, in units of 4 rows or 4 columns. First, for Figure 26 Blocks P and Q are shown, as follows Figure 27 That determines the Bs (Boundary Strength) value.

[0434] according to Figure 27 The Bs value determines whether deblocking filtering of different intensities is applied even to block boundaries belonging to the same image. A Bs value of 2 performs deblocking filtering on the chrominance signal. A Bs value of 1 or higher, meeting specified conditions, performs deblocking filtering on the luminance signal. Furthermore, the criteria for determining the Bs value are not limited to... Figure 27 The conditions shown can also be determined based on other parameters.

[0435] [Prediction Unit (Intra-frame Prediction Unit / Inter-frame Prediction Unit / Prediction Control Unit)]

[0436] Figure 28 This is a flowchart illustrating an example of the processing performed by the prediction unit of the coding apparatus 100. Furthermore, as an example, the prediction unit is composed of all or part of the components of the intra-frame prediction unit 124, the inter-frame prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra-frame prediction unit 124 and the inter-frame prediction unit 126.

[0437] The prediction unit generates a prediction image for the current block (step Sb_1). Furthermore, the prediction image may include, for example, an intra-frame prediction image (intra-frame prediction signal) or an inter-frame prediction image (inter-frame prediction signal). Specifically, the prediction unit generates a prediction image for the current block using a reconstructed image obtained by generating prediction images for other blocks, generating prediction residuals, generating quantization coefficients, restoring prediction residuals, and adding prediction images.

[0438] The reconstructed image can be, for example, an image of a reference picture or an image of a picture that contains the coded block (i.e., the other block described above) within a current picture, i.e., the picture of the current block.

[0439] Figure 29 is a flowchart of another example of the processing performed by the prediction section of the encoding apparatus 100.

[0440] The prediction section generates a prediction image by the first method (step Sc la), generates a prediction image by the second method (step Sc lb), and generates a prediction image by the third method (step Sc lc). The first method, the second method, and the third method are mutually different methods for generating a prediction image, and can be, for example, an inter prediction method, an intra prediction method, and another prediction method, respectively. In such a prediction method, the reconstructed image described above can also be used.

[0441] Next, the prediction section evaluates the prediction images generated in steps Sc la, Sc lb, and Sc lc, respectively (step Sc 2). For example, the prediction section evaluates the prediction images by calculating the cost C for the prediction images generated in each of steps Sc la, Sc lb, and Sc lc, and comparing the costs C of the prediction images. In addition, the cost C is calculated by the equation of the R-D optimization model, for example, C = D + λ x R. In this equation, D is the coding distortion of the prediction image, and is represented by, for example, the sum of the absolute values of the differences between the pixel values of the current block and the pixel values of the prediction image. Furthermore, R is the bit rate of the stream. In addition, λ is an undetermined multiplier of, for example, Lagrange.

[0442] Next, the prediction section selects one of the prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 3). That is, the prediction section selects the method or mode for obtaining the final prediction image. For example, the prediction section selects the prediction image of the smallest cost C on the basis of the costs C calculated for the prediction images. Alternatively, the evaluation of step Sc 2 and the selection of the prediction image in step Sc 3 can also be performed on the basis of the parameters used in the encoding processing. The encoding apparatus 100 can signal information for determining the selected prediction image, method, or mode to the stream. The information can be, for example, a flag or the like. As a result, the decoding apparatus 200 can generate a prediction image in the method or mode selected in the encoding apparatus 100 on the basis of the information. Furthermore, in the example shown in FIG. 8, the prediction section selects one of the prediction images after generating the prediction images by the respective methods. However, the prediction section can select the method or mode on the basis of the parameters for the encoding processing described above before generating the prediction images, and can generate the prediction images in accordance with the method or mode. Figure 29 In the example shown in FIG. 8, the prediction section selects one of the prediction images after generating the prediction images by the respective methods. However, the prediction section can select the method or mode on the basis of the parameters for the encoding processing described above before generating the prediction images, and can generate the prediction images in accordance with the method or mode.

[0443] For example, the first and second modes are intra prediction and inter prediction, respectively, and the prediction section can select a final prediction image for the current block from among prediction images generated in accordance with these prediction modes.

[0444] Figure 30 is a flowchart of another example of processing performed by the prediction section of the encoding apparatus 100.

[0445] First, the prediction section generates a prediction image by intra prediction (step Sd la), and generates a prediction image by inter prediction (step Sd lb). Further, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0446] Next, the prediction section evaluates each of the intra prediction image and the inter prediction image (step Sd 2). The above-described cost C can also be used in this evaluation. Then, the prediction section can select, as the final prediction image of the current block, the prediction image for which the minimum cost C is calculated from among the intra prediction image and the inter prediction image (step Sd 3). That is, the prediction mode or the mode used to generate the prediction image of the current block is selected.

[0447] [Intra Prediction Section]

[0448] The intra prediction section 124 performs intra prediction (also referred to as in-picture prediction) of the current block with reference to the blocks within the current picture saved in the block memory 118, thereby generating a prediction image (i.e., an intra prediction image) of the current block. Specifically, the intra prediction section 124 generates an intra prediction image by performing intra prediction with reference to the pixel values (e.g., luminance values, color difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control section 128.

[0449] For example, the intra prediction section 124 performs intra prediction using one of a plurality of intra prediction modes specified in advance. The plurality of intra prediction modes generally includes one or more non-directional prediction modes and a plurality of directional prediction modes.

[0450] The one or more non-directional prediction modes include, for example, a Planar prediction mode and a DC prediction mode specified in the H.265 / HEVC specification.

[0451] The plurality of directional prediction modes include, for example, 33 directional prediction modes specified in the H.265 / HEVC specification. Alternatively, the plurality of directional prediction modes can include 32 directional prediction modes in addition to the 33 directional prediction modes (a total of 65 directional prediction modes). Figure 31is a diagram representing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows represent 33 directions specified by the H.265 / HEVC specification, and dotted arrows represent 32 additional directions (2 non-directional prediction modes are not shown in Figure 31

[0452] In various installation examples, in intra prediction of a color difference block, a luminance block can also be referred to. That is, a color difference component of a current block can also be predicted based on a luminance component of the current block. Such intra prediction is a case where it is called CCLM (cross-component linear model) prediction. Such an intra prediction mode of a color difference block (for example, called a CCLM mode) that refers to a luminance block can also be added as one of the intra prediction modes of the color difference block.

[0453] The intra prediction section 124 can also correct a pixel value after intra prediction based on a gradient of a reference pixel in a horizontal / vertical direction. Intra prediction with such correction is a case where it is called PDPC (position dependent intra prediction combination). Information indicating whether or not PDPC is adopted (for example, called a PDPC flag) is generally signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and can be another level (for example, a sequence level, a picture level, a slice level, a tile level, or a CTU level).

[0454] Figure 32 is a flowchart representing an example of the processing performed by the intra prediction section 124.

[0455] The intra prediction section 124 selects one intra prediction mode from among a plurality of intra prediction modes (step Sw_1). Then, the intra prediction section 124 generates a prediction image in accordance with the selected intra prediction mode (step Sw_2). Next, the intra prediction section 124 determines MPMs (Most Probable Modes) (step Sw_3). The MPMs are, for example, composed of 6 intra prediction modes. Two of the 6 intra prediction modes can be Planar and DC prediction modes, and the remaining 4 modes can be directional prediction modes. Then, the intra prediction section 124 determines whether or not the intra prediction mode selected in step Sw_1 is included in the MPMs (step Sw_4).

[0456] ​When it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction section 124 sets the MPM flag to 1 (step Sw_5), generates information indicating the selected intra prediction mode in the MPM (step Sw_6). The MPM flag set to 1 and the information indicating the intra prediction mode are encoded as prediction parameters by the entropy encoding section 110, respectively.

[0457] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction section 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction section 124 does not set the MPM flag. Then, the intra prediction section 124 generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). The MPM flag set to 0 and the information indicating the intra prediction mode are encoded as prediction parameters by the entropy encoding section 110, respectively. The information indicating the intra prediction mode indicates, for example, any one of 0 to 60.

[0458] [Inter prediction section]

[0459] The inter prediction section 126 performs inter prediction (also referred to as inter picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a predicted image (inter prediction image). Inter prediction is performed in units of the current block or a current sub-block within the current block. A sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, can be 8x8 pixels, or can be a size other than these. The size of the sub-block can also be switched in units of a slice, a tile, or a picture.

[0460] For example, the inter prediction section 126 performs motion search within the reference picture for the current block or the current sub-block, and searches for a reference block or sub-block most consistent with the current block or the current sub-block. Then, the inter prediction section 126 acquires motion information (e.g., a motion vector) that compensates for motion or change from the reference block or sub-block to the current block or sub-block. The inter prediction section 126 performs motion compensation (or motion prediction) based on the motion information, thereby generating an inter prediction image of the current block or sub-block. Then, the inter prediction section 126 outputs the generated inter prediction image to the prediction control section 128.

[0461] The motion information used in the motion compensation is signaled as the inter prediction image in various forms. For example, a motion vector can be signaled. As another example, a difference between a motion vector and a prediction motion vector (motion vector predictor) can be signaled.

[0462] [Reference picture list]

[0463] Figure 33 is a diagram showing an example of each reference picture, Figure 34 is a conceptual diagram showing an example of a reference picture list. The reference picture list is a list showing one or more reference pictures stored in the frame memory 122. In addition, in Figure 33 , the rectangles represent pictures, the arrows represent the reference relationship of the pictures, the horizontal axis represents time, I, P, and B in the rectangles respectively represent an intra prediction picture, a single prediction picture, and a double prediction picture, and the numbers in the rectangles represent the decoding order. As shown in Figure 33 , the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As shown in Figure 34 , the reference picture list is a list showing candidates of reference pictures, and for example, one picture (or slice) can have one or more reference picture lists. For example, if the current picture is a single prediction picture, one reference picture list is used, and if the current picture is a double prediction picture, two reference picture lists are used. In the example of Figure 33 and Figure 34 , the picture B3 as the current picture currPic has two reference picture lists, L0 list and L1 list. In the case where the current picture currPic is the picture B3, the candidates of the reference pictures of the current picture currPic are I0, P1, and B2, and each reference picture list (i.e., L0 list and L1 list) shows these pictures. The inter prediction section 126 or the prediction control section 128 specifies which picture in each reference picture list is actually referred to by the reference picture index refidxLx. In Figure 34 , the reference pictures P1 and B2 are specified by the reference picture indexes refIdxL0 and refIdxL1.

[0464] Such a reference picture list can be generated in a sequence unit, a picture unit, a slice unit, a tile unit, a CTU unit, or a CU unit. In addition, the reference picture index of the reference picture referred to in inter prediction among the reference pictures shown in the reference picture list can be encoded in a sequence level, a picture level, a slice level, a tile level, a CTU level, or a CU level. In addition, a common reference picture list can be used in a plurality of inter prediction modes.

[0465] [Basic flow of inter prediction]

[0466] Figure 35 is a flowchart showing the basic flow of inter prediction.

[0467] The inter prediction section 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction section 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0468] Here, in the generation of the prediction image, the inter prediction section 126 generates the prediction image, for example, by performing determination of a motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Further, in the determination of the MV, the inter prediction section 126 determines the MV, for example, by performing selection of a candidate motion vector (candidate MV) (step Se_1) and derivation of the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter prediction section 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. In addition, in the candidate MV list, a past-derived MV can be added as a candidate MV. Further, in the derivation of the MV, the inter prediction section 126 can determine the MV of the current block by further selecting at least one candidate MV from the at least one candidate MV. Alternatively, the inter prediction section 126 can determine the MV of the current block by searching a region of a reference picture indicated by each of the selected at least one candidate MV. In addition, the action of searching the region of the reference picture can be referred to as motion estimation.

[0469] Further, in the above example, steps Se_1 to Se_3 are performed by the inter prediction section 126, but the processing of step Se_1 or step Se_2, or the like can be performed by another constituent element included in the encoding apparatus 100.

[0470] In addition, the candidate MV list can be made for each of the respective inter prediction modes, or a common candidate MV list can be used in a plurality of inter prediction modes. Further, the processing of step Se_3 and step Se_4 corresponds to the processing of steps Sa_3 and Sa_4, respectively, illustrated in FIG. 8. Figure 9 Further, the processing of step Se_3 corresponds to the processing of step Sd_1b of FIG. 9. Figure 30

[0471] [Flow of MV derivation]

[0472] Figure 36 is a flowchart illustrating an example of MV derivation.

[0473] The inter prediction section 126 can derive the MV of the current block in a mode in which motion information (e.g., MV) is encoded. In this case, for example, the motion information can be encoded as a prediction parameter, and be signaled. That is, the encoded motion information is included in a stream. ​

[0474] Alternatively, the inter prediction section 126 can derive the MV in a mode in which the motion information is not coded. In this case, the motion information is not included in the stream.

[0475] Here, the mode of the MV derivation has a normal inter mode, a normal merge mode, an FRUC mode, and an affine mode, and the like, which will be described later. Among these modes, the mode in which the motion information is coded has the normal inter mode, the normal merge mode, and the affine mode (specifically, an affine inter mode and an affine merge mode), and the like. Further, the motion information can include not only the MV but also prediction MV selection information, which will be described later. Further, the mode in which the motion information is not coded has the FRUC mode, and the like. The inter prediction section 126 selects the mode for deriving the MV of the current block from among these multiple modes, and derives the MV of the current block using the selected mode.

[0476] Figure 37 is a flowchart showing another example of the MV derivation.

[0477] The inter prediction section 126 can derive the MV of the current block in a mode in which the differential MV is coded. In this case, for example, the differential MV is coded as a prediction parameter, and is signaled. That is, the coded differential MV is included in the stream. The differential MV is the difference between the MV of the current block and a prediction MV thereof. Further, the prediction MV is a predicted motion vector.

[0478] Alternatively, the inter prediction section 126 can derive the MV in a mode in which the differential MV is not coded. In this case, the coded differential MV is not included in the stream.

[0479] Here, as described above, the mode of the MV derivation has the normal inter mode, the normal merge mode, the FRUC mode, and the affine mode, and the like, which will be described later. Among these modes, the mode in which the differential MV is coded has the normal inter mode, and the affine mode (specifically, the affine inter mode), and the like. Further, the mode in which the differential MV is not coded has the FRUC mode, the normal merge mode, and the affine mode (specifically, the affine merge mode), and the like. The inter prediction section 126 selects the mode for deriving the MV of the current block from among these multiple modes, and derives the MV of the current block using the selected mode.

[0480] [MV derivation mode]

[0481] Figure 38A and Figure 38B is a diagram showing an example of the classification of each mode of the MV derivation. For example, as Figure 38AAs shown, the MV derivation modes are classified into three modes according to whether or not the motion information is coded and whether or not the differential MV is coded. The three modes are an inter mode, a merge mode, and an FRUC (frame rate up-conversion) mode. The inter mode is a mode in which the motion search is performed and in which the motion information and the differential MV are coded. For example, as shown in Figure 38B As shown, the inter mode includes an affine inter mode and a normal inter mode. The merge mode is a mode in which the motion search is not performed and in which the MV is selected from the peripheral coded blocks and the MV of the current block is derived using the MV. The merge mode is basically a mode in which the motion information is coded and the differential MV is not coded. For example, as shown in Figure 38B As shown, the merge mode includes a normal merge mode (sometimes also referred to as a usual merge mode or a regular merge mode), an MMVD (Merge with Motion Vector Difference) mode, a CIIP (Combined inter merge / intra prediction) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, in the MMVD mode among the modes included in the merge mode, the differential MV is exceptionally coded. Further, the affine merge mode and the affine inter mode described above are modes included in the affine mode. The affine mode is a mode in which the MVs of a plurality of sub-blocks constituting the current block are assumed to be affine-transformed and the MVs of the sub-blocks are derived as the MV of the current block. The FRUC mode is a mode in which the MV of the current block is derived by searching between the coded regions and is a mode in which neither the motion information nor the differential MV is coded. Details of these modes will be described later.

[0482] Further, Figure 38A and Figure 38B The classification of the modes shown is an example and is not limited thereto. For example, in a case where the differential MV is coded in the CIIP mode, the CIIP mode is classified as the inter mode.

[0483] [MV derivation > normal inter mode]

[0484] The normal inter mode is an inter prediction mode in which the MV of the current block is derived based on a region of a reference picture represented by a candidate MV by finding a block similar to the image of the current block in the region. Further, in the normal inter mode, the differential MV is coded.

[0485] Figure 39 is a flowchart showing an example of the inter prediction based on the normal inter mode.

[0486] First, the inter prediction section 126 acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of coded blocks located around the current block in time or space (step Sg_1). That is, the inter prediction section 126 creates a list of candidate MVs.

[0487] Next, the inter prediction section 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs acquired in step Sg_1 as prediction MV candidates in a predetermined priority order (step Sg_2). Note that the priority order is predetermined for each of the N candidate MVs.

[0488] Next, the inter prediction section 126 selects one of the N prediction MV candidates as a prediction MV for the current block (step Sg_3). At this time, the inter prediction section 126 encodes prediction MV selection information for identifying the selected prediction MV into a stream. That is, the inter prediction section 126 outputs the prediction MV selection information as a prediction parameter to the entropy coding section 110 via the prediction parameter generation section 130.

[0489] Next, the inter prediction section 126 refers to a coded reference picture to derive the MV of the current block (step Sg_4). At this time, the inter prediction section 126 also encodes a difference value between the derived MV and the prediction MV as a difference MV into a stream. That is, the inter prediction section 126 outputs the difference MV as a prediction parameter to the entropy coding section 110 via the prediction parameter generation section 130. Note that the coded reference picture is a picture composed of a plurality of blocks reconstructed after coding.

[0490] Finally, the inter prediction section 126 generates a prediction image of the current block by performing motion compensation on the current block using the derived MV and the coded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are performed for each block. For example, when the processes of steps Sg_1 to Sg_5 are performed for all blocks included in a slice, respectively, the inter prediction using the normal inter mode ends for the slice. Also, when the processes of steps Sg_1 to Sg_5 are performed for all blocks included in a picture, respectively, the inter prediction using the normal inter mode ends for the picture. Further, it can also be that the processes of steps Sg_1 to Sg_5, when performed for a part of the blocks included in a slice, end the inter prediction using the normal inter mode for the slice. Similarly, it can also be that the processes of steps Sg_1 to Sg_5, when performed for a part of the blocks included in a picture, end the inter prediction using the normal inter mode for the picture.

[0491] In addition, the prediction picture is the above-described inter prediction signal. Furthermore, information indicating the inter prediction mode (in the above-described example, the normal inter mode) used in the generation of the prediction picture, which is included in the coded signal, is coded as, for example, the prediction parameter.

[0492] In addition, the candidate MV list can also be commonly used with the list used in other modes. Furthermore, the processing related to the candidate MV list can be applied to the processing related to the list used in other modes. The processing related to the candidate MV list is, for example, extraction or selection of the candidate MV from the candidate MV list, rearrangement of the candidate MV, or deletion of the candidate MV, and the like.

[0493] [MV derivation> normal merge mode]

[0494] The normal merge mode is an inter prediction mode in which the MV of the current block is derived by selecting a candidate MV from the candidate MV list. In addition, the normal merge mode is a narrow sense of the merge mode, and is sometimes simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.

[0495] Figure 40 is a flowchart indicating an example of inter prediction based on the normal merge mode.

[0496] First, the inter prediction section 126 acquires a plurality of candidate MVs for the current block based on information of a plurality of coded block MVs and the like located around the current block in time or space (step Sh_1). That is, the inter prediction section 126 creates a candidate MV list.

[0497] Next, the inter prediction section 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs acquired in step Sh_1 (step Sh_2). At this time, the inter prediction section 126 codes MV selection information for identifying the selected candidate MV to the stream. That is, the inter prediction section 126 outputs the MV selection information as the prediction parameter to the entropy coding section 110 via the prediction parameter generation section 130.

[0498] Finally, the inter prediction section 126 generates a prediction image of the current block by performing motion compensation on the current block using the derived MV and the coded reference picture (step Sh_3). The processes of steps Sh_l to Sh_3 are performed, for example, for each block. For example, when the processes of steps Sh_l to Sh_3 are performed for all the blocks included in a slice, respectively, the inter prediction using the normal merge mode for the slice ends. Also, when the processes of steps Sh_l to Sh_3 are performed for all the blocks included in a picture, respectively, the inter prediction using the normal merge mode for the picture ends. Further, the processes of steps Sh_l to Sh_3 can also be such that, when performed for a part of the blocks included in a slice, the inter prediction using the normal merge mode for the slice ends. Likewise, the processes of steps Sh_l to Sh_3 can also be such that, when performed for a part of the blocks included in a picture, the inter prediction using the normal merge mode for the picture ends.

[0499] Further, information indicating the inter prediction mode (in the above example, the normal merge mode) used in the generation of the prediction image is coded as, for example, a prediction parameter, in the stream.

[0500] Figure 41 is a diagram for explaining an example of the MV derivation process for the current picture based on the normal merge mode.

[0501] First, the inter prediction section 126 generates a candidate MV list in which candidate MVs are registered. As the candidate MVs, there are: a spatial neighboring candidate MV, which is an MV possessed by a plurality of coded blocks located in the vicinity of the current block in space; a temporal neighboring candidate MV, which is an MV possessed by a block in the vicinity of the position of the current block in the coded reference picture; a combined candidate MV, which is an MV generated by combining the MV values of the spatial neighboring candidate MV and the temporal neighboring candidate MV; and a zero candidate MV, which is an MV having a value of zero, and the like.

[0502] Next, the inter prediction section 126 determines, as the MV of the current block, one candidate MV selected from among the plurality of candidate MVs registered in the candidate MV list.

[0503] Further, in the entropy coding section 110, a signal indicating which candidate MV is selected, i.e., merge_idx, is coded by being described in the stream.

[0504] Further, the candidate MVs registered in the candidate MV list explained in Figure 41 may be a number different from that in the drawing, or a structure not including a part of the kinds of the candidate MVs in the drawing, or a structure in which a candidate MV other than the kinds of the candidate MVs in the drawing is added.

[0505] The MV of the current block derived by the normal merge mode can also be used to determine the final MV by performing the DMVR (dynamic motion vector refreshing) described later. In addition, in the normal merge mode, the difference MV is not encoded, but in the MMVD mode, the difference MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list as in the normal merge mode, but encodes the difference MV. As shown in Figure 38B FIG. 16, such MMVD can also be classified as the merge mode together with the normal merge mode. In addition, the difference MV in the MMVD mode can also be different from the difference MV used in the inter mode, for example, the derivation of the difference MV in the MMVD mode can also be a process with less processing amount than the derivation of the difference MV in the inter mode.

[0506] In addition, the prediction image generated in the inter prediction can also be made to coincide with the prediction image generated in the intra prediction, and the CIIP (Combined inter merge / intra prediction) mode of generating the prediction image of the current block can be performed.

[0507] In addition, the candidate MV list can also be referred to as the candidate list. In addition, merge_idx is the MV selection information.

[0508] [MV derivation > HMVP mode]

[0509] Figure 42 FIG. 16 is a diagram for explaining an example of the MV derivation process of the current picture based on the HMVP mode.

[0510] In the normal merge mode, one candidate MV is selected from the candidate MV list generated by referring to the already encoded block (for example, CU), and thereby the MV of the current block such as the CU is determined. In this case, other candidate MVs can also be registered in the candidate MV list. The mode of registering such other candidate MVs is referred to as the HMVP mode.

[0511] In the HMVP mode, the candidate MVs for the HMVP are managed using a FIFO (First-In First-Out) buffer separately from the candidate MV list of the normal merge mode.

[0512] In the FIFO buffer, the motion information such as the MV of the block processed in the past is sequentially saved from the new FIFO buffer. In the management of the FIFO buffer, every time the processing of one block is performed, the MV of the latest block (that is, the CU processed immediately before) is saved in the FIFO buffer, and in place of this, the MV of the earliest CU (that is, the CU processed first) in the FIFO buffer is deleted from the FIFO buffer. In this way, the FIFO buffer is managed so that the MV of the CU processed immediately before is always saved in the FIFO buffer. Figure 42In the example shown, HMVP1 is the MV of the most recent block, and HMVP5 is the MV of the earliest block.

[0513] Then, for example, the inter prediction section 126 checks, for each MV managed in the FIFO buffer, in order from HMVP1, whether the MV is different from all the candidate MVs already registered in the candidate MV list for the normal merge mode. Also, the inter prediction section 126 can add the MV managed in the FIFO buffer to the candidate MV list for the normal merge mode as a candidate MV in the case where it is determined that the MV is different from all the candidate MVs. At this time, the candidate MV registered from the FIFO buffer can be one or a plurality of candidates.

[0514] In this way, by using the HMVP mode, it is possible to add not only the MVs of the spatially or temporally adjacent blocks of the current block to the candidates, but also the MVs of the blocks processed in the past. As a result, by expanding the changes in the candidate MVs for the normal merge mode, it is possible to increase the likelihood of improving the coding efficiency.

[0515] In addition, the above MV can also be motion information. That is, the information saved in the candidate MV list and the FIFO buffer can include not only the value of the MV, but also information indicating the reference picture, the direction and the number of references, and the like. In addition, the above block is, for example, a CU.

[0516] In addition, Figure 42 The candidate MV list and the FIFO buffer described above are an example, and the candidate MV list and the FIFO buffer can be a list or a buffer of a different size from Figure 42 , or a structure in which the candidate MVs are registered in a different order from Figure 42 . In addition, the processing described here is common in the encoding device 100 and in the decoding device 200.

[0517] Furthermore, the HMVP mode can also be applied to modes other than the normal merge mode. For example, it is also possible to sequentially save, from a new FIFO buffer, motion information such as the MVs of the blocks processed in the past in the affine mode, as candidate MVs. The mode in which the HMVP mode is applied in the affine mode can be referred to as a history affine mode.

[0518] [MV derivation > FRUC mode]

[0519] The motion information can also not be signaled from the encoding apparatus 100 side, but derived at the decoding apparatus 200 side. For example, the motion information can also be derived by performing a motion search at the decoding apparatus 200 side. In such a case, the motion search is performed at the decoding apparatus 200 side without using the pixel values of the current block. Such a mode of performing a motion search at the decoding apparatus 200 side is an FRUC (frame rate up-conversion) mode or a PMMVD (pattern matched motion vector derivation) mode, or the like.

[0520] Figure 43 An example of FRUC processing is shown in FIG. 12. First, referring to the MVs of each of the coded blocks that are spatially or temporally adjacent to the current block, a list is generated that represents these MVs as candidate MVs (i.e., it is a candidate MV list, and can also be common to the candidate MV list of the normal merge mode) (step Si_1). Next, a best candidate MV is selected from among the plurality of candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each of the candidate MVs included in the candidate MV list is calculated, and based on this evaluation value, one candidate is selected as the best candidate MV. Also, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is derived as it is as the MV for the current block. Further, for example, the MV for the current block can also be derived by performing pattern matching in the surrounding area of the position within the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching in the reference picture and the evaluation value can be performed on the surrounding area of the best candidate MV, and in the case where there is an MV with a better value of the evaluation value, the best candidate MV is updated to this MV, and this is taken as the final MV of the current block. The update to the MV with a better evaluation value can not be performed.

[0521] Finally, the inter prediction section 126 generates a prediction image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_l to Si_5 are performed, for example, for each block. For example, when the processes of steps Si_l to Si_5 are performed for all blocks included in a slice respectively, the inter prediction using the FRUC mode for the slice ends. Also, when the processes of steps Si_l to Si_5 are performed for all blocks included in a picture respectively, the inter prediction using the FRUC mode for the picture ends. Also, the processes of steps Si_l to Si_5 can be such that, when the processes are performed for a part of the blocks included in a slice, the inter prediction using the FRUC mode for the slice ends. Likewise, the processes of steps Si_l to Si_5 can be such that, when the processes are performed for a part of the blocks included in a picture, the inter prediction using the FRUC mode for the picture ends.

[0522] The same processes as those described above can also be performed in the case where the processes are performed in a sub-block unit.

[0523] The evaluation value can also be calculated by various methods. For example, a reconstructed image of a region within a reference picture corresponding to an MV is compared with a reconstructed image of a prescribed region (for example, the region can be a region of another reference picture or a region of a neighboring block of the current picture). Then, a difference in pixel values of the two reconstructed images can be calculated as the evaluation value for the MV. Alternatively, other information can be used in addition to the difference value to calculate the evaluation value.

[0524] Next, pattern matching is described in detail. First, one of the candidate MVs included in the candidate MV list (also referred to as a merge list) is selected as a starting point of search based on pattern matching. As the pattern matching, first pattern matching or second pattern matching can be used. The first pattern matching and the second pattern matching are respectively referred to as bilateral matching and template matching.

[0525] [MV derivation > FRUC > bilateral matching]

[0526] In the first pattern matching, pattern matching is performed between two blocks within different two reference pictures along the motion trajectory of the current block. Thus, in the first pattern matching, as the prescribed region for the calculation of the evaluation value for the candidate MV described above, a region within the other reference picture along the motion trajectory of the current block is used.

[0527] Figure 44is a diagram for explaining an example of the first style matching (bi-directional matching) between two blocks in two reference pictures along a motion trajectory. As shown in Figure 44 In the first style matching, two MVs (MV0, MV1) are derived by searching the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block). Specifically, for the current block, the difference between the reconstructed image at the specified position in the first coded reference picture (Ref0) designated by the candidate MV and the reconstructed image at the specified position in the second coded reference picture (Ref1) designated by the symmetric MV of the candidate MV scaled by the display time interval is derived, and the evaluation value is calculated using the resulting difference value. The candidate MV whose evaluation value is the best value can be selected as the best candidate MV among a plurality of candidate MVs.

[0528] Under the assumption of the continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, in the case where the current picture is located between the two reference pictures in time and the temporal distance from the current picture to the two reference pictures is equal, in the first style matching, the mirror-symmetric bi-directional MVs are derived.

[0529] [MV derivation > FRUC > template matching]

[0530] In the second style matching (template matching), the style matching is performed between a template (a block adjacent to the current block (e.g., the upper and / or left adjacent block) in the current picture) in the current picture and a block in the reference picture. Thus, in the second style matching, as the specified region for the calculation of the evaluation value for the candidate MV, the block adjacent to the current block in the current picture is used.

[0531] Figure 45 is a diagram for explaining an example of the style matching (template matching) between a template in the current picture and a block in the reference picture. As shown in Figure 45 In the second style matching, the MV of the current block is derived by searching the most matching block in the reference picture (Ref0) to the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed images of the left and upper adjacent blocks or one of them and the reconstructed image at the equivalent position in the coded reference picture (Ref0) designated by the candidate MV is derived, and the evaluation value is calculated using the resulting difference value. The candidate MV whose evaluation value is the best value can be selected as the best candidate MV among a plurality of candidate MVs.

[0532] Information indicating whether FRUC mode is used (e.g., referred to as the FRUC flag) is signaled at the CU level. Furthermore, when FRUC mode is used (e.g., when the FRUC flag is true), information indicating the available pattern matching method (first pattern matching or second pattern matching) is signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).

[0533] [MV Export > Affine Mode]

[0534] Affine mode is a mode that uses affine transformations to generate motion representations (MVs). For example, MVs can also be derived on a sub-block basis based on the MVs of multiple adjacent blocks. This mode is sometimes referred to as affine motion compensation prediction mode.

[0535] Figure 46A This is a diagram illustrating an example of deriving the MV from sub-block units based on the MV of multiple adjacent blocks. Figure 46A In this context, the current block may consist of 16 sub-blocks, each composed of 4×4 pixels. Here, the motion vector v0 of the top-left control point of the current block is derived based on the MV of adjacent blocks, and similarly, the motion vector v1 of the top-right control point of the current block is derived based on the MV of adjacent sub-blocks. Then, according to the following equation (1A), the two motion vectors v0 and v1 are projected, and the motion vectors (v0, v1, v1) of each sub-block within the current block are derived. x v y ).

[0536]

Formula 1

[0537]

[0538] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents a predetermined weighting coefficient.

[0539] Information representing this affine pattern (e.g., referred to as an affine flag) can be signaled at the CU level. Furthermore, the signaling of information representing this affine pattern is not limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).

[0540] Furthermore, such affine modes can also include several modes with different methods for deriving the MV of the upper left and upper right control points. For example, in affine mode, there are two modes: affine inter-frame (also known as affine normal inter-frame) mode and affine merge mode.

[0541] Figure 46B is a diagram for explaining an example of the derivation of the MV of a sub-block unit in the affine mode using 3 control points. In Figure 46B , the current block includes 16 sub-blocks of 4x4 pixels. Here, the motion vector v0 of the top-left corner control point of the current block is derived based on the MVs of the neighboring blocks. Likewise, the motion vector v1 of the top-right corner control point of the current block is derived based on the MVs of the neighboring blocks, and the motion vector v2 of the bottom-left corner control point of the current block is derived based on the MVs of the neighboring blocks. Then, the 3 motion vectors v0, v1, and v2 are projected to derive the motion vectors (v x , v y ) of the respective sub-blocks within the current block according to the following equation (1B).

[0542] [Equation 2]

[0543]

[0544] Here, x and y represent the horizontal position and the vertical position of the center of the sub-block, respectively, and w and h represent predetermined weight coefficients. Also, w can represent the width of the current block, and h can represent the height of the current block.

[0545] The affine mode using mutually different numbers of control points (for example, 2 and 3) can also be signaled by switching at the CU level. Also, information indicating the number of control points of the affine mode used at the CU level can be signaled at another level (for example, the sequence level, the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0546] Also, in such an affine mode having 3 control points, several modes different in the derivation method of the MV of the top-left, top-right, and bottom-left corner control points can be included. For example, in the affine mode having 3 control points, there are 2 modes, the affine inter mode and the affine merge mode, as in the affine mode having 2 control points.

[0547] Also, in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels, and can be another size. For example, the size of each sub-block can be 8x8 pixels.

[0548] [MV derivation > affine mode > control point]

[0549] Figure 47A 、 Figure 47B and Figure 47C is a conceptual diagram for explaining an example of the derivation of the MV of a control point in the affine mode.

[0550] In the affine mode, as in Figure 47AAs shown, for example, the prediction MV of each of the control points of the current block is calculated based on the MVs corresponding to the coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) neighboring the current block. Specifically, the coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) are examined in this order to determine the first valid block coded in affine mode. The MV of the control point of the current block is calculated based on the MV corresponding to the determined block.

[0551] For example, as shown in FIG. 6, in the case that the block A neighboring the left side of the current block is coded in affine mode with 2 control points, the motion vectors v3 and v4 projected to the positions of the top-left corner and the top-right corner of the coded block containing the block A are derived. Then, the motion vector v0 of the top-left control point and the motion vector v1 of the top-right control point of the current block are calculated according to the derived motion vectors v3 and v4. Figure 47B

[0552] For example, as shown in FIG. 7, in the case that the block A neighboring the left side of the current block is coded in affine mode with 3 control points, the motion vectors v3, v4, and v5 projected to the positions of the top-left corner, the top-right corner, and the bottom-left corner of the coded block containing the block A are derived. Then, the motion vector v0 of the top-left control point, the motion vector v1 of the top-right control point, and the motion vector v2 of the bottom-left control point of the current block are calculated according to the derived motion vectors v3, v4, and v5. Figure 47C

[0553] In addition, the method of deriving the MVs shown in FIG. 6 and FIG. 7 can be used in the derivation of the prediction MVs of the control points of the current block in the step Sj_1 shown in FIG. 8. Figure 47A-47C Figure 50 Figure 51

[0554] Figure 48A Figure 48B is a conceptual diagram for explaining another example of the derivation of the control point MVs in affine mode.

[0555] Figure 48A is a diagram for explaining affine mode with 2 control points.

[0556] In this affine mode, as shown in FIG. 4, the prediction MV of each of the control points of the current block is calculated based on the MVs corresponding to the coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) neighboring the current block. Specifically, the coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) are examined in this order to determine the first valid block coded in affine mode. The MV of the control point of the current block is calculated based on the MV corresponding to the determined block. Figure 48A ​​​​​​As shown, the MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block.

[0557] Figure 48B This is a diagram used to illustrate an affine pattern with three control points.

[0558] In this affine mode, such as Figure 48B As shown, the MV selected from the MVs of the previously encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the top-left control point of the current block. Similarly, the MV selected from the MVs of the previously encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the top-right control point of the current block. Furthermore, the MV selected from the MVs of the previously encoded blocks F and G adjacent to the current block is used as the motion vector v2 of the bottom-left control point of the current block.

[0559] in addition, Figure 48A and Figure 48B The method for exporting the MV shown can be used in the following sections. Figure 50 The deriving of the MV of each control point of the current block in step Sk_1 shown can also be used for the following description. Figure 51 The step Sj_1 is to derive the predicted MV for each control point of the current block.

[0560] Here, for example, in the case of signaling by switching different numbers of control points (e.g., 2 and 3) in affine mode at the CU level, the number of control points may sometimes differ depending on the encoded block and the current block.

[0561] Figure 49A and Figure 49B This is a conceptual diagram illustrating an example of a method for deriving the MV of control points when the number of control points differs between an encoded block and the current block.

[0562] For example, such as Figure 49A As shown, the current block has three control points: the top-left corner, the top-right corner, and the bottom-left corner. Block A, adjacent to the left of the current block, is encoded in an affine pattern with two control points. In this case, motion vectors v3 and v4 are derived, projected onto the top-left and top-right corners of the encoded block containing block A. Then, based on the derived motion vectors v3 and v4, the motion vector v0 for the top-left control point and the motion vector v1 for the top-right control point of the current block are calculated. Furthermore, based on the derived motion vectors v0 and v1, the motion vector v2 for the bottom-left control point is calculated.

[0563] For example, such asFigure 49B As shown, the current block has two control points, the top left and the top right. Block A, which is adjacent to the left of the current block, is encoded in an affine pattern with three control points. In this case, motion vectors v3, v4, and v5 are derived, projected onto the top left, top right, and bottom left corners of the encoded block containing block A. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the top left control point and the motion vector v1 of the top right control point of the current block are calculated.

[0564] in addition, Figure 49A and Figure 49B The method for exporting the MV shown can be used in the following sections. Figure 50 The deriving of the MV of each control point of the current block in step Sk_1 shown can also be used for the following description. Figure 51 The step Sj_1 is to derive the predicted MV for each control point of the current block.

[0565] [MV Export > Affine Mode > Affine Merge Mode]

[0566] Figure 50 This is a flowchart representing an example of an affine merge pattern.

[0567] In affine merging mode, firstly, the inter-frame prediction unit 126 derives the MV of each control point for the current block (step Sk_1). The control points are as follows: Figure 46A As shown, these are the top left and top right corners of the current block, or as... Figure 46B The diagram shows the top-left, top-right, and bottom-left corners of the current block. At this time, the inter-frame prediction unit 126 can also encode MV selection information used to identify the derived two or three MVs into the stream.

[0568] For example, in use Figure 47A-47C In the case of the MV export method shown, such as Figure 47A As shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) and determines the initial valid blocks encoded in affine mode.

[0569] The inter-frame prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control points. For example, when block A is determined and block A has 2 control points, as follows... Figure 47BAs shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the top-left control point and the motion vector v1 of the top-right control point of the current block based on the motion vectors v3 and v4 of the top-left and top-right corners of the encoded block containing block A. For example, by projecting the motion vectors v3 and v4 of the top-left and top-right corners of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the top-left control point and the motion vector v1 of the top-right control point of the current block.

[0570] Alternatively, if block A is determined and block A has 3 control points, such as Figure 47C As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block based on the motion vectors v3, v4, and v5 of the upper left, upper right, and lower left corners of the encoded block containing block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left, upper right, and lower left corners of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block.

[0571] Alternatively, it can be as described above. Figure 49A As shown, block A is determined, and the MV of the three control points is calculated when block A has two control points. This can also be done as described above. Figure 49B As shown, block A is determined, and if block A has 3 control points, the MV of 2 control points is calculated.

[0572] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks contained in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 126 calculates the MV of the sub-block as an affine MV using two motion vectors v0 and v1 and the above-described equation (1A), or using three motion vectors v0, v1, and v2 and the above-described equation (1B) (step Sk_2). Then, the inter-frame prediction unit 126 performs motion compensation on the sub-block using the affine MV and the encoded reference image (step Sk_3). When steps Sk_2 and Sk_3 are performed on all sub-blocks contained in the current block respectively, the process of generating the prediction image of the current block using the affine merging mode ends. That is, motion compensation is performed on the current block, and the prediction image of the current block is generated.

[0573] Furthermore, in step Sk_1, the aforementioned candidate MV list can also be generated. The candidate MV list could, for example, be a list containing candidate MVs exported using multiple MV export methods for each control point. Multiple MV export methods could be... Figure 47A-47C The method for exporting the MV shown. Figure 48A andFigure 48B the derivation method of the MV shown in Figure 49A and Figure 49B any combination of the derivation method of the MV shown in and the derivation method of other MVs.

[0574] In addition, the candidate MV list can also include a candidate MV of a mode other than the affine mode that performs prediction in a sub-block unit.

[0575] In addition, as the candidate MV list, for example, a candidate MV list including a candidate MV of the affine merge mode with 2 control points and a candidate MV of the affine merge mode with 3 control points can be generated. Alternatively, a candidate MV list including a candidate MV of the affine merge mode with 2 control points and a candidate MV list including a candidate MV of the affine merge mode with 3 control points can be generated respectively. Alternatively, a candidate MV list including a candidate MV of one of the affine merge mode with 2 control points and the affine merge mode with 3 control points can be generated. The candidate MV can be, for example, an MV of a coded block A (left), a block B (above), a block C (upper right), a block D (lower left), and a block E (upper left), and an MV of a valid block among these blocks.

[0576] Further, as the MV selection information, an index indicating which candidate MV in the candidate MV list can be transmitted.

[0577] [MV derivation > affine mode > affine inter mode]

[0578] Figure 51 is a flowchart indicating an example of the affine inter mode.

[0579] In the affine inter mode, first, the inter prediction section 126 derives a prediction MV (v0, v1) or (v0, v1, v2) of each of 2 or 3 control points of the current block (step Sj_1). As shown in Figure 46A or Figure 46B The control points are, for example, a point of the upper left corner, the upper right corner, or the lower left corner of the current block.

[0580] For example, in the case of using the derivation method of the MV shown in Figure 48A and Figure 48B The inter prediction section 126 derives the prediction MV (v0, v1) or (v0, v1, v2) of each of the control points of the current block by selecting an MV of a certain block among coded blocks in the vicinity of each of the control points of the current block shown in Figure 48A or Figure 48B At this time, the inter prediction section 126 encodes prediction MV selection information for identifying the selected 2 or 3 prediction MVs into the stream.

[0581] For example, the inter prediction section 126 can decide which MV of the coded blocks neighboring the current block is used as the prediction MV of the control point by using a cost evaluation or the like, and can describe a flag indicating which prediction MV is selected in the bit stream. That is, the inter prediction section 126 outputs the flag or the like prediction MV selection information as the prediction parameter to the entropy encoding section 110 via the prediction parameter generation section 130.

[0582] Next, the inter prediction section 126 performs a motion search (steps Sj_3 and Sj_4) while updating the prediction MVs selected or derived in the respective update steps Sj_1 (step Sj_2). That is, the inter prediction section 126 calculates (step Sj_3) the MVs of the respective sub-blocks corresponding to the prediction MVs to be updated as affine MVs using the above-described equation (1A) or equation (1B). Then, the inter prediction section 126 performs motion compensation on the respective sub-blocks using these affine MVs and the coded reference picture (step Sj_4). The processing of steps Sj_3 and Sj_4 is performed on all the blocks within the current block every time the prediction MVs are updated in step Sj_2. As a result, in the motion search loop, the inter prediction section 126 decides, for example, the prediction MV that can result in the minimum cost as the MV of the control point (step Sj_5). At this time, the inter prediction section 126 also encodes the difference between this decided MV and the prediction MV as the difference MV into the stream. That is, the inter prediction section 126 outputs the difference MV as the prediction parameter to the entropy encoding section 110 via the prediction parameter generation section 130.

[0583] Finally, the inter prediction section 126 generates a prediction image of the current block by performing motion compensation on the current block using the decided MV and the coded reference picture (step Sj_6).

[0584] In addition, in step Sj_1, the above-described candidate MV list can also be generated. The candidate MV list can be, for example, a list containing candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be Figures 47A-47C the MV derivation method shown in FIG. 1, Figure 48A and Figure 48B the MV derivation method shown in FIG. 1, Figure 49A and Figure 49B the MV derivation method shown in FIG. 1, and an arbitrary combination of other MV derivation methods.

[0585] In addition, the candidate MV list can contain candidate MVs of modes other than the affine mode that perform prediction in the sub-block unit.

[0586] In addition, as the candidate MV list, a candidate MV list including a candidate MV of the affine inter mode with 2 control points and a candidate MV of the affine inter mode with 3 control points can be generated. Alternatively, a candidate MV list including a candidate MV of the affine inter mode with 2 control points and a candidate MV of the affine inter mode with 3 control points can be generated respectively. Alternatively, a candidate MV list including a candidate MV of one of the affine inter mode with 2 control points and the affine inter mode with 3 control points can be generated. The candidate MV can be, for example, an MV of the coded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), and an MV of an effective block among these blocks.

[0587] In addition, as the prediction MV selection information, an index indicating which of the candidate MVs in the candidate MV list can be sent out.

[0588] [MV derivation > triangle mode]

[0589] In the above example, the inter prediction section 126 generates one prediction image of a rectangle for the rectangular current block. However, the inter prediction section 126 can generate a plurality of prediction images of shapes other than a rectangle for the rectangular current block, and generate a final prediction image of a rectangle by combining the plurality of prediction images. The shape other than a rectangle can be, for example, a triangle.

[0590] Figure 52A is a diagram for explaining generation of prediction images of 2 triangles.

[0591] The inter prediction section 126 generates a prediction image of a triangle for a first partition of the triangle within the current block using a first MV of the first partition. Likewise, the inter prediction section 126 generates a prediction image of a triangle for a second partition of the triangle within the current block using a second MV of the second partition. Then, the inter prediction section 126 combines the prediction images to generate a prediction image of a rectangle identical to the current block.

[0592] In addition, as the prediction image of the first partition, a first prediction image of a rectangle corresponding to the current block can be generated using the first MV. In addition, as the prediction image of the second partition, a second prediction image of a rectangle corresponding to the current block can be generated using the second MV. The prediction image of the current block can be generated by weightedly adding the first prediction image and the second prediction image. In addition, the portion to be weightedly added can be only a partial region sandwiching the boundary between the first partition and the second partition.

[0593] Figure 52Bis a conceptual diagram indicating an example of a first portion of a first partition that overlaps with a second partition, and a first sample set and a second sample set that are obtained by weighting as part of a correction process. The first portion can be, for example, one fourth of the width or height of the first partition. In another example, the first portion can have a width corresponding to N samples adjacent to an edge of the first partition. Here, N is an integer greater than zero, and for example, N can be the integer 2. Figure 52B indicates a rectangular partition of a rectangular portion having a width of one fourth of the width of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion. Figure 52B the center of indicates a triangular partition of a polygonal portion having a height of one fourth of the height of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion. Figure 52B the right of indicates a triangular partition of a polygonal portion having a height corresponding to 2 samples. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion.

[0594] The first portion can be a portion of the first partition that overlaps with an adjacent partition. Figure 52C is a conceptual diagram indicating a first portion of a first partition that is a portion of the first partition overlapping with a portion of an adjacent partition. For simplicity of explanation, a rectangular partition having a portion overlapping with a spatially adjacent rectangular partition is shown. A partition having another shape such as a triangular partition can be used, and the overlapping portion can overlap with a spatially or temporally adjacent partition.

[0595] In addition, an example in which a prediction image is generated for each of the two partitions using inter prediction is shown, but a prediction image can be generated for at least one of the partitions using intra prediction.

[0596] Figure 53 is a flowchart indicating an example of a triangular mode.

[0597] In the triangular mode, first, the inter prediction section 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter prediction section 126 can encode partition information related to the division into each partition as a prediction parameter to the stream. That is, the inter prediction section 126 can output the partition information as a prediction parameter to the entropy encoding section 110 via the prediction parameter generation section 130.

[0598] Next, the inter prediction section 126 first acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of coded blocks located around the current block in time or in space (step Sx_2). That is, the inter prediction section 126 creates a candidate MV list.

[0599] Then, the inter prediction section 126 selects a candidate MV of the first partition and a candidate MV of the second partition as the first MV and the second MV, respectively, from the plurality of candidate MVs acquired in step Sx_1 (step Sx_3). At this time, the inter prediction section 126 can also encode MV selection information for identifying the selected candidate MVs as the prediction parameters to the stream. That is, the inter prediction section 126 can output the MV selection information as the prediction parameters to the entropy coding section 110 via the prediction parameter generation section 130.

[0600] Next, the inter prediction section 126 generates a first prediction image by performing motion compensation using the selected first MV and the coded reference picture (step Sx_4). Similarly, the inter prediction section 126 generates a second prediction image by performing motion compensation using the selected second MV and the coded reference picture (step Sx_5).

[0601] Finally, the inter prediction section 126 generates a prediction image of the current block by weightedly adding the first prediction image and the second prediction image (step Sx_6).

[0602] In addition, in the example shown in FIG. 8, the first partition and the second partition are each a triangle, but can be a trapezoid, and can each be a different shape from the other. Also, in the example shown in FIG. 8, the current block is composed of two partitions, but can be composed of three or more partitions. Figure 52A Figure 52A In addition, the first partition and the second partition can be repeated. That is, the first partition and the second partition can include the same pixel region. In this case, a prediction image of the current block can be generated using a prediction image in the first partition and a prediction image in the second partition.

[0603] In addition, in this example, an example in which a prediction image is generated by inter prediction in both the first partition and the second partition is shown, but a prediction image can be generated by intra prediction for at least one of the partitions.

[0604] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV can be different, or can be the same candidate MV list.

[0605] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV can be different, or can be the same candidate MV list.

[0606] ​Further, the partition information can include an index indicating a split direction in which the current block is split into a plurality of partitions. The MV selection information can also include an index indicating the selected first MV and an index indicating the selected second MV. One index can indicate a plurality of information. For example, one index that collectively indicates a part or all of the partition information and a part or all of the MV selection information can be encoded.

[0607] [MV derivation > ATMVP mode]

[0608] Figure 54 is a diagram indicating an example of the ATMVP mode in which the MV is derived in the sub-block unit.

[0609] The ATMVP mode is a mode classified as the merge mode. For example, in the ATMVP mode, a candidate MV in the sub-block unit is registered in the candidate MV list used for the normal merge mode.

[0610] Specifically, in the ATMVP mode, first, as shown in Figure 54 , a temporal MV reference block that has correspondence with the current block is determined in an encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block. Next, for each sub-block within the current block, an MV used at the time of encoding of a region within the temporal MV reference block corresponding to the sub-block is determined. The MV thus determined is included in the candidate MV list as a candidate MV of the sub-block of the current block. In a case where such a candidate MV of each sub-block is selected from the candidate MV list, motion compensation is performed on the sub-block using the candidate MV as the MV of the sub-block. Thereby, a predicted image of each sub-block is generated.

[0611] In addition, in the example shown in Figure 54 , the block adjacent to the lower left of the current block is used as the peripheral MV reference block, but a block other than this can also be used. In addition, the size of the sub-block can be 4x4 pixels, 8x8 pixels, or a size other than this. The size of the sub-block can also be switched in units of a slice, a tile, or a picture, and the like.

[0612] [Motion search > DMVR]

[0613] Figure 55 is a diagram indicating the relationship of the merge mode and the DMVR.

[0614] The inter prediction section 126 derives the MV of the current block in the merge mode (step Sl_l). Next, the inter prediction section 126 determines whether or not to perform the MV search, that is, the motion search (step Sl_2). Here, when it is determined not to perform the motion search (NO in step Sl_2), the inter prediction section 126 decides the MV derived in step Sl_l as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is decided in the merge mode.

[0615] On the other hand, when it is determined to perform the motion search in step Sl_l (YES in step Sl_2), the inter prediction section 126 derives the final MV for the current block by searching the surrounding area of the reference block included in the reference picture represented by the MV derived in step Sl_l (step Sl_3). That is, in this case, the MV of the current block is decided by the DMVR.

[0616] Figure 56 is a conceptual diagram for explaining an example of the DMVR for deciding the MV.

[0617] First, for example in the merge mode, candidate MVs (L0 and Ll) are selected for the current block. Then, in accordance with the candidate MV (L0), the reference pixels are determined from the coded picture of the L0 list, that is, the 1st reference picture (L0). Likewise, in accordance with the candidate MV (Ll), the reference pixels are determined from the coded picture of the Ll list, that is, the 2nd reference picture (Ll). A template is generated by taking the average of these reference pixels.

[0618] Next, using the template, the surrounding area of the candidate MV of the 1st reference picture (L0) and the 2nd reference picture (Ll) is searched, respectively, and the MV with the minimum cost is decided as the final MV of the current block. Further, the cost can be calculated using, for example, the difference value of each pixel value of the template and each pixel value of the search area, and the candidate MV value, and the like.

[0619] Even if it is not the process itself explained here, as long as it is a process capable of deriving the final MV by searching the surrounding of the candidate MV, any process can be used.

[0620] Figure 57 is a conceptual diagram for explaining another example of the DMVR for deciding the MV. Figure 57 The present example shown in Figure 56 differs from the example of the DMVR shown in

[0621] First, the inter prediction section 126 searches the surrounding of the reference block included in the reference picture of the L0 list and the Ll list, based on the candidate MV, that is, the initial MV, acquired from the candidate MV list. For example, as shown in Figure 57The initial MV corresponding to the reference block of the L0 list is denoted as InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is denoted as InitMV_L1. In the motion search, the inter prediction section 126 first sets the search position for the reference picture in the L0 list. The difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to the search position is denoted as MVd_L0. Then, the inter prediction section 126 determines the search position in the reference picture of the L1 list. The search position is denoted by the difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. Specifically, the inter prediction section 126 determines the difference vector as MVd_L1 by mirroring MVd_L0. That is, the inter prediction section 126 sets the position symmetrical to the position indicated by the initial MV as the search position in the reference picture of each of the L0 list and the L1 list. The inter prediction section 126 calculates the sum of absolute differences (SAD) or the like of the pixel values in the block at each search position as the cost, and finds the search position with the minimum cost.

[0622] Figure 58A is a diagram illustrating an example of the motion search in DMVR, Figure 58B is a flowchart illustrating an example of the motion search.

[0623] First, the inter prediction section 126 calculates the costs of the search position indicated by the initial MV (also referred to as the starting point) and eight search positions located around the search position in Step 1. Then, the inter prediction section 126 determines whether the cost of the search position other than the starting point is the minimum. Here, when it is determined that the cost of the search position other than the starting point is the minimum, the inter prediction section 126 moves to the search position with the minimum cost and performs the processing of Step 2. On the other hand, if the cost of the starting point is the minimum, the inter prediction section 126 skips the processing of Step 2 and performs the processing of Step 3.

[0624] In Step 2, the inter prediction section 126 performs the same search as the processing of Step 1 using the search position moved according to the result of the processing of Step 1 as a new starting point. Then, the inter prediction section 126 determines whether the cost of the search position other than the starting point is the minimum. Here, if the cost of the search position other than the starting point is the minimum, the inter prediction section 126 performs the processing of Step 4. On the other hand, if the cost of the starting point is the minimum, the inter prediction section 126 performs the processing of Step 3.

[0625] In Step 4, the inter prediction section 126 processes the search position of the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0626] In Step 3, the inter prediction section 126 decides a decimal precision pixel position in which the cost is the smallest, based on the costs at the 4 points located above, below, left and right of the start point in Step 1 or Step 2, and sets the pixel position as a final search position. The decimal precision pixel position is decided by weightedly adding vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) at the 4 points above, below, left and right, with the costs at the respective search positions of the 4 points as weights. Then, the inter prediction section 126 decides a difference between the position indicated by the initial MV and the final search position as a difference vector.

[0627] [Motion compensation > BIO / OBMC / LIC]

[0628] In the motion compensation, there are modes of generating a prediction image and correcting the prediction image. The modes are, for example, BIO, OBMC and LIC described later.

[0629] Figure 59 is a flowchart showing an example of generation of a prediction image.

[0630] The inter prediction section 126 generates a prediction image (Step Sm_1), and corrects the prediction image by any one of the above modes (Step Sm_2).

[0631] Figure 60 is a flowchart showing another example of generation of a prediction image.

[0632] The inter prediction section 126 derives an MV of a current block (Step Sn_1). Next, the inter prediction section 126 generates a prediction image using the MV (Step Sn_2), and determines whether or not to perform a correction process (Step Sn_3). Here, when it is determined to perform the correction process (Yes in Step Sn_3), the inter prediction section 126 generates a final prediction image by correcting the prediction image (Step Sn_4). In addition, in LIC described later, it is also possible to correct the luminance and color differences in Step Sn_4. On the other hand, when it is determined not to perform the correction process (No in Step Sn_3), the inter prediction section 126 outputs the prediction image as a final prediction image without correcting the prediction image (Step Sn_5).

[0633] [Motion compensation > OBMC]

[0634] Not only the motion information of the current block obtained by the motion search, but also the motion information of the neighboring block can be used to generate the inter prediction image. Specifically, the inter prediction image can also be generated in the sub-block unit within the current block by weighted adding a prediction image based on the motion information obtained by the motion search (within the reference picture) and a prediction image based on the motion information of the neighboring block (within the current picture). Such inter prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or OBMC mode.

[0635] In the OBMC mode, information indicating the size of the sub-block for OBMC (e.g., referred to as OBMC block size) can also be signaled at the sequence level. Also, information indicating whether to apply the OBMC mode (e.g., referred to as OBMC flag) can also be signaled at the CU level. In addition, the level of the signaling of these information does not need to be limited to the sequence level and the CU level, but can be other levels (e.g., the picture level, the slice level, the tile level, the CTU level, or the sub-block level).

[0636] The OBMC mode is described in more detail. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process based on OBMC.

[0637] First, as shown in Figure 62 , a prediction image (Pred) based on the usual motion compensation is obtained using the MV assigned to the current block. In Figure 62 , the arrow "MV" points to the reference picture, and indicates which block of the current picture refers to in order to obtain the prediction image.

[0638] Next, the MV (MV_L) already derived for the coded left neighboring block is applied (reused) to the current block to obtain a prediction image (Pred_L). The MV (MV_L) is indicated by the arrow "MV_L" pointing from the current block to the reference picture. Then, the 1st correction of the prediction image is performed by overlapping the 2 prediction images Pred and Pred_L. This has the effect of blending the boundaries between the neighboring blocks.

[0639] Likewise, the MV (MV_U) that has been derived for the coded upper neighboring block is applied (reused) to the current block, resulting in a prediction picture (Pred_U). The MV (MV_U) is indicated by the arrow "MV_U" pointing from the current block to the reference picture. Then, a second correction of the prediction picture is performed by overlaying the prediction picture Pred_U with the prediction picture that has been corrected for the first time (e.g. Pred and Pred_L). This has the effect of blending the boundaries between neighboring blocks. The prediction picture resulting from the second correction is the final prediction picture for the current block with the boundaries to the neighboring blocks being blended (smoothed).

[0640] Further, the above-described example is a two-path correction method using the left and upper neighboring blocks, but the correction method can also be a three-path or more path correction method using the right and / or lower neighboring blocks as well.

[0641] In addition, the region to which the overlay is performed can also not be the entire pixel region of the block, but only a partial region near the block boundary.

[0642] In addition, the prediction picture correction processing of the OBMC is described herein, which is used to obtain one prediction picture Pred by overlaying one reference picture with the additional prediction pictures Pred_L and Pred_U. However, in the case where the prediction picture is corrected based on multiple reference pictures, the same processing can also be applied to the multiple reference pictures respectively. In this case, by performing the OBMC picture correction based on multiple reference pictures, after the corrected prediction pictures are obtained from the respective reference pictures, the final prediction picture is obtained by further overlaying the obtained multiple corrected prediction pictures.

[0643] In addition, in the OBMC, the unit of the current block can be the PU unit, or the sub-block unit obtained by further dividing the PUs.

[0644] As a method of determining whether or not to apply the OBMC, for example, there is a method of using a signal indicating whether or not to apply the OBMC, i.e. obmc_flag. As a specific example, the encoding apparatus 100 can determine whether or not the current block belongs to a region of complex motion. The encoding apparatus 100 encodes by setting the value 1 to the obmc_flag and applying the OBMC in the case where it belongs to the region of complex motion, and encodes without applying the OBMC in the case where it does not belong to the region of complex motion, by setting the value 0 to the obmc_flag. On the other hand, in the decoding apparatus 200, by decoding the obmc_flag described in the stream, whether or not to apply the OBMC is switched according to the value.

[0645] [Motion compensation > BIO]

[0646] Next, a method of deriving the MV will be described. First, a mode of deriving the MV based on a model assuming constant velocity straight line motion will be described. This mode is sometimes referred to as a bi-directional optical flow (BIO) mode. In addition, this bi-directional optical flow can be expressed as BDOF instead of BIO.

[0647] Figure 63 is a diagram for explaining the model assuming constant velocity straight line motion. In Figure 63 , (v x , v y ) represents a velocity vector, τ0, τ1 represent distances in time between a current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.

[0648] At this time, under the assumption of constant velocity straight line motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (2) holds.

[0649]

Equation 3

[0650]

[0651] Here, I(k) represents a luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation represents that the sum of (i) a time differential of the luminance value, (ii) a product of a horizontal component of a velocity in the horizontal direction and a spatial gradient of the reference image, and (iii) a product of a vertical component of a velocity in the vertical direction and a spatial gradient of the reference image is equal to zero. It can also be that, based on a combination of this optical flow equation and Hermite interpolation, a block unit motion vector obtained from a candidate MV list or the like is corrected in pixel units.

[0652] In addition, the MV can also be derived on the decoding apparatus 200 side by a method different from the derivation of the motion vector based on the model assuming constant velocity straight line motion. For example, the motion vector can also be derived in sub-block units based on MVs of a plurality of neighboring blocks.

[0653] Figure 64 is a flowchart representing an example of inter prediction according to BIO. In addition, Figure 65is a diagram showing an example of a functional configuration of the inter prediction section 126 that performs the inter prediction according to BIO.

[0654] As shown in Figure 65 , the inter prediction section 126 includes, for example, a memory 126a, an interpolated image derivation section 126b, a gradient image derivation section 126c, an optical flow derivation section 126d, a correction value derivation section 126e, and a prediction image correction section 126f. Note that the memory 126a can also be the frame memory 122.

[0655] The inter prediction section 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) that includes the current block. Then, the inter prediction section 126 derives a prediction image of the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0656] Next, the interpolated image derivation section 126b refers to the memory 126a and derives an interpolated image I 0 of the current block using the motion vector M0 and the reference picture L0. In addition, the interpolated image derivation section 126b refers to the memory 126a and derives an interpolated image I 1 of the current block using the motion vector M1 and the reference picture L1 (step Sy_2). Here, the interpolated image I 0 is an image included in the reference picture Ref0, which is derived for the current block, and the interpolated image I 1 is an image included in the reference picture Ref1, which is derived for the current block. The interpolated image I 0 and the interpolated image I 1 may each be the same size as the current block. Alternatively, in order to appropriately derive a gradient image described later, the interpolated image I 0 and the interpolated image I 1 may each be an image larger than the current block. Furthermore, the interpolated image I 0 and I 1 may include a prediction image derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1) and a motion compensation filter.

[0657] In addition, the gradient image derivation section 126c derives gradient images (Ix 0 , Ix 1 , Iy 0 , and Iy 1 of the current block from the interpolated image I 0 and the interpolated image I 1) (step Sy_3). Further, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivation section 126c can derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image only needs to indicate the amount of spatial variation of pixel values in the horizontal direction or the vertical direction.

[0658] Next, the optical flow derivation section 126d derives the optical flow (vx, vy) as the above-mentioned velocity vector using the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) in units of the plurality of sub-blocks that constitute the current block (step Sy_4). The optical flow is a coefficient that corrects the amount of spatial movement of a pixel and can also be referred to as a local motion estimation value, a correction motion vector, or a correction weight vector. As an example, the sub-block can be a sub-CU of 4x4 pixels. In addition, the derivation of the optical flow can not be performed in units of sub-blocks, but can be performed in units of pixels or other units.

[0659] Next, the inter prediction section 126 corrects the prediction image of the current block using the optical flow (vx, vy). For example, the correction value derivation section 126e derives a correction value of the value of a pixel included in the current block using the optical flow (vx, vy) (step Sy_5). Also, the prediction image correction section 126f can correct the prediction image of the current block using the correction value (step Sy_6). In addition, the correction value can be derived in units of each pixel, or can be derived in units of a plurality of pixels or sub-blocks.

[0660] Further, the processing flow of the BIO is not limited to the processing disclosed in Figure 64 . The processing can implement only a part of the processing disclosed in Figure 64 , can add or replace a different processing, or can be executed in a different processing order.

[0661] [Motion compensation > LIC]

[0662] Next, an example of a mode in which a prediction image (prediction) is generated using LIC (local illumination compensation) will be described.

[0663] Figure 66A is a diagram for explaining an example of a prediction image generation method in which LIC-based luminance correction processing is used. In addition, Figure 66Bis a flowchart showing an example of a method of generating a prediction image using the LIC.

[0664] First, the inter prediction section 126 derives MVs from the coded reference pictures, and acquires a reference image corresponding to the current block (step Sz_1).

[0665] Next, the inter prediction section 126 extracts information indicating how luminance values change in the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on luminance pixel values of the coded left-adjacent reference region (peripheral reference region) and the coded upper-adjacent reference region (peripheral reference region) in the current picture, and luminance pixel values at equivalent positions in the reference picture specified by the derived MVs. Then, the inter prediction section 126 calculates a luminance correction parameter using the information indicating how luminance values change (step Sz_3).

[0666] The inter prediction section 126 generates a prediction image for the current block by performing luminance correction processing on the reference image in the reference picture specified by the MVs using the luminance correction parameter (step Sz_4). That is, the prediction image as the reference image in the reference picture specified by the MVs is corrected based on the luminance correction parameter. In this correction, luminance can be corrected, or color difference can be corrected. That is, a correction parameter of color difference can be calculated using information indicating how color difference changes, and color difference correction processing can be performed.

[0667] In addition, Figure 66A The shape of the peripheral reference region in the above example can be used, or a shape other than this can be used.

[0668] Further, the processing of generating a prediction image from one reference picture is described here, but the same applies to the case of generating a prediction image from a plurality of reference pictures, and a prediction image can be generated after performing luminance correction processing on the reference images acquired from each of the reference pictures in the same manner as described above.

[0669] As a method of determining whether or not to apply the LIC, for example, there is a method of using lic_flag as a signal indicating whether or not to apply the LIC. As a specific example, in the encoding device 100, it is determined whether or not the current block belongs to a region in which luminance changes have occurred, and in the case where it belongs to a region in which luminance changes have occurred, a value of 1 is set as lic_flag, and the LIC is applied to perform encoding, and in the case where it does not belong to a region in which luminance changes have occurred, a value of 0 is set as lic_flag, and the LIC is not applied to perform encoding. On the other hand, in the decoding device 200, lic_flag described in the stream can be decoded, and whether or not to apply the LIC is switched according to the value to perform decoding.

[0670] Other methods for determining whether to apply LIC include determining whether LIC has been applied to surrounding blocks. As a specific example, when the current block is processed in merge mode, the inter-frame prediction unit 126 determines whether the surrounding encoded blocks selected during the MV derivation in merge mode have been encoded using LIC. Based on this result, the inter-frame prediction unit 126 switches between applying LIC and encoding. Furthermore, in this example, the same processing also applies to the decoding device 200.

[0671] use Figure 66A and Figure 66B The LIC (Limited Light Correction Processing) has been explained; its details are explained below.

[0672] First, the inter-frame prediction unit 126 derives the MV from the reference image, which is an encoded image, to obtain the reference image corresponding to the current block.

[0673] Next, the inter-frame prediction unit 126 uses the luminance pixel values ​​of the coded peripheral reference regions (left and top adjacent) and the luminance pixel values ​​at the same position in the reference image specified by MV to extract information indicating how the luminance values ​​change in the reference image and the current image, and calculates luminance correction parameters for the current block. For example, the luminance pixel value of a pixel in the peripheral reference region of the current image is set to p0, and the luminance pixel value of a pixel in the peripheral reference region of the reference image at the same position is set to p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1+B=p0 as luminance correction parameters for multiple pixels in the peripheral reference region.

[0674] Next, the inter-frame prediction unit 126 performs brightness correction processing on the reference image within the reference image specified by MV using brightness correction parameters, and generates a prediction image for the current block. For example, the brightness pixel value in the reference image is set as p2, and the brightness pixel value of the brightness-corrected prediction image is set as p3. The inter-frame prediction unit 126 generates the brightness-corrected prediction image by calculating A×p2+B=p3 for each pixel in the reference image.

[0675] In addition, it can also be used Figure 66A A portion of the surrounding reference region shown. For example, a region containing a predetermined number of pixels, each spaced out from the upper adjacent pixel and the left adjacent pixel, can also be used as the surrounding reference region. Furthermore, the surrounding reference region is not limited to the region adjacent to the current block; it can also be a region not adjacent to the current block. Additionally, in Figure 66AIn the illustrated example, the peripheral reference region in the reference picture is a region specified by the MV of the current picture from the peripheral reference region in the current picture, but can also be a region specified by another MV. For example, the other MV can also be the MV of the peripheral reference region in the current picture.

[0676] In addition, the operation in the encoding device 100 is described here, but the operation in the decoding device 200 is the same.

[0677] Further, LIC can be applied not only to luminance but also to color difference. At this time, a correction parameter can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0678] Further, LIC can also be applied in sub-block units. For example, a correction parameter can be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0679] [Prediction control section]

[0680] The prediction control section 128 selects one of the intra-predicted image (pixels or signals output from the intra-prediction section 124) and the inter-predicted image (pixels or signals output from the inter-prediction section 126), and outputs the selected predicted image to the subtraction section 104 and the addition section 116.

[0681] [Prediction parameter generation section]

[0682] The prediction parameter generation section 130 can output, as a prediction parameter, information related to intra-prediction, inter-prediction, and selection of a predicted image in the prediction control section 128, to the entropy encoding section 110. The entropy encoding section 110 can generate a stream based on the prediction parameter input from the prediction parameter generation section 130 and the quantization coefficients input from the quantization section 108. The prediction parameter can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream and perform the same processing as the prediction processing performed in the intra-prediction section 124, the inter-prediction section 126, and the prediction control section 128. The prediction parameter can include a selected prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra-prediction section 124 or the inter-prediction section 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction section 124, the inter-prediction section 126, and the prediction control section 128.

[0683] [Decoding device]

[0684] Next, a decoding device 200 capable of decoding a stream output from the above-described encoding device 100 will be described. Figure 67is a block diagram showing an example of a functional configuration of a decoding apparatus 200 according to an embodiment. The decoding apparatus 200 is an apparatus that decodes an encoded image, i.e., a stream, in a block unit.

[0685] As shown in Figure 67 , the decoding apparatus 200 includes an entropy decoding section 202, an inverse quantization section 204, an inverse transform section 206, an addition section 208, a block memory 210, a loop filter section 212, a frame memory 214, an intra prediction section 216, an inter prediction section 218, a prediction control section 220, a prediction parameter generation section 222, and a partition decision section 224. In addition, the intra prediction section 216 and the inter prediction section 218 each constitute a part of a prediction processing section.

[0686] [Example of Installation of Decoding Apparatus]

[0687] Figure 68 is a block diagram showing an example of installation of the decoding apparatus 200. The decoding apparatus 200 includes a processor b1 and a memory b2. For example, as shown in Figure 67 , a plurality of constituent elements of the decoding apparatus 200 shown in Figure 68 are installed by the processor b1 and the memory b2.

[0688] The processor b1 is a circuit that performs information processing, and is a circuit that can access the memory b2. For example, the processor b1 is a special-purpose or general-purpose electronic circuit that decodes a stream. The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be a collection of a plurality of electronic circuits. In addition, for example, the processor b1 can also function as a plurality of constituent elements of the decoding apparatus 200 shown in Figure 67 and the like, except for constituent elements for storing information.

[0689] The memory b2 is a special-purpose or general-purpose memory that stores information used for the processor b1 to decode a stream. The memory b2 can be an electronic circuit, or can be connected to the processor b1. In addition, the memory b2 can be included in the processor b1. In addition, the memory b2 can be a collection of a plurality of electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disk, or can be represented as a storage or a recording medium. In addition, the memory b2 can be a nonvolatile memory, or can be a volatile memory.

[0690] For example, the memory b2 can store an image, or can store a stream. In addition, a program for the processor b1 to decode a stream can also be stored in the memory b2.

[0691] In addition, for example, the memory b2 can also function as a plurality of constituent elements of the decoding apparatus 200 shown in Figure 67the role of the constituent element for storing information among the plurality of constituent elements of the decoding apparatus 200 illustrated in FIG. 2. Specifically, the memory b2 can function as a memory for storing a reconstructed image (specifically, a reconstructed block or a reconstructed picture, or the like) and a memory for storing a decoded image (specifically, a decoded block or a decoded picture, or the like). Figure 67 the roles of the block memory 210 and the frame memory 214 illustrated in FIG. 2. More specifically, in the memory b2, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, or the like) can be stored.

[0692] In addition, in the decoding apparatus 200, all of the plurality of constituent elements illustrated in FIG. 2 can not be installed, and all of the plurality of processes described above can not be performed. Figure 67 In addition, in the decoding apparatus 200, all of the plurality of constituent elements illustrated in FIG. 2 can not be installed, and all of the plurality of processes described above can not be performed. Figure 67 In addition, in the decoding apparatus 200, all of the plurality of constituent elements illustrated in FIG. 2 can not be installed, and all of the plurality of processes described above can not be performed.

[0693] Hereinafter, after the flow of the overall process of the decoding apparatus 200 is described, each constituent element included in the decoding apparatus 200 is described. Furthermore, for each constituent element included in the decoding apparatus 200, a detailed description is omitted for the constituent element that performs the same process as the constituent element included in the encoding apparatus 100. For example, the inverse quantization section 204, the inverse transform section 206, the addition section 208, the block memory 210, the frame memory 214, the intra prediction section 216, the inter prediction section 218, the prediction control section 220, and the loop filter section 212 included in the decoding apparatus 200 respectively perform the same processes as the inverse quantization section 112, the inverse transform section 114, the addition section 116, the block memory 118, the frame memory 122, the intra prediction section 124, the inter prediction section 126, the prediction control section 128, and the loop filter section 120 included in the encoding apparatus 100.

[0694] [Overall flow of decoding process]

[0695] Figure 69 is a flowchart showing an example of the overall decoding process performed by the decoding apparatus 200.

[0696] First, the split decision section 224 of the decoding apparatus 200 decides the split pattern of each of a plurality of fixed-size blocks (128 x 128 pixels) included in a picture based on the parameters input from the entropy decoding section 202 (step Sp_1). The split pattern is the split pattern selected by the encoding apparatus 100. Then, the decoding apparatus 200 performs the processes of steps Sp_2 to Sp_6 on each of the plurality of blocks constituting the split pattern.

[0697] The entropy decoding section 202 decodes (specifically, entropy decodes) the encoded quantized coefficients and the prediction parameters of the current block (step Sp_2).

[0698] Next, the inverse quantization section 204 and the inverse transform section 206 reproduce the prediction residual of the current block by inverse quantizing and inverse transforming the plurality of quantized coefficients (step Sp_3).

[0699] Next, a prediction processing section composed of the intra prediction section 216, the inter prediction section 218, and the prediction control section 220 generates a prediction image of the current block (step Sp_4).

[0700] Next, the addition section 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction image to the prediction residual (step Sp_5).

[0701] Also, when the reconstructed image is generated, the loop filter 212 filters the reconstructed image (step Sp_6).

[0702] Then, the decoding apparatus 200 determines whether or not the decoding of the picture as a whole has been completed (step Sp_7), and in the case where it is determined that the decoding has not been completed (NO in step Sp_7), the processing from step Sp_1 is repeated.

[0703] Further, the processing of these steps Sp_1 to Sp_7 can be sequentially performed by the decoding apparatus 200, a plurality of processes of a part of these processes can be performed in parallel, or the order can be changed.

[0704] [Partitioning decision section]

[0705] Figure 70 is a diagram showing the relationship of the partitioning decision section 224 with other constituent elements. As an example, the partitioning decision section 224 can also perform the following processing.

[0706] The partitioning decision section 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further acquires parameters from the entropy decoding section 202. Also, the partitioning decision section 224 can decide the partitioning pattern of the fixed-size block on the basis of the block information and the parameters. Also, the partitioning decision section 224 can output information indicating the decided partitioning pattern to the inverse transform section 206, the intra prediction section 216, and the inter prediction section 218. The inverse transform section 206 can perform inverse transform of the transform coefficients on the basis of the partitioning pattern indicated by the information from the partitioning decision section 224. The intra prediction section 216 and the inter prediction section 218 can generate a prediction image on the basis of the partitioning pattern indicated by the information from the partitioning decision section 224.

[0707] [Entropy decoding section]

[0708] Figure 71 is a block diagram showing an example of the functional structure of the entropy decoding section 202.

[0709] The entropy decoding section 202 generates quantization coefficients, prediction parameters, and parameters related to a division pattern, and the like by entropy-decoding the stream. In this entropy decoding, CABAC is used, for example. Specifically, the entropy decoding section 202 has, for example, a binary arithmetic decoding section 202a, a context control section 202b, and a multivaluing section 202c. The binary arithmetic decoding section 202 arithmetically decodes the stream in binary signals using a context value derived by the context control section 202b. As with the context control section 110b of the encoding apparatus 100, the context control section 202b derives a context value, that is, a probability of occurrence of a binary signal, corresponding to a feature of a syntax element or a surrounding condition. The multivaluing section 202c performs multivaluing that transforms the binary signal output from the binary arithmetic decoding section 202a into a multivalued signal representing the quantization coefficients and the like described above. This multivaluing is performed in the manner described above with respect to binarization.

[0710] The entropy decoding section 202 outputs the quantization coefficients to the inverse quantization section 204 in a block unit. The entropy decoding section 202 can also output the prediction parameters included in the stream (refer to FIG. 6) to the intra prediction section 216, the inter prediction section 218, and the prediction control section 220. The intra prediction section 216, the inter prediction section 218, and the prediction control section 220 can perform the same prediction processing as the processing performed by the intra prediction section 124, the inter prediction section 126, and the prediction control section 128 on the encoding apparatus 100 side. Figure 1 ) in a block unit. The entropy decoding section 202 can also output the prediction parameters included in the stream (refer to FIG. 6) to the intra prediction section 216, the inter prediction section 218, and the prediction control section 220. The intra prediction section 216, the inter prediction section 218, and the prediction control section 220 can perform the same prediction processing as the processing performed by the intra prediction section 124, the inter prediction section 126, and the prediction control section 128 on the encoding apparatus 100 side.

[0711] [Entropy decoding section]

[0712] Figure 72 is a diagram showing the flow of CABAC in the entropy decoding section 202.

[0713] First, in CABAC in the entropy decoding section 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding section 202c and setting of an initial context value are performed. Then, the binary arithmetic decoding section 202c and the multivaluing section 202c perform arithmetic decoding and multivaluing on the encoded data of the CTU, for example. At this time, the context control section 202b performs updating of the context value each time arithmetic decoding is performed. Then, the context control section 202b causes the context value to be backed up as post-processing. This backed-up context value is used as an initial value for the context value for the next CTU, for example.

[0714] [Inverse quantization section]

[0715] The inverse quantization section 204 inversely quantizes the quantization coefficients of the current block that are input from the entropy decoding section 202. Specifically, the inverse quantization section 204 inversely quantizes the quantization coefficients of the current block on a per-quantization coefficient basis based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization section 204 outputs the inversely quantized quantization coefficients (i.e., the transform coefficients) of the current block to the inverse transform section 206.

[0716] Figure 73 is a block diagram that shows an example of the functional structure of the inverse quantization section 204.

[0717] The inverse quantization section 204 includes, for example, a quantization parameter generation section 204a, a predicted quantization parameter generation section 204b, a quantization parameter storage section 204d, and an inverse quantization processing section 204e.

[0718] Figure 74 is a flowchart that shows an example of the inverse quantization performed by the inverse quantization section 204.

[0719] As an example, the inverse quantization section 204 can perform the inverse quantization processing on a per-CU basis according to the flow shown in Figure 74 Specifically, the quantization parameter generation section 204a determines whether or not to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation section 204a acquires the differential quantization parameter of the current block from the entropy decoding section 202 (step Sv_12).

[0720] Next, the predicted quantization parameter generation section 204b acquires the quantization parameter of a processing unit other than the current block from the quantization parameter storage section 204d (step Sv_13). The predicted quantization parameter generation section 204b generates the predicted quantization parameter of the current block based on the acquired quantization parameter (step Sv_14).

[0721] Then, the quantization parameter generation section 204a adds the differential quantization parameter of the current block acquired from the entropy decoding section 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generation section 204b (step Sv_15). By this addition, the quantization parameter of the current block is generated. Furthermore, the quantization parameter generation section 204a stores the quantization parameter of the current block in the quantization parameter storage section 204d (step Sv_16).

[0722] Next, the inverse quantization processing section 204e inversely quantizes the quantization coefficients of the current block into the transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).

[0723] Further, the difference quantization parameter can be decoded at a bit sequence level, a picture level, a slice level, a tile level, or a CTU level. In addition, an initial value of the quantization parameter can be decoded at a sequence level, a picture level, a slice level, a tile level, or a CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the difference quantization parameter.

[0724] Further, the inverse quantization section 204 can have a plurality of inverse quantizers, and can perform inverse quantization of the quantized coefficients using an inverse quantization method selected from among a plurality of inverse quantization methods.

[0725] [Inverse transform section]

[0726] The inverse transform section 206 restores the prediction residual by performing inverse transform on the transform coefficients that are input from the inverse quantization section 204.

[0727] For example, in a case where the information read from the stream indicates that EMT or AMT is applied (for example, the AMT flag is true), the inverse transform section 206 performs inverse transform on the transform coefficients of the current block based on the information read indicating the transform type.

[0728] Further, for example, in a case where the information read from the stream indicates that NSST is applied, the inverse transform section 206 applies inverse retransform on the transform coefficients.

[0729] Figure 75 is a flowchart indicating an example of the processing performed by the inverse transform section 206.

[0730] For example, the inverse transform section 206 determines whether or not information indicating that orthogonal transform is not performed is present in the stream (step St_11). Here, when it is determined that the information is not present (NO in step St_11), the inverse transform section 206 acquires information indicating the transform type decoded by the entropy decoding section 202 (step St_12). Next, the inverse transform section 206 decides the transform type used in the orthogonal transform of the encoding apparatus 100 based on the information (step St_13). Then, the inverse transform section 206 performs inverse orthogonal transform using the decided transform type (step St_14).

[0731] Figure 76 is a flowchart indicating another example of the processing performed by the inverse transform section 206.

[0732] For example, the inverse transform unit 206 determines whether the transform size is equal to or smaller than a predetermined value (step Su 11). When it is determined that the transform size is equal to or smaller than the predetermined value (Yes in step Su 11), the inverse transform unit 206 acquires information indicating which one of one or more transform types included in the first transform type group is used by the encoding apparatus 100 from the entropy decoding unit 202 (step Su 12). In addition, such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.

[0733] The inverse transform unit 206 determines a transform type used in orthogonal transform in the encoding apparatus 100 on the basis of the information (step Su 13). Then, the inverse transform unit 206 inversely orthogonal-transforms the transform coefficients of the current block using the determined transform type (step Su 14). On the other hand, when it is determined that the transform size is not equal to or smaller than the predetermined value (No in step Su 11), the inverse transform unit 206 inversely orthogonal-transforms the transform coefficients of the current block using the second transform type group (step Su 15).

[0734] In addition, as an example, the inverse orthogonal transform performed by the inverse transform unit 206 can be implemented in accordance with the flowchart shown in FIG. 6. In addition, the information indicating the transform type used in orthogonal transform can not be decoded, and the inverse orthogonal transform can be performed using a predetermined transform type. In addition, specifically, the transform type is DST7 or DCT8 or the like, and an inverse transform basis function corresponding to the transform type is used in the inverse orthogonal transform. Figure 75 Figure 76

[0735] [Addition unit]

[0736] The addition unit 208 reconstructs the current block by adding the prediction residual as input from the inverse transform unit 206 to the prediction image as input from the prediction control unit 220. That is, a reconstructed image of the current block is generated. Furthermore, the addition unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.

[0737] [Block memory]

[0738] The block memory 210 is a storage unit that stores a block within the current picture referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the addition unit 208.

[0739] [Loop filter unit]

[0740] The loop filter unit 212 applies loop filtering to the reconstructed image generated by the addition unit 208, and outputs the reconstructed image on which the filtering has been performed to the frame memory 214 and a display apparatus or the like.

[0741] ​​In a case where the information indicating the on / off of the ALF read from the stream indicates the on of the ALF, one filter is selected from among the plurality of filters based on the direction and the activity of the gradient of the locality, and the selected filter is applied to the reconstructed image.

[0742] Figure 77 is a block diagram indicating an example of a functional structure of the loop filter 212. Further, the loop filter 212 has the same structure as the loop filter 120 of the encoding apparatus 100.

[0743] The loop filter 212, for example, as shown in Figure 77 , has a deblocking filter processing section 212a, an SAO processing section 212b, and an ALF processing section 212c. The deblocking filter processing section 212a applies the above-described deblocking filter processing to the reconstructed image. The SAO processing section 212b applies the above-described SAO processing to the reconstructed image after the deblocking filter processing. In addition, the ALF processing section 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. Further, the loop filter 212 can not have all the processing sections disclosed in Figure 77 , but can have only a part of the processing sections. In addition, the loop filter 212 can be a structure that performs the above-described each processing in a different order from the processing order disclosed in Figure 77

[0744] [Frame Memory]

[0745] The frame memory 214 is a storage section for storing a reference picture used in inter prediction, and is also called a frame buffer. Specifically, the frame memory 214 stores the reconstructed image after the filtering by the loop filter 212.

[0746] [Prediction Section (Intra Prediction Section / Inter Prediction Section / Prediction Control Section)]

[0747] Figure 78 is a flowchart indicating an example of processing performed by the prediction section of the decoding apparatus 200. Further, as an example, the prediction section is constituted by all or a part of the intra prediction section 216, the inter prediction section 218, and the prediction control section 220. The prediction processing section, for example, includes the intra prediction section 216 and the inter prediction section 218.

[0748] ​The prediction unit generates a prediction image of the current block (step Sq_1). The prediction image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, for example, there is an intra prediction signal or an inter prediction signal. Specifically, the prediction unit generates the prediction image of the current block using a reconstructed image that has been obtained by performing generation of a prediction image for another block, restoration of a prediction residual, and addition of the prediction images. The prediction unit of the decoding apparatus 200 generates the same prediction image as the prediction image generated by the prediction unit of the encoding apparatus 100. That is, the generation method of the prediction image used in these prediction units is common or corresponds to each other.

[0749] The reconstructed image can be, for example, an image of a reference picture or an image of a picture that contains the decoded block (i.e., the above-described another block) in the current picture, i.e., the current picture. The decoded block in the current picture is, for example, a neighboring block of the current block.

[0750] Figure 79 is a flowchart showing another example of the processing performed by the prediction unit of the decoding apparatus 200.

[0751] The prediction unit determines a method or mode for generating a prediction image (step Sr_1). For example, the method or mode can be determined based on, for example, a prediction parameter or the like.

[0752] In a case where it is determined that the first method is the mode for generating a prediction image, the prediction unit generates a prediction image in accordance with the first method (step Sr_2a). In addition, in a case where it is determined that the second method is the mode for generating a prediction image, the prediction unit generates a prediction image in accordance with the second method (step Sr_2b). In addition, in a case where it is determined that the third method is the mode for generating a prediction image, the prediction unit generates a prediction image in accordance with the third method (step Sr_2c).

[0753] The first method, the second method, and the third method are mutually different methods for generating a prediction image, and can be, for example, an inter prediction method, an intra prediction method, and another prediction method. In such a prediction method, the above-described reconstructed image can also be used.

[0754] Figure 80A and Figure 80B is a flowchart showing another example of the processing performed in the prediction unit of the decoding apparatus 200.

[0755] As an example, the prediction unit can also perform the prediction processing in accordance with the flow shown in Figure 80A and Figure 80B In addition, Figure 80A and Figure 80BThe intra block copy shown is one mode belonging to inter prediction, and is a mode in which a block included in the current picture is referred to as a reference image or a reference block. That is, in the intra block copy, no reference is made to a picture different from the current picture. In addition, Figure 80A The PCM mode shown is one mode belonging to intra prediction, and is a mode in which no transform and quantization are performed.

[0756] [Intra prediction section]

[0757] The intra prediction section 216 performs intra prediction with reference to the block within the current picture saved in the block memory 210 based on the intra prediction mode read out from the stream, thereby generating a prediction image (i.e., an intra prediction image) of the current block. Specifically, the intra prediction section 216 performs intra prediction by referring to the pixel values (e.g., luminance values, color difference values) of the blocks adjacent to the current block, thereby generating an intra prediction image, which is output to the prediction control section 220.

[0758] In addition, in a case where the intra prediction mode of referring to the luminance block is selected in the intra prediction of the color difference block, the intra prediction section 216 can also predict the color difference components of the current block based on the luminance component of the current block.

[0759] Further, in a case where the information read out from the stream indicates the application of PDPC, the intra prediction section 216 corrects the pixel values after the intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0760] Figure 81 is a diagram showing an example of the processing performed by the intra prediction section 216 of the decoding apparatus 200.

[0761] The intra prediction section 216 first determines whether or not the MPM flag indicating 1 is present in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 is present (Yes in step Sw_11), the intra prediction section 216 acquires the information indicating the intra prediction mode selected in the encoding apparatus 100 among the MPMs from the entropy decoding section 202 (step Sw_12). Further, this information is decoded by the entropy decoding section 202 and output to the intra prediction section 216. Next, the intra prediction section 216 decides the MPM (step Sw_13). The MPM is constituted by, for example, six intra prediction modes. Then, the intra prediction section 216 decides the intra prediction mode indicated by the information acquired in step Sw_12 from among the plurality of intra prediction modes included in the MPM (step Sw_14).

[0762] On the other hand, when it is determined in step Sw_11 that the MPM flag indicating 1 is not present in the stream (NO in step Sw_11), the intra prediction section 216 acquires information indicating the intra prediction mode selected in the encoding apparatus 100 (step Sw_15). That is, the intra prediction section 216 acquires information indicating the intra prediction mode selected in the encoding apparatus 100 from among one or more intra prediction modes not included in the MPM from the entropy decoding section 202. Further, this information is decoded by the entropy decoding section 202 and output to the intra prediction section 216. Then, the intra prediction section 216 decides the intra prediction mode indicated by the information acquired in step Sw_15 from among one or more intra prediction modes not included in the MPM (step Sw_17).

[0763] The intra prediction section 216 generates a prediction image in accordance with the intra prediction mode decided in step Sw_14 or step Sw_17 (step Sw_18).

[0764] [Inter prediction section]

[0765] The inter prediction section 218 predicts a current block with reference to a reference picture stored in the frame memory 214. Prediction is performed in units of a current block or a sub-block within the current block. In addition, a sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or a size other than these. The size of the sub-block can also be switched in units of a slice, a tile, or a picture.

[0766] For example, the inter prediction section 218 performs motion compensation using motion information (e.g., an MV) read from the stream (e.g., a prediction parameter output from the entropy decoding section 202), thereby generating an inter prediction image of the current block or the sub-block, and outputs the inter prediction image to the prediction control section 220.

[0767] In a case where the information read from the stream indicates that the OBMC mode is applied, the inter prediction section 218 generates an inter prediction image using not only the motion information of the current block obtained by the motion search but also the motion information of a neighboring block.

[0768] Further, in a case where the information read from the stream indicates that the FRUC mode is applied, the inter prediction section 218 performs a motion search in accordance with a pattern matching method (bi-directional matching or template matching) read from the stream, thereby deriving motion information. Also, the inter prediction section 218 performs motion compensation (prediction) using the derived motion information.

[0769] Further, the inter prediction section 218 derives the MV based on a model assuming constant velocity straight line motion in a case where the BIO mode is applied. Further, in a case where information read out from the stream indicates that the affine mode is applied, the inter prediction section 218 derives the MV in a sub-block unit based on MVs of a plurality of neighboring blocks.

[0770] [Flow of MV derivation]

[0771] Figure 82 is a flowchart showing an example of MV derivation in the decoding apparatus 200.

[0772] The inter prediction section 218 determines, for example, whether or not to decode motion information (e.g., MV). For example, the inter prediction section 218 can determine based on a prediction mode included in the stream, or can determine based on other information included in the stream. Here, when it is determined to decode the motion information, the inter prediction section 218 derives the MV of the current block in a mode in which the motion information is decoded. On the other hand, when it is determined not to decode the motion information, the inter prediction section 218 derives the MV in a mode in which the motion information is not decoded.

[0773] Here, the modes of MV derivation are a normal inter mode, a normal merge mode, an FRUC mode, an affine mode, and the like, which will be described later. Among these modes, the modes in which the motion information is decoded are the normal inter mode, the normal merge mode, and the affine mode (specifically, the affine inter mode and the affine merge mode), and the like. Further, the motion information can include not only the MV but also prediction MV selection information, which will be described later. Further, the modes in which the motion information is not decoded are the FRUC mode, and the like. The inter prediction section 218 selects a mode for deriving the MV of the current block from among these plurality of modes, and derives the MV of the current block using the selected mode.

[0774] Figure 83 is a flowchart showing another example of MV derivation in the decoding apparatus 200.

[0775] The inter prediction section 218 determines, for example, whether or not to decode the differential MV, for example, the inter prediction section 218 can determine based on a prediction mode included in the stream, or can determine based on other information included in the stream. Here, when it is determined to decode the differential MV, the inter prediction section 218 can derive the MV of the current block in a mode in which the differential MV is decoded. In this case, the differential MV included in the stream is decoded as a prediction parameter, for example.

[0776] On the other hand, when it is determined not to decode the differential MV, the inter prediction section 218 derives the MV in a mode in which the differential MV is not decoded. In this case, the differential MV after encoding is not included in the stream.

[0777] Here, as described above, the derivation mode of the MV has a normal inter mode, a normal merge mode, an FRUC mode, an affine mode, and the like, which will be described later. Among these modes, the mode in which the differential MV is encoded has the normal inter mode, the affine mode (specifically, affine inter mode), and the like. Further, the mode in which the differential MV is not encoded has the FRUC mode, the normal merge mode, the affine mode (specifically, affine merge mode), and the like. The inter prediction section 218 selects the mode for deriving the MV of the current block from among these multiple modes, and derives the MV of the current block using the selected mode.

[0778] [MV derivation > normal inter mode]

[0779] For example, in a case where the information read out from the stream indicates that the normal inter mode is applied, the inter prediction section 218 derives the MV in the normal merge mode based on the information read out from the stream, and performs motion compensation (prediction) using the MV.

[0780] Figure 84 is a flowchart indicating an example of inter prediction by the normal inter mode in the decoding apparatus 200.

[0781] The inter prediction section 218 of the decoding apparatus 200 performs motion compensation for each block. At this time, the inter prediction section 218 first acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded blocks located in the surroundings of the current block in time or in space (step Sg_11). That is, the inter prediction section 218 creates a list of candidate MVs.

[0782] Next, the inter prediction section 218 extracts N (N is an integer of 2 or more) candidate MVs as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined priority order from among the plurality of candidate MVs acquired in step Sg_11 (step Sg_12). Note that the priority order can be predetermined for each of the N prediction MV candidates.

[0783] Next, the inter prediction section 218 decodes prediction MV selection information from the input stream, and selects one prediction MV candidate as a prediction MV for the current block from among the N prediction MV candidates using the decoded prediction MV selection information (step Sg_13).

[0784] Next, the inter prediction section 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the differential value of the decoded differential MV to the selected prediction MV (step Sg_14).

[0785] Finally, the inter prediction section 218 generates a prediction image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are performed for each block. For example, when the processes of steps Sg_11 to Sg_15 are performed for all blocks included in a slice respectively, the inter prediction using the normal inter mode for the slice ends. Also, when the processes of steps Sg_11 to Sg_15 are performed for all blocks included in a picture respectively, the inter prediction using the normal inter mode for the picture ends. Further, the processes of steps Sg_11 to Sg_15 can also be such that, when the processes are performed for a part of the blocks included in a slice, the inter prediction using the normal inter mode for the slice ends. Similarly, the processes of steps Sg_11 to Sg_15 can also be such that, when the processes are performed for a part of the blocks included in a picture, the inter prediction using the normal inter mode for the picture ends.

[0786] [MV derivation] normal merge mode

[0787] For example, in a case where the information read out from the stream indicates the application of the normal merge mode, the inter prediction section 218 derives an MV in the normal merge mode, and performs motion compensation (prediction) using the MV.

[0788] Figure 85 is a flowchart indicating an example of inter prediction based on the normal merge mode in the decoding apparatus 200.

[0789] The inter prediction section 218 first acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded blocks located in the surroundings of the current block in time or in space (step Sh_11). That is, the inter prediction section 218 creates a list of candidate MVs.

[0790] Next, the inter prediction section 218 derives an MV of the current block by selecting one candidate MV from among the plurality of candidate MVs acquired in step Sh_11 (step Sh_12). Specifically, the inter prediction section 218, for example, acquires MV selection information included as a prediction parameter in the stream, and selects a candidate MV identified by the MV selection information as the MV of the current block.

[0791] Finally, the inter prediction section 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processes of steps Sh_11 to Sh_13 are performed, for example, for each block. For example, when the processes of steps Sh_11 to Sh_13 are performed for all blocks included in a slice respectively, the inter prediction using the normal merge mode for the slice ends. Also, when the processes of steps Sh_11 to Sh_13 are performed for all blocks included in a picture respectively, the inter prediction using the normal merge mode for the picture ends. Further, the processes of steps Sh_11 to Sh_13 can also be that, when the processes are performed for a part of the blocks included in a slice, the inter prediction using the normal merge mode for the slice ends. The processes of steps Sh_11 to Sh_13 can also be that, likewise, when the processes are performed for a part of the blocks included in a picture, the inter prediction using the normal merge mode for the picture ends.

[0792] [MV derivation > FRUC mode]

[0793] For example, in a case where the information read out from the stream indicates the application of the FRUC mode, the inter prediction section 218 derives an MV in the FRUC mode, and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding apparatus 100 side, but is derived at the decoding apparatus 200 side. For example, the decoding apparatus 200 can also derive the motion information by performing a motion search. In this case, the decoding apparatus 200 does not use the pixel values of the current block to perform the motion search.

[0794] Figure 86 is a flowchart indicating an example of the inter prediction based on the FRUC mode in the decoding apparatus 200.

[0795] First, the inter prediction section 218 refers to MVs of each of the decoded blocks that are spatially or temporally adjacent to the current block, generates a list in which these MVs are expressed as candidate MVs (i.e., is a candidate MV list, and can also be common to the candidate MV list of the normal merge mode) (step Si_11). Next, the inter prediction section 218 selects the best candidate MV from among the plurality of candidate MVs registered in the candidate MV list (step Si_12). For example, the inter prediction section 218 calculates evaluation values of each of the candidate MVs included in the candidate MV list, and selects one candidate MV as the best candidate MV on the basis of the evaluation values. Then, the inter prediction section 218 derives the MV for the current block on the basis of the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block can be derived by performing pattern matching in a peripheral region of the position within the reference picture corresponding to the selected best candidate MV. That is, for the region around the best candidate MV, search using pattern matching and evaluation values in the reference picture is performed, and further, in a case where there is an MV for which the evaluation value becomes a good value, the best candidate MV can be updated to this MV, and this can be used as the final MV for the current block. The update to the MV having a better evaluation value can not be performed.

[0796] Finally, the inter prediction section 218 generates a prediction image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processes of steps Si_11 to Si_15 are performed, for example, for each block. For example, when the processes of steps Si_11 to Si_15 are performed for all the blocks included in a slice respectively, the inter prediction using the FRUC mode for the slice ends. Alternatively, when the processes of steps Si_11 to Si_15 are performed for all the blocks included in a picture respectively, the inter prediction using the FRUC mode for the picture ends. The processes can be performed in the sub-block unit as well as the above-described block unit.

[0797] [MV derivation > affine merge mode]

[0798] For example, in a case where the information read out from the stream indicates application of the affine merge mode, the inter prediction section 218 derives the MV in the affine merge mode, and performs motion compensation (prediction) using the MV.

[0799] Figure 87 is a flowchart indicating an example of inter prediction based on the affine merge mode in the decoding apparatus 200.

[0800] In the affine merge mode, the inter prediction section 218 first derives the MVs of the control points of the current block respectively (step Sk_11). As described above, the MVs of the control points are derived on the basis of the MVs of the decoded blocks that are spatially or temporally adjacent to the current block. Figure 46AAs shown, the control points are the top-left and top-right corners of the current block, or as... Figure 46B The image shows the top left, top right, and bottom left corners of the current block.

[0801] For example, in use Figures 47A-47C In the case of the MV export method shown, such as Figure 47A As shown, the inter-frame prediction unit 218 checks these blocks in the order of decoded block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left) to determine the first valid block decoded in affine mode.

[0802] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to derive the MV of the control points. For example, if block A is determined and block A has 2 control points, then... Figure 47B As shown, the inter-frame prediction unit 218 calculates the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block by projecting the motion vectors v3 and v4 of the upper left and upper right corners of the decoded block containing block A onto the current block. Thus, the MV of each control point is derived.

[0803] In addition, such as Figure 49A As shown, given that block A has 2 control points, the MV of 3 control points can also be calculated, or as follows: Figure 49B Define block A as shown. If block A has 3 control points, calculate the MV of 2 control points.

[0804] In addition, when the stream contains MV selection information as a prediction parameter, the inter-frame prediction unit 218 can also use the MV selection information to derive the MV of each control point of the current block.

[0805] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks contained in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 218 uses two motion vectors v0 and v1 and the above-mentioned equation (1A), or uses three motion vectors v0, v1, and v2 and the above-mentioned equation (1B), to calculate the MV of the sub-block as an affine MV (step Sk_12). Then, the inter-frame prediction unit 218 uses these affine MVs and the decoded reference image to perform motion compensation on the sub-block (step Sk_13). When the processing of steps Sk_12 and Sk_13 is performed on all the sub-blocks contained in the current block, the inter-frame prediction using the affine merging mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0806] Further, in step Sk_11, the above-described candidate MV list can also be generated. The candidate MV list can be, for example, a list including candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be Figures 47A-47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in, and an arbitrary combination of other MV derivation methods.

[0807] Further, the candidate MV list can include candidate MVs of a mode other than the affine mode that performs prediction in a sub-block unit.

[0808] Further, as the candidate MV list, for example, a candidate MV list including candidate MVs of the affine merge mode having 2 control points and candidate MVs of the affine merge mode having 3 control points can be generated. Alternatively, a candidate MV list including candidate MVs of the affine merge mode having 2 control points and a candidate MV list including candidate MVs of the affine merge mode having 3 control points can be generated, respectively. Alternatively, a candidate MV list including candidate MVs of one of the affine merge mode having 2 control points and the affine merge mode having 3 control points can be generated.

[0809] [MV derivation > affine inter mode]

[0810] For example, in a case where information read out from the stream indicates application of the affine inter mode, the inter prediction section 218 derives an MV in the affine inter mode, and performs motion compensation (prediction) using the MV.

[0811] Figure 88 is a flowchart indicating an example of inter prediction based on the affine inter mode in the decoding apparatus 200.

[0812] In the affine inter mode, first, the inter prediction section 218 derives a prediction MV (v0, v1) or (v0, v1, v2) of each of 2 or 3 control points of the current block (step Sj_11). The control points are, for example, points of the upper left corner, the upper right corner, or the lower left corner of the current block as shown in Figure 46A or Figure 46B

[0813] The inter prediction section 218 acquires prediction MV selection information included as a prediction parameter in the stream, and derives a prediction MV of each control point of the current block using an MV identified by the prediction MV selection information. For example, in a case where the MV derivation method shown in Figure 48A and Figure 48B Figure 48A ​​or Figure 48B The MV of the block in the decoded block near the control point of the current block indicated by the prediction MV selection information is added to the prediction MV (v0, v1) or (v0, v1, v2) of the control point of the current block.

[0814] Next, the inter prediction section 218 adds the prediction MV of each control point of the current block to the differential MV corresponding to the prediction MV (step Sj_12), for example. Thus, the MV of each control point of the current block is derived.

[0815] Next, the inter prediction section 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. That is, the inter prediction section 218 calculates the MV of each of the plurality of sub-blocks as an affine MV using the two motion vectors v0 and v1 and the above-described equation (1A), or using the three motion vectors v0, v1, and v2 and the above-described equation (1B) (step Sj_13). Then, the inter prediction section 218 performs motion compensation on the sub-block using these affine MVs and the decoded reference picture (step Sj_14). When the processes of steps Sj_13 and Sj_14 are performed on each of all the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block ends. That is, the current block is motion-compensated, and a prediction image of the current block is generated.

[0816] Further, in step Sj_11, the above-described candidate MV list can also be generated similarly to step Sk_11.

[0817] [MV derivation > triangle mode]

[0818] For example, in a case where the information read out from the stream indicates the application of the triangle mode, the inter prediction section 218 derives the MV in the triangle mode, and performs motion compensation (prediction) using the MV.

[0819] Figure 89 is a flowchart indicating an example of inter prediction based on the triangle mode in the decoding apparatus 200.

[0820] In the triangle mode, first, the inter prediction section 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter prediction section 218 can acquire partition information as prediction parameters from the stream in association with the division into the partitions. Further, the inter prediction section 218 can divide the current block into the first partition and the second partition in accordance with the partition information.

[0821] Next, the inter prediction section 218 first acquires a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded blocks located around the current block in time or in space (step Sx_12). That is, the inter prediction section 218 creates a list of candidate MVs.

[0822] Then, the inter prediction section 218 selects a candidate MV of the first partition and a candidate MV of the second partition as the first MV and the second MV, respectively, from the plurality of candidate MVs acquired in step Sx_11 (step Sx_13). At this time, the inter prediction section 218 can also acquire, from the stream, MV selection information for identifying the selected candidate MVs as the prediction parameters. Then, the inter prediction section 218 can select the first MV and the second MV in accordance with the MV selection information.

[0823] Next, the inter prediction section 218 generates the first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Likewise, the inter prediction section 218 generates the second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).

[0824] Finally, the inter prediction section 218 generates the prediction image of the current block by weightedly adding the first prediction image and the second prediction image (step Sx_16).

[0825] [Motion search > DMVR]

[0826] For example, in a case where the information read out from the stream indicates the application of DMVR, the inter prediction section 218 performs motion search by DMVR.

[0827] Figure 90 is a flowchart showing an example of motion search based on DMVR in the decoding apparatus 200.

[0828] The inter prediction section 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter prediction section 218 derives the final MV for the current block by searching a surrounding area of a reference picture indicated by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is decided by DMVR.

[0829] Figure 91 is a flowchart showing an example of motion search based on DMVR in the decoding apparatus 200.

[0830] First, the inter prediction section 218 derives the MV of the current block in the merge mode (step Sl_11). Next, the inter prediction section 218 derives the final MV for the current block by searching a surrounding area of a reference picture indicated by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is decided by DMVR. Figure 58AIn Step 1, the inter prediction section 218 calculates the cost of the search position indicated by the initial MV (also referred to as the start point) and eight search positions located around it. Then, the inter prediction section 218 determines whether the cost of the search position other than the start point is the minimum. In this case, when it is determined that the cost of the search position other than the start point is the minimum, the inter prediction section 218 moves to the search position of which the cost is the minimum, and performs the processing of Step 2. Figure 58A On the other hand, if the cost of the start point is the minimum, the inter prediction section 218 skips the processing of Step 2 and performs the processing of Step 3. Figure 58A

[0831] In Step 2, the inter prediction section 218 performs the same search as the processing of Step 1, using the search position moved according to the result of the processing of Step 1 as a new start point. Then, the inter prediction section 218 determines whether the cost of the search position other than the start point is the minimum. In this case, when the cost of the search position other than the start point is the minimum, the inter prediction section 218 performs the processing of Step 4. On the other hand, when the cost of the start point is the minimum, the inter prediction section 218 performs the processing of Step 3. Figure 58A

[0832] In Step 4, the inter prediction section 218 processes the search position of the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0833] In Step 3, the inter prediction section 218 determines the decimal-precision pixel position of which the cost is the minimum, based on the costs of the four points above, below, left, and right of the start point in Step 1 or Step 2, and processes the pixel position as the final search position. The decimal-precision pixel position is determined by performing weighted addition of the vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) located above, below, left, and right of the four points, using the costs of the search positions of the four points as the weights. Then, the inter prediction section 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector. Figure 58A [Motion compensation > BIO / OBMC / LIC]

[0834] For example, in the case where the information read out from the stream indicates the application of a modification to the prediction image, the inter prediction section 218 modifies the prediction image in accordance with the mode of the modification when generating the prediction image. The mode is, for example, BIO, OBMC, and LIC described above.

[0835]

[0836] Figure 92 is a flowchart indicating an example of the generation of a prediction image in the decoding apparatus 200. ​​​

[0837] The inter prediction section 218 generates a prediction image (step Sm_11), and corrects the prediction image by any of the above-described modes (step Sm_12).

[0838] Figure 93 is a flowchart showing another example of generation of a prediction image in the decoding apparatus 200.

[0839] The inter prediction section 218 derives the MV of the current block (step Sn_11). Next, the inter prediction section 218 generates a prediction image using the MV (step Sn_12), and determines whether or not to perform the correction process (step Sn_13). For example, the inter prediction section 218 acquires a prediction parameter included in the stream, and determines whether or not to perform the correction process based on the prediction parameter. The prediction parameter is, for example, a flag indicating whether or not to apply each of the above-described modes. Here, when it is determined to perform the correction process (Yes in step Sn_13), the inter prediction section 218 generates a final prediction image by correcting the prediction image (step Sn_14). Further, in LIC, the luminance and color difference of the prediction image can be corrected in step Sn_14. On the other hand, when it is determined not to perform the correction process (No in step Sn_13), the inter prediction section 218 outputs the prediction image as it is as a final prediction image without correcting it (step Sn_15).

[0840] [OBMC]

[0841] For example, in a case where the information read out from the stream indicates the application of OBMC, the inter prediction section 218 corrects the prediction image in accordance with OBMC when generating the prediction image.

[0842] Figure 94 is a flowchart showing an example of correction of a prediction image based on OBMC in the decoding apparatus 200. In addition, Figure 94 The flowchart of Figure 62 shows a flow of correction of a prediction image using the current picture and the reference picture as shown in

[0843] First, as shown in Figure 62 , the inter prediction section 218 acquires a prediction image (Pred) based on normal motion compensation using the MV allocated to the current block.

[0844] Next, the inter prediction section 218 applies (reuses) the MV (MV_L) already derived with respect to the decoded left neighboring block to the current block, and acquires a prediction image (Pred_L). Then, the inter prediction section 218 performs the 1st correction of the prediction image by superimposing the two prediction images Pred and Pred_L. This has an effect of mixing the boundaries between the neighboring blocks.

[0845] Similarly, the inter prediction section 218 applies (reuses) the MV (MV_U) that has been derived for the decoded upper neighboring block to the current block, to obtain a prediction image (Pred_U). Then, the inter prediction section 218 performs a second correction of the prediction image by superimposing the prediction image Pred_U and the prediction image that has been subjected to the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between the neighboring blocks. The prediction image obtained by the second correction is the final prediction image of the current block that is blended (smoothed) with the boundaries of the neighboring blocks.

[0846] [Motion compensation > BIO]

[0847] For example, in a case where the information read out from the stream indicates the application of BIO, the inter prediction section 218 corrects the prediction image in accordance with BIO when generating the prediction image.

[0848] Figure 95 is a flowchart that shows an example of the correction of the prediction image based on BIO in the decoding apparatus 200.

[0849] As shown in Figure 63 , the inter prediction section 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (CurPic) that contains the current block. Then, the inter prediction section 218 derives the prediction image of the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0850] Next, the inter prediction section 218 derives an interpolated image I 0 for the current block using the motion vector M0 and the reference picture L0 (step Sy_12). Here, the interpolated image I 1 is an image contained in the reference picture Ref0 that is derived for the current block, and the interpolated image I 0 is an image contained in the reference picture Ref1 that is derived for the current block. The interpolated image I 1 and the interpolated image I 0 may each be an image of the same size as the current block. Alternatively, in order to appropriately derive a gradient image to be described later, the interpolated image I 1 and the interpolated image I 0 may each be an image larger than the current block. Furthermore, the interpolated image I 1 and I 0 may each be an image larger than the current block. Furthermore, the interpolated image I 1A prediction picture can be derived including the application of motion vectors (M0, M1) and reference pictures (L0, L1), and a motion compensated filter.

[0851] In addition, the inter prediction section 218 derives gradient pictures (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) of the current block from the interpolated pictures I 0 and I 1 (step Sy_13). Further, the gradient pictures in the horizontal direction are (Ix 0 , Ix 1 ), and the gradient pictures in the vertical direction are (Ix 0 , Ix 1 ). The inter prediction section 218 can also derive the gradient pictures by applying a gradient filter to the interpolated pictures, for example. The gradient pictures are images that indicate the amount of spatial change in pixel values in the horizontal direction or the vertical direction.

[0852] Next, the inter prediction section 218 derives the optical flow (vx, vy) as the above-mentioned velocity vector using the interpolated pictures (I 0 , I 1 ) and the gradient pictures (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) in units of the plurality of sub-blocks that constitute the current block (step Sy_14). As an example, the sub-blocks can be sub-CUs of 4x4 pixels.

[0853] Next, the inter prediction section 218 corrects the prediction picture of the current block using the optical flow (vx, vy). For example, the inter prediction section 218 derives a correction value of the value of a pixel included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter prediction section 218 can correct the prediction picture of the current block using the correction value (step Sy_16). In addition, the correction value can be derived in units of each pixel, or can be derived in units of a plurality of pixels or sub-blocks.

[0854] Further, the processing flow of the BIO is not limited to the disclosed processing. The processing can be performed only for a part of the disclosed processing, or different processing can be added or replaced, or the processing order can be changed. Figure 95 Figure 95 [Motion compensation > LIC]

[0855] [Motion compensation > LIC]

[0856] ​For example, in a case where the information read out from the stream indicates an application of LIC, when a prediction image is generated, the inter prediction section 218 corrects the prediction image in accordance with the LIC.

[0857] Figure 96 is a flowchart indicating an example of correction of a prediction image based on LIC in the decoding device 200.

[0858] First, the inter prediction section 218 uses the MV to acquire a reference image corresponding to the current block from a decoded reference picture (step Sz_11).

[0859] Next, the inter prediction section 218 extracts information indicating how luminance values change in the reference picture and the current picture with respect to the current block (step Sz_12). As shown in Figure 66A is performed based on luminance pixel values of a decoded left neighboring reference region (a peripheral reference region) and a decoded upper neighboring reference region (a peripheral reference region) in the current picture, and luminance pixel values of equivalent positions within the reference picture specified by the derived MV. Then, the inter prediction section 218 calculates a luminance correction parameter using the information indicating how luminance values change (step Sz_13).

[0860] The inter prediction section 218 generates a prediction image with respect to the current block by performing luminance correction processing of applying the luminance correction parameter to the reference image within the reference picture specified by the MV (step Sz_14). That is, the prediction image as the reference image within the reference picture specified by the MV is corrected based on the luminance correction parameter. In this correction, luminance can be corrected, or color difference can be corrected.

[0861] [Prediction control section]

[0862] The prediction control section 220 selects one of the intra prediction image and the inter prediction image, and outputs the selected prediction image to the addition section 208. Overall, the structure, functions, and processing of the prediction control section 220, the intra prediction section 216, and the inter prediction section 218 on the decoding device 200 side can correspond to the structure, functions, and processing of the prediction control section 128, the intra prediction section 124, and the inter prediction section 126 on the encoding device 100 side.

[0863] [1st mode regarding allocation of slice index]

[0864] For example, in VVC, there are two kinds of slice indexes, a picture-level slice index and a sub-picture-level slice index, for a slice. The picture-level slice index is an index for identifying a slice from among a plurality of slices included in a picture. The sub-picture-level slice index is an index for identifying a slice from among one or more slices included in a sub-picture.

[0865] slice_address signaled in the slice header represents the subpicture level slice index.

[0866] It can be allowed that the picture level slice index is completely different from the subpicture level slice index. In the case that there is no simple queuing rule between the 2 slice indexes, an additional permutation for mapping the subpicture level slice index to the picture level slice index can be used.

[0867] For example, such an additional permutation can be represented by SliceSubpicToPicIdx[i][k]. Here, i represents the subpicture index, and k represents the subpicture level slice index. And SliceSubpicToPicIdx[i][k] represents the picture level slice index of the slice whose subpicture level slice index is k in the subpicture whose subpicture index is i.

[0868] Figure 97 is a conceptual diagram representing multiple picture level slice indexes within a subpicture with discontinuity. In Figure 97 , x of slice #x represents the picture level slice index, and x of tile #x represents the tile index. In Figure 97 , subpicture A contains slice #0, slice #1, slice #2, slice #3, slice #4, and slice #6, and subpicture B contains slice #5 and slice #7.

[0869] Although not shown, the subpicture level slice indexes of slice #0, slice #1, slice #2, slice #3, slice #4, and slice #6 in subpicture A are 0, 1, 2, 3, 4, and 5, respectively. The subpicture level slice indexes of slice #5 and slice #7 in subpicture B are 0 and 1, respectively. The subpicture level slice indexes are signaled in slice_address within the slice header.

[0870] In Figure 97 , the picture level slice index is not aligned with the subpicture level slice index signaled in slice_address within the slice header.

[0871] In addition, all the slices within a subpicture are permuted together in the bitstream. Therefore, for example, slice #6 is configured before slice #5 in the bitstream. That is, Figure 97 the picture level slice index of does not represent the decoding order of multiple slices and the order of multiple NAL (Network Abstraction Layer) units in the bitstream.

[0872] Figure 98is a conceptual diagram indicating a plurality of picture-level slice indexes within a sub-picture without discontinuity. This modality involves that a plurality of pictur...

Claims

1. An encoding apparatus, comprising: Possessing: circuitry; and a memory connected to the circuitry, a picture contains a plurality of sub-pictures, the picture contains a plurality of slices, each of the plurality of slices is contained in one of the plurality of sub-pictures, the circuitry acts on, assigning, to each of the plurality of slices, a picture-level slice index assigned at a picture level and a sub-picture-level slice index assigned at a sub-picture level, encoding the plurality of sub-picture-level slice indexes into a plurality of slice headers respectively corresponding to the plurality of slices, encoding each of the plurality of slices into a bitstream, the picture-level slice index assigned to a target slice contained in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice and (2) a total number of slices contained in a sub-picture that is encoded before the target sub-picture, the sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is incremented by 1 from 0, the bitstream contains the plurality of slices as a plurality of NAL units, an order of the plurality of NAL units is determined based on an order of sub-picture-level slice indexes.

2. A decoding apparatus, wherein, Possessing: circuitry; and a memory connected to the circuitry, a picture contains a plurality of sub-pictures, the picture contains a plurality of slices, each of the plurality of slices is contained in one of the plurality of sub-pictures, the circuitry acts on, decoding, from a plurality of slice headers respectively corresponding to the plurality of slices, a plurality of sub-picture-level slice indexes respectively, decoding each of the plurality of slices from a bitstream, each of the plurality of slices is assigned a picture-level slice index assigned at a picture level and the sub-picture-level slice index assigned at a sub-picture level, the picture-level slice index assigned to a target slice contained in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice and (2) a total number of slices contained in a sub-picture that is encoded before the target sub-picture, the sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is incremented by 1 from 0, the bitstream contains the plurality of slices as a plurality of NAL units, an order of the plurality of NAL units is determined based on an order of sub-picture-level slice indexes.

3. A bitstream generation apparatus, wherein, Possessing: circuitry; and a memory connected to the circuitry, a picture contains a plurality of sub-pictures, the picture contains a plurality of slices, each of the plurality of slices is contained in one of the plurality of sub-pictures, the circuitry acts on, assigning, to each of the plurality of slices, a picture-level slice index assigned at a picture level and a sub-picture-level slice index assigned at a sub-picture level, including the plurality of sub-picture-level slice indexes in a plurality of slice headers respectively corresponding to the plurality of slices, generating a bitstream containing the plurality of slice headers and the plurality of slices, The picture-level slice index given to a target slice included in a target sub-picture is calculated by adding (1) a value of a sub-picture-level slice index of the target slice and (2) a total number of slices included in a sub-picture that is encoded before the target sub-picture, The sub-picture-level slice index is slice_address stored in a slice header and is an integer value that is increased by 1 from 0, The bitstream includes the plurality of slices as a plurality of NAL units, An order of the plurality of NAL units is determined based on an order of the sub-picture-level slice index.