Encoding device, decoding device, encoding method, and decoding method

By designing an encoding device that can encode and decode the performance requirements information of the decoding device in the video encoding technology, the difficulty of selecting video encoding elements in the prior art is solved, and the encoding efficiency and picture quality are improved, as well as the reduction of processing volume and circuit scale are achieved.

CN114902679BActive Publication Date: 2025-06-24PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080085136.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-13
Filing Date
2020-12-10
Publication Date
2025-06-24
Estimated Expiration
2040-12-10

AI Technical Summary

Technical Problem

While the existing video encoding technology improves encoding efficiency, picture quality and reduces processing volume and circuit scale, there is difficulty in appropriately selecting elements such as filters, blocks, sizes, motion vectors, reference pictures or reference blocks.

Method used

By designing an encoding device, the device has a circuit and a memory, it is able to encode a set of layers including at least one output layer, generate a bit stream, and encoded data of the image of the output layer shared by the set of layers. The device can encode performance requirement information of the decoding device to the header shared by the layer set when the number of layers is 1 in the layer set.

Benefits of technology

Improved coding efficiency, improved picture quality, reduced processing volume and circuit scale, improved processing speed and appropriate selection of elements or actions are achieved, and the system for processing multi-layer bitstreams is simplified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902679B_ABST
    Figure CN114902679B_ABST
Patent Text Reader

Abstract

The encoding device (100) includes a circuit (a1) and a memory (a2) connected to the circuit (a1). During operation, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, the circuit (a1) encodes performance requirement information representing the performance requirements of the decoding device (200) into a header shared by the layers included in the layer set and generates a bitstream. The bitstream includes the header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method. Background Art

[0002] Video encoding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to process the continuously increasing amount of digital video data in various applications, there has always been a need to provide improvements and optimizations to video encoding technology. The present disclosure relates to further progress, improvements, and optimizations in video encoding.

[0003] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above-described video encoding technology.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] Regarding the encoding method as described above, for the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, or the appropriate selection of elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, it is desired to propose a new method.

[0009] The present disclosure provides a structure or method that can contribute to one or more of, for example, the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, the improvement of processing speed, and the appropriate selection of elements or operations. In addition, the present disclosure may include a structure or method that can contribute to benefits other than the above.

[0010] Means for Solving the Problems

[0011] For example, an encoding device according to an aspect of the present disclosure includes a circuit and a memory connected to the circuit. During operation, the circuit encodes performance requirement information representing the performance requirements of a decoding device into a header shared by the layers included in a layer set, even if the number of layers included in the layer set is 1, for a layer set including at least one output layer, and generates a bitstream. The bitstream includes the header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

[0012] In video encoding technology, in order to improve encoding efficiency, improve picture quality, reduce circuit scale, etc., it is desirable to propose new methods.

[0013] Each embodiment or a part of the structure or method in the present disclosure can respectively achieve at least any one of, for example, improvement of encoding efficiency, improvement of picture quality, reduction of encoding / decoding processing amount, reduction of circuit scale, or improvement of encoding / decoding processing speed. Or, each embodiment or a part of the structure or method in the present disclosure can respectively make an appropriate selection of elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, etc. during encoding and decoding. In addition, the present disclosure also includes disclosure of structures or methods that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in processing amount.

[0014] Based on the specification and the drawings, further advantages and effects in an aspect of the present disclosure are clarified. These advantages and / or effects are respectively obtained by several embodiments and the features described in the specification and the drawings, but it is not necessary to provide all of them in order to obtain one or more advantages and / or effects.

[0015] In addition, these general or specific aspects can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0016] Advantages of the Invention

[0017] The structure or method according to an aspect of the present disclosure can contribute to, for example, one or more of improvement of encoding efficiency, improvement of picture quality, reduction of processing amount, reduction of circuit scale, improvement of processing speed, and appropriate selection of elements or actions. In addition, the structure or method according to an aspect of the present disclosure can also contribute to benefits other than the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram showing an example of the structure of a transmission system according to an embodiment.

[0019] Figure 2 This is a diagram showing an example of the hierarchical structure of data in a stream.

[0020] Figure 3 This is a diagram showing an example of the structure of a slice.

[0021] Figure 4 This is a diagram showing an example of the structure of a tile.

[0022] Figure 5 This is a diagram showing an example of the coding structure in scalable coding.

[0023] Figure 6 This is a diagram showing an example of the coding structure in scalable coding.

[0024] Figure 7 This is a block diagram showing an example of the functional structure of an encoding device in an embodiment.

[0025] Figure 8 This is a block diagram showing an example of the installation of an encoding device.

[0026] Figure 9 This is a flowchart showing an example of the overall encoding process performed by an encoding device.

[0027] Figure 10 This is a diagram showing an example of block segmentation.

[0028] Figure 11 This is a diagram showing an example of the functional structure of a segmentation unit.

[0029] Figure 12 This is a diagram showing an example of a segmentation pattern.

[0030] Figure 13A This is a diagram showing an example of the syntax tree of a segmentation pattern.

[0031] Figure 13B This is a diagram showing another example of the syntax tree of a segmentation pattern.

[0032] Figure 14 This is a table showing transform basis functions corresponding to each transform type.

[0033] Figure 15 This is a diagram showing an example of SVT.

[0034] Figure 16 This is a flowchart showing an example of the process performed by a transform unit.

[0035] Figure 17 This is another flowchart showing an example of the process performed by a transform unit.

[0036] Figure 18It is a block diagram showing an example of the functional structure of a quantization unit.

[0037] Figure 19 It is a flowchart showing an example of quantization performed by a quantization unit.

[0038] Figure 20 It is a block diagram showing an example of the functional structure of an entropy encoding unit.

[0039] Figure 21 It is a diagram showing the process of CABAC in an entropy encoding unit.

[0040] Figure 22 It is a block diagram showing an example of the functional structure of a loop filtering unit.

[0041] Figure 23A It is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).

[0042] Figure 23B It is a diagram showing another example of the shape of a filter used in an ALF.

[0043] Figure 23C It is a diagram showing another example of the shape of a filter used in an ALF.

[0044] Figure 23D It is a diagram showing an example where a Y sample (first component) is used for CCALF of Cb and Cr (multiple components different from the first component).

[0045] Figure 23E It is a diagram showing a diamond-shaped filter.

[0046] Figure 23F It is a diagram showing an example of JC-CCALF.

[0047] Figure 23G It is a diagram showing an example of the weight_index candidate of JC-CCALF.

[0048] Figure 24 It is a block diagram showing an example of the detailed structure of a loop filtering unit that functions as a DBF.

[0049] Figure 25 It is a diagram showing an example of deblocking filtering with filtering characteristics symmetric with respect to a block boundary.

[0050] Figure 26 It is a diagram for explaining an example of a block boundary where deblocking filtering processing is performed.

[0051] Figure 27 It is a diagram showing an example of a Bs value.

[0052] Figure 28 It is a flowchart showing an example of the processing performed by the prediction unit of the encoding device.

[0053] Figure 29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.

[0054] Figure 30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.

[0055] Figure 31 It is a diagram showing an example of 67 intra prediction modes in intra prediction.

[0056] Figure 32 It is a flowchart showing an example of the processing performed by the intra prediction unit.

[0057] Figure 33 It is a diagram showing an example of each reference picture.

[0058] Figure 34 It is a conceptual diagram showing an example of a reference picture list.

[0059] Figure 35 It is a flowchart showing the flow of the basic processing of inter prediction.

[0060] Figure 36 It is a flowchart showing an example of MV derivation.

[0061] Figure 37 It is a flowchart showing another example of MV derivation.

[0062] Figure 38A It is a diagram showing an example of the classification of each mode of MV derivation.

[0063] Figure 38B It is a diagram showing an example of the classification of each mode of MV derivation.

[0064] Figure 39 It is a flowchart showing an example of inter prediction based on the normal inter mode.

[0065] Figure 40 It is a flowchart showing an example of inter prediction based on the normal merge mode.

[0066] Figure 41 It is a diagram for explaining an example of the MV derivation process based on the normal merge mode.

[0067] Figure 42 It is a diagram for explaining an example of the MV derivation process based on the HMVP mode.

[0068] Figure 43It is a flowchart showing an example of FRUC (frame rate up conversion).

[0069] Figure 44 It is a diagram showing an example for explaining pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0070] Figure 45 It is a diagram showing an example for explaining pattern matching (template matching) between a template in the current picture and a block in a reference picture.

[0071] Figure 46A It is a diagram showing an example for explaining the derivation of MV in units of sub - blocks in an affine mode using two control points.

[0072] Figure 46B It is a diagram showing an example for explaining the derivation of MV in units of sub - blocks in an affine mode using three control points.

[0073] Figure 47A It is a conceptual diagram showing an example for explaining the derivation of MV of control points in an affine mode.

[0074] Figure 47B It is a conceptual diagram showing an example for explaining the derivation of MV of control points in an affine mode.

[0075] Figure 47C It is a conceptual diagram showing an example for explaining the derivation of MV of control points in an affine mode.

[0076] Figure 48A It is a diagram showing an affine mode with two control points.

[0077] Figure 48B It is a diagram showing an affine mode with three control points.

[0078] Figure 49A It is a conceptual diagram showing an example for explaining the method of deriving MV of control points when the number of control points in an encoded block and the current block is different.

[0079] Figure 49B It is a conceptual diagram showing another example for explaining the method of deriving MV of control points when the number of control points in an encoded block and the current block is different.

[0080] Figure 50 It is a flowchart showing an example of the process of an affine merge mode.

[0081] Figure 51 It is a flowchart showing an example of the process of an affine inter - frame mode.

[0082] Figure 52AIt is a diagram for explaining the generation of predicted images of two triangles.

[0083] Figure 52B It is a conceptual diagram showing an example of the first part of the first partition, the first sample set, and the second sample set.

[0084] Figure 52C It is a conceptual diagram showing the first part of the first partition.

[0085] Figure 53 It is a flowchart showing an example of a triangular pattern.

[0086] Figure 54 It is a diagram showing an example of the ATMVP mode for deriving MV in sub-block units.

[0087] Figure 55 It is a diagram showing the relationship between the merge mode and DMVR (dynamic motion vector refreshing).

[0088] Figure 56 It is a conceptual diagram for explaining an example of DMVR.

[0089] Figure 57 It is a conceptual diagram for explaining another example of DMVR for determining MV.

[0090] Figure 58A It is a diagram showing an example of motion search in DMVR.

[0091] Figure 58B It is a flowchart showing an example of motion search in DMVR.

[0092] Figure 59 It is a flowchart showing an example of the generation of predicted images.

[0093] Figure 60 It is a flowchart showing another example of the generation of predicted images.

[0094] Figure 61 It is a flowchart for explaining an example of the predicted image correction process based on OBMC (overlapped block motion compensation).

[0095] Figure 62 It is a conceptual diagram for explaining an example of the predicted image correction process based on OBMC.

[0096] Figure 63 It is a diagram for explaining a model assuming uniform linear motion.

[0097] Figure 64It is a flowchart showing an example of inter-frame prediction according to BIO.

[0098] Figure 65 It is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.

[0099] Figure 66A It is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC (local illumination compensation).

[0100] Figure 66B It is a flowchart showing an example of a method for generating a predicted image using a luminance correction process based on LIC.

[0101] Figure 67 It is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0102] Figure 68 It is a block diagram showing an example of the installation of a decoding device.

[0103] Figure 69 It is a flowchart showing an example of the overall decoding process performed by a decoding device.

[0104] Figure 70 It is a diagram showing the relationship between a segmentation determination unit and other components.

[0105] Figure 71 It is a block diagram showing an example of the functional structure of an entropy decoding unit.

[0106] Figure 72 It is a diagram showing the process of CABAC in an entropy decoding unit.

[0107] Figure 73 It is a block diagram showing an example of the functional structure of an inverse quantization unit.

[0108] Figure 74 It is a flowchart showing an example of inverse quantization performed by an inverse quantization unit.

[0109] Figure 75 It is a flowchart showing an example of the process performed by an inverse transform unit.

[0110] Figure 76 It is a flowchart showing another example of the process performed by an inverse transform unit.

[0111] Figure 77 It is a block diagram showing an example of the functional structure of a loop filter unit.

[0112] Figure 78 It is a flowchart showing an example of the process performed by the prediction unit of a decoding device.

[0113] Figure 79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device.

[0114] Figure 80A It is a flowchart showing a part of another example of the processing performed by the prediction unit of the decoding device.

[0115] Figure 80B It is a flowchart showing the remaining part of another example of the processing performed by the prediction unit of the decoding device.

[0116] Figure 81 It is a diagram showing an example of the processing performed by the intra prediction unit of the decoding device.

[0117] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device.

[0118] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device.

[0119] Figure 84 It is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding device.

[0120] Figure 85 It is a flowchart showing an example of inter prediction based on the normal merge mode in the decoding device.

[0121] Figure 86 It is a flowchart showing an example of inter prediction based on the FRUC mode in the decoding device.

[0122] Figure 87 It is a flowchart showing an example of inter prediction based on the affine merge mode in the decoding device.

[0123] Figure 88 It is a flowchart showing an example of inter prediction based on the affine inter mode in the decoding device.

[0124] Figure 89 It is a flowchart showing an example of inter prediction based on the triangular mode in the decoding device.

[0125] Figure 90 It is a flowchart showing an example of motion search based on DMVR in the decoding device.

[0126] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device.

[0127] Figure 92 It is a flowchart showing an example of the generation of a predicted image in the decoding device.

[0128] Figure 93 It is a flowchart showing another example of the generation of a predicted image in a decoding device.

[0129] Figure 94 It is a flowchart showing an example of the correction of a predicted image based on OBMC in a decoding device.

[0130] Figure 95 It is a flowchart showing an example of the correction of a predicted image based on BIO in a decoding device.

[0131] Figure 96 It is a flowchart showing an example of the correction of a predicted image based on LIC in a decoding device.

[0132] Figure 97 It is a diagram showing an example of the syntax in which an encoding device notifies one or more PTL (Profile / Tier / Level) parameters and one or more HRD (Hypothetical Reference Decoder) parameters in a VPS (Video Parameter Set).

[0133] Figure 98 It is a diagram showing an example of the syntax of PTL parameters.

[0134] Figure 99 It is a diagram showing an example of the syntax of HRD parameters.

[0135] Figure 100 It is a flowchart showing the process in which a decoding device analyzes the PTL parameters and HRD parameters notified in a VPS.

[0136] Figure 101 It is a diagram showing an example of the syntax in which one or more PTL parameters are notified in a VPS.

[0137] Figure 102 It is a flowchart showing the process in which an encoding device notifies PTL parameters in a VPS and an SPS (Sequence Parameter Set).

[0138] Figure 103 It is a diagram showing an example of the syntax in which one or more HRD parameters are notified in a VPS.

[0139] Figure 104 It is a flowchart showing the process in which an encoding device notifies HRD parameters in a VPS and an SPS.

[0140] Figure 105This is a diagram showing an example of the syntax for notifying one or more DPB (Decoded Picture Buffer) parameters in the VPS.

[0141] Figure 106 This is a flowchart showing the process by which an encoding device notifies DPB parameters in the VPS and SPS.

[0142] Figure 107 This is a diagram showing an example of the syntax for enabling the switching of whether to notify all PTL parameters, HRD parameters, and DPB parameters using the VPS in the second and third aspects of the present disclosure.

[0143] Figure 108 This is a flowchart showing an example of the operation of an encoding device in an embodiment.

[0144] Figure 109 This is a flowchart showing an example of the operation of a decoding device in an embodiment.

[0145] Figure 110 This is an overall structural diagram of a content supply system for implementing a content distribution service.

[0146] Figure 111 This is a diagram showing an example of the display screen of a web page.

[0147] Figure 112 This is a diagram showing an example of the display screen of a web page.

[0148] Figure 113 This is a diagram showing an example of a smart phone.

[0149] Figure 114 This is a block diagram showing an example of the structure of a smart phone. Detailed Embodiments

[0150] [Introduction]

[0151] The encoding device in the embodiment of the present disclosure includes a circuit and a memory connected to the circuit. During operation, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, performance requirement information indicating the performance requirements of the decoding device is encoded into the header shared by the layers included in the layer set, and a bitstream is generated. The bitstream includes the header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

[0152] Accordingly, the encoding device in the embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it may be possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0153] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the circuit encodes the performance requirement information into both the header shared by the plurality of layers and the header of a specific layer among the plurality of layers.

[0154] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes a bitstream including a plurality of layers.

[0155] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the circuit encodes the performance requirement information with the same content into the header shared by the plurality of layers and the header of a specific layer among the plurality of layers, respectively.

[0156] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes a bitstream including a plurality of layers.

[0157] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the layer set is an OLS (Output LayerSet).

[0158] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes a bitstream including a plurality of layers.

[0159] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the header shared by the plurality of layers is a VPS (Video Parameter Set).

[0160] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes a bitstream including a plurality of layers.

[0161] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the header of a specific layer is an SPS (Sequence Parameter Set).

[0162] Accordingly, the encoding device according to an embodiment of the present disclosure can use various filters and can obtain an image with better image quality.

[0163] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the performance requirement information represents PTL (Profile / Tier / Level) parameters indicating grades, ranks, and levels as performance requirements.

[0164] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it may be possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0165] In addition, for example, in the encoding device according to an embodiment of the present disclosure, when the number of layers in the layer set is 1, the circuit does not encode HRD (Hypothetical Reference Decoder) parameters into the header shared by multiple layers, but encodes the performance requirement information into the header shared by multiple layers.

[0166] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it may be possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0167] In addition, for example, in the encoding device according to an embodiment of the present disclosure, the circuit and a memory connected to the circuit are provided. During operation, for a layer included in an OLS with 1 layer, even if the number of layers in the OLS is 1, the information indicating the performance requirements of the decoding device is encoded into the VPS.

[0168] Accordingly, the encoding device according to an embodiment of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it may be possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0169] In addition, for example, in the decoding device according to an embodiment of the present disclosure, the circuit and a memory connected to the circuit are provided. During operation, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, the performance requirement information indicating the performance requirements of the decoding device is decoded from the header shared by the layers included in the layer set to generate a bitstream, and the bitstream includes the header shared by the layers included in the layer set and the data obtained by encoding at least one image included in the output layer.

[0170] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0171] In addition, in the decoding device in the embodiments of the present disclosure, the circuit decodes the performance requirement information from both the header shared by multiple layers and the header of a specific layer among the multiple layers.

[0172] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0173] In addition, for example, in the decoding device in the embodiments of the present disclosure, the layer set is OLS.

[0174] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0175] In addition, for example, in the decoding device in the embodiments of the present disclosure, the circuit decodes the performance requirement information of the same content from the header shared by multiple layers and the header of a specific layer among the multiple layers respectively.

[0176] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0177] In addition, for example, in the decoding device in the embodiments of the present disclosure, the header shared by multiple layers is the VPS.

[0178] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0179] In addition, for example, in the decoding device in the embodiments of the present disclosure, the header of a specific layer is the SPS.

[0180] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0181] In addition, for example, in the decoding device in the embodiments of the present disclosure, the performance requirement information is a PTL parameter that represents a profile, a level, or a tier as the performance requirement.

[0182] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0183] In addition, for example, in the decoding device in the embodiments of the present disclosure, for a layer set, when the number of layers in the layer set is 1, the circuit does not decode the HRD parameter from the header shared by multiple layers, but decodes the performance requirement information.

[0184] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0185] In addition, for example, in the decoding device in the embodiments of the present disclosure, there is a circuit and a memory connected to the circuit. During operation, for the layer included in the OLS (Output Layer Set) with the number of layers being 1, even if the number of layers in the OLS is 1, the circuit decodes the information representing the performance requirement of the decoding device from the VPS (Video Parameter Set).

[0186] Accordingly, the decoding device in the embodiments of the present disclosure can generate a bitstream that can obtain PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0187] In addition, for example, in the encoding method in the embodiments of the present disclosure, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, the performance requirement information representing the performance requirement of the decoding device is encoded into the header shared by the layers included in the layer set, and a bitstream is generated. The bitstream includes the header shared by the layers included in the layer set and the data obtained by encoding at least one image included in the output layer.

[0188] Thus, the encoding method in the embodiments of the present disclosure can achieve the same effect as the above-mentioned encoding device.

[0189] In addition, for example, in the decoding method in the embodiments of the present disclosure, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, performance requirement information representing the performance requirements of the decoding device is decoded from the header shared by the multiple layers included in the layer set to generate a bitstream, and the bitstream includes the header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

[0190] Thus, the decoding method in the embodiments of the present disclosure can achieve the same effect as the above-mentioned decoding device.

[0191] Moreover, for example, an encoding device according to an aspect of the present disclosure may include a splitting unit, an intra prediction unit, an inter prediction unit, a loop filter unit, a transformation unit, a quantization unit, and an entropy encoding unit.

[0192] The splitting unit may split a picture into a plurality of blocks. The intra prediction unit may perform intra prediction on the blocks included in the plurality of blocks. The inter prediction unit may perform inter prediction on the blocks. The transformation unit may transform the prediction error between the prediction image obtained by the intra prediction or the inter prediction and the original image to generate transform coefficients. The quantization unit may quantize the transform coefficients to generate quantization coefficients. The entropy encoding unit may encode the quantization coefficients to generate an encoded bitstream. The loop filter unit may apply a filter to the reconstructed image of the blocks.

[0193] In addition, for example, the encoding device may be an encoding device that encodes a moving image including a plurality of pictures.

[0194] Furthermore, for example, the entropy encoding unit may encode performance requirement information representing the performance requirements of the decoding device into the header shared by the layers included in the layer set to generate a bitstream, even if the number of layers included in the layer set is 1, and the bitstream includes the header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

[0195] Moreover, for example, a decoding device according to an aspect of the present disclosure may include an entropy decoding unit, an inverse quantization unit, an inverse transformation unit, an intra prediction unit, an inter prediction unit, and a loop filter unit.

[0196] The entropy decoding unit can decode the quantization coefficients of blocks within a picture from the encoded bitstream. The inverse quantization unit can perform inverse quantization on the quantization coefficients to obtain transform coefficients. The inverse transform unit can perform inverse transform on the transform coefficients to obtain prediction errors. The intra prediction unit can perform intra prediction on the block. The inter prediction unit can perform inter prediction on the block. The loop filtering unit can apply a filter to a reconstructed image generated using a predicted image obtained by the intra prediction or the inter prediction and the prediction error.

[0197] In addition, for example, the decoding device may be a decoding device that decodes a moving image including a plurality of pictures.

[0198] Moreover, for example, the entropy decoding unit, for a layer set including at least one output layer, even if the number of layers included in the layer set is 1, decodes performance requirement information indicating performance requirements of the decoding device from a header shared by the plurality of layers included in the layer set and generates a bitstream, the bitstream including a header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

[0199] Moreover, these inclusive or specific forms may also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.

[0200] [Definition of Terms]

[0201] As an example, each term may be defined as follows.

[0202] (1) Image

[0203] Is a unit of data composed of a set of pixels, composed of a picture or a block smaller than a picture, and includes still images in addition to moving images.

[0204] (2) Picture

[0205] Is a processing unit of an image composed of a set of pixels, and is sometimes referred to as a frame or a field.

[0206] (3) Block

[0207] Is a processing unit including a set of a specific number of pixels, and as listed in the following examples, the name is not limited. In addition, the shape is not limited either. For example, it naturally includes a rectangle composed of M×N pixels, a square composed of M×M pixels, and also includes triangles, circles, and other shapes.

[0208] (Examples of Blocks)

[0209] ·Slice / Tile / Brick

[0210] ·CTU / Superblock / Base Partition Unit

[0211] ·VPDU / Processing Partition Unit of Hardware

[0212] ·CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transform Block Unit (TU) / Unit

[0213] ·Sub-block

[0214] (4) Pixel / Sample

[0215] It is the point that is the smallest unit constituting an image, including not only pixels at integer positions but also pixels at fractional positions generated based on pixels at integer positions.

[0216] (5) Pixel Value / Sample Value

[0217] It is the inherent value that a pixel has, of course including luminance value, chromaticity difference value, grayscale of RGB, and also including depth value, or binary values of 0 and 1.

[0218] (6) Flag

[0219] In addition to 1 bit, it also includes cases of multiple bits. For example, it can also be a parameter or index of 2 bits or more. Additionally, not only binary-valued two-values are used, but also multi-values using other number systems can be used.

[0220] (7) Signal

[0221] It is symbolized and encoded for information transmission. In addition to discrete digital signals, it also includes analog signals taking continuous values.

[0222] (8) Stream / Bitstream

[0223] It refers to a data string of digital data or a stream of digital data. A stream / bitstream can be divided into multiple layers and composed of multiple streams in addition to a single stream. Additionally, in addition to the case of being transmitted on a single transmission path through serial communication, it also includes the case of being transmitted through packet communication in multiple transmission paths.

[0224] (9) Difference / Differential

[0225] In the case of a scalar, in addition to a simple difference (x - y), as long as it includes the operation of difference, it includes the absolute value of the difference (|x - y|), the square difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a and b are constants), and the offset difference (x - y + a: a is the offset).

[0226] (10) Sum

[0227] In the case of a scalar, in addition to the simple sum (x + y), any operation involving a sum is acceptable, including the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted sum (ax + by where a and b are constants), and the offset sum (x + y + a where a is the offset).

[0228] (11) Based on

[0229] This also includes cases where elements other than those that become the object on which it is based are added. Additionally, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0230] (12) Used, Using

[0231] This also includes cases where elements other than those that become the object being used are added. Additionally, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0232] (13) Prohibit, Forbid

[0233] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation.

[0234] (14) Limit, Restriction / Restrict / Restricted

[0235] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation. Also, as long as a part is prohibited in terms of quantity or quality, it includes cases where it is prohibited comprehensively.

[0236] (15) Chroma

[0237] It is an adjective represented by the notations Cb and Cr that designates one of two color difference signals associated with the primary colors for a sample arrangement or a single sample representation. Additionally, the term chrominance can also be used instead of the term chroma.

[0238] (16) Luma

[0239] It is an adjective represented by the notation or subscript Y or L that designates a monochrome signal associated with the primary colors for a sample arrangement or a single sample representation. The term luminance can also be used instead of the term luma.

[0240] [Regarding the explanations in the description]

[0241] In the accompanying drawings, the same reference numerals denote the same or similar components. In addition, the dimensions and relative positions of the components in the accompanying drawings are not necessarily drawn to a certain scale.

[0242] Hereinafter, embodiments will be specifically described with reference to the accompanying drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and do not limit the meaning of the claims.

[0243] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processing and / or structure described in each aspect of the present disclosure. The processing and / or structure can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processing and / or structure applied to the embodiments, any of the following can be performed.

[0244] (1) Among a plurality of components of the encoding device or decoding device of the embodiment described in each aspect of the present disclosure, a certain component can be replaced with another component described in a certain aspect of the present disclosure, or they can be combined;

[0245] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can be made to the functions or processes performed by a part of the components of the encoding device or decoding device. For example, any function or process can be replaced with another function or process described in a certain aspect of the present disclosure, or they can be combined;

[0246] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can be made to a part of the plurality of processes included in the method. For example, any process in the method can be replaced with another process described in a certain aspect of the present disclosure, or they can be combined;

[0247] (4) A part of the components constituting the encoding device or decoding device of the embodiment can be combined with the components described in a certain aspect of the present disclosure, or can be combined with the components having a part of the functions described in a certain aspect of the present disclosure, or can be combined with the components that perform a part of the processes performed by the components described in each aspect of the present disclosure;

[0248] (5) A component that is part of the function of the encoding device or decoding device of the embodiment, or a component that performs part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced by a component described in one of the various forms of the present disclosure, a component that is part of the function described in one of the various forms of the present disclosure, or a component that performs part of the processing described in one of the various forms of the present disclosure;

[0249] (6) In the method implemented by the encoding device or decoding device of the embodiment, one of the multiple processes included in the method is replaced by a process described in one of the various forms of the present disclosure or a similar process, or they are combined;

[0250] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various forms of the present disclosure.

[0251] (8) The manner of implementing the processes and / or structures described in the various forms of the present disclosure is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the moving image encoding or moving image decoding disclosed in the embodiment.

[0252] [System Structure]

[0253] Figure 1 It is a schematic diagram showing an example of the structure of the transmission system of the present embodiment.

[0254] The transmission system Trs is a system that transmits the stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs is, for example, as Figure 1 shown, including an encoding device 100, a network Nw, and a decoding device 200.

[0255] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream contains, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding.

[0256] In addition, the original image input to the encoding device 100 before being encoded is also referred to as the original image, original signal, or original sample. Additionally, the image can be a moving image or a still image. Furthermore, an image is a superordinate concept of sequences, pictures, blocks, etc., and is not restricted by spatial and temporal regions unless otherwise specified. Also, an image is composed of an arrangement of pixels or pixel values, and the signal or pixel value representing the image is also called a sample. Moreover, a stream can be referred to as a bitstream, encoded bitstream, compressed bitstream, or encoded signal. Furthermore, the encoding device can also be called an image encoding device or a moving image encoding device, and the encoding method of the encoding device 100 can also be called an encoding method, image encoding method, or moving image encoding method.

[0257] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Additionally, the network Nw can also be replaced by a storage medium that records the stream such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0258] The decoding device 200 generates a decoded image, such as an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method of the encoding device 100.

[0259] Additionally, the decoding device can also be called an image decoding device or a moving image decoding device, and the decoding method of the decoding device 200 can also be called a decoding method, image decoding method, or moving image decoding method.

[0260] [Data Structure]

[0261] Figure 2 is a diagram showing an example of the hierarchical structure of the data in the stream. The stream includes, for example, a video sequence. This video sequence, as shown in (a) of Figure 2 , contains a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and a plurality of pictures.

[0262] In a moving image in which a VPS includes a plurality of layers, the encoding parameters common to the plurality of layers and the plurality of layers included in the moving image, or the encoding parameters associated with each layer.

[0263] An SPS includes parameters used for a sequence, that is, encoding parameters that a decoding device 200 refers to for decoding the sequence. For example, the encoding parameters may also represent the width or height of a picture. In addition, there may be multiple SPSs.

[0264] A PPS includes parameters used for a picture, that is, encoding parameters that a decoding device 200 refers to for decoding each picture in the sequence. For example, the encoding parameters may also include a reference value of a quantization width used in decoding the picture and a flag indicating the application of weighted prediction. In addition, there may be multiple PPSs. In addition, SPS and PPS are sometimes simply referred to as parameter sets.

[0265] As Figure 2 shown in (b) of [], a picture may include a picture header and one or more slices. The picture header includes encoding parameters that a decoding device 200 refers to for decoding the one or more slices.

[0266] As Figure 2 shown in (c) of [], a slice includes a slice header and one or more tiles. The slice header includes encoding parameters that a decoding device 200 refers to for decoding the one or more tiles.

[0267] As Figure 2 shown in (d) of [], a tile includes one or more CTUs (Coding Tree Units).

[0268] In addition, a picture may not include a slice and include a tile group instead of the slice. In this case, the tile group includes one or more tiles. In addition, a slice may be included in a tile.

[0269] A CTU is also referred to as a superblock or a basic segmentation unit. As Figure 2 shown in (e) of [], such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes encoding parameters that a decoding device 200 refers to for decoding the one or more CUs.

[0270] A CU may also be divided into a plurality of small CUs. In addition, as Figure 2As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual described later. In addition, the CU is basically the same as the PU (Prediction Unit) and the TU (Transform Unit). However, for example, in the SBT described later, it may also include a plurality of TUs smaller than the CU. In addition, the CU can also process each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. The VPDU is, for example, a fixed unit that can be processed in one stage during pipeline processing in hardware.

[0271] In addition, the stream may not have Figure 2 any part of the hierarchies shown. In addition, the order of these hierarchies can be swapped, and any one hierarchy can be replaced with another hierarchy. In addition, the picture that is the object of the processing performed by the device such as the encoding device 100 or the decoding device 200 at the current time point is called the current picture. If the processing is encoding, the current picture is synonymous with the encoding target picture. If the processing is decoding, the current picture is synonymous with the decoding target picture. In addition, a block such as a CU or a CU that is the object of the processing performed by the device such as the encoding device 100 or the decoding device 200 at the current time point is called the current block. If the processing is encoding, the current block is synonymous with the encoding target block. If the processing is decoding, the current block is synonymous with the decoding target block.

[0272] [Structural Slices / Tiles of Picture]

[0273] For parallel decoding of a picture, the picture may sometimes be composed of slices or tiles.

[0274] A slice is the basic encoding unit that constitutes a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs.

[0275] Figure 3This is a diagram showing an example of the structure of a slice. For example, the picture contains 11×8 CTUs and is divided into 4 slices (Slice 1 - 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU in the picture belongs to a certain slice. The shape of the slice is the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can be anywhere among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. In addition, the slice contains a slice header and encoded data. In the slice header, features of the slice such as the starting CTU address and slice type of the slice can also be described.

[0276] A tile is a unit of a rectangular area that makes up a picture. Numbers called TileId can also be assigned to each tile in the raster scan order.

[0277] Figure 4 This is a diagram showing an example of the structure of a tile. For example, the picture contains 11×8 CTUs and is divided into 4 rectangular area tiles (Tile 1 - 4). When tiles are used, the processing order of the CTUs is changed compared to the case where tiles are not used. When tiles are not used, multiple CTUs in the picture are processed in the raster scan order, for example. When tiles are used, in each of the multiple tiles, at least 1 CTU is processed in the raster scan order, for example. For example, as Figure 4 shown, the processing order of the multiple CTUs contained in Tile 1 is the order from the left end of the first column of Tile 1 towards the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 towards the right end of the second column of Tile 1.

[0278] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.

[0279] Furthermore, a picture can also be composed of tile set units. A tile set can contain more than one tile group and can also contain more than one tile. A picture can be composed of only one of a tile set, a tile group, and a tile. For example, the order of scanning multiple tiles in the raster order for each tile set is set as the basic encoding order of the tiles. A set of one or more tiles that are consecutive in the basic encoding order within each tile set is set as a tile group. Such a picture can also be composed of the later-described dividing unit 102 (refer to Figure 7 ).

[0280] [Scalable Coding]

[0281] Figure 5 and Figure 6This is a diagram showing an example of the structure of a scalable stream.

[0282] As Figure 5 shown, the encoding device 100 can generate a temporally / spatially scalable stream by encoding multiple pictures into a certain layer among multiple layers. For example, the encoding device 100 achieves scalability where the enhancement layer exists above the base layer by encoding pictures for each layer. The encoding of each such picture is called scalable encoding. Thereby, the decoding device 200 can switch the picture quality of the image displayed by decoding this stream. That is, the decoding device 200 determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth. As a result, the decoding device 200 can freely switch the same content between low-resolution content and high-resolution content for decoding. For example, a user of this stream, while on the move, uses a smartphone to view and listen to a moving image of this stream halfway, and after returning home, uses a device such as an Internet TV to view and listen to the subsequent part of this moving image. Additionally, decoding devices 200 with the same or different performances are respectively assembled in the above-mentioned smartphone and device. In this case, if the device decodes to the upper layer in this stream, the user can view and listen to a high-quality moving image after returning home. Thus, the encoding device 100 does not need to generate multiple streams with the same content but different picture qualities, and can reduce the processing load.

[0283] Furthermore, the enhancement layer can also include meta-information such as based on the statistical information of the image. It can also be that the decoding device 200 generates a high-quality moving image by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can be either improving the SNR at the same resolution or expanding the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in filtering processes, machine learning, or least squares operations used in the super-resolution process, etc.

[0284] Alternatively, the picture can also be segmented into tiles, etc. according to the meaning of each object, etc. within the picture. In this case, the decoding device 200 can decode only a partial area in the picture by selecting the tile to be decoded. Moreover, the attributes of the object (person, car, ball, etc.) and the position within the picture (coordinate position in the same picture, etc.) can be saved as meta-information. In this case, the decoding device 200 can determine the position of the desired object based on the meta-information and decide the tile containing the object. For example, as Figure 6 shown, a data storage structure different from the image data, such as SEI in HEVC, can also be used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0285] In addition, the meta-information may be stored in units composed of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding device 200 can obtain the moments when a specific person appears in the moving image, etc. By using this moment and the information of the picture unit, the picture in which the target exists and the position of the target in that picture can be determined.

[0286] [Encoding Device]

[0287] Next, the encoding device 100 of the embodiment will be described. Figure 7 It is a block diagram showing an example of the functional structure of the encoding device 100 of the embodiment. The encoding device 100 encodes an image in units of blocks.

[0288] As Figure 7 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. In addition, the intra prediction unit 124 and the inter prediction unit 126 are respectively configured as a part of the prediction processing unit.

[0289] [Installation Example of Encoding Device]

[0290] Figure 8 It is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 7 as shown, the multiple components of the encoding device 100 are implemented by the Figure 8 shown processor a1 and memory a2.

[0291] The processor a1 is a circuit for information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding an image. The processor a1 may also be a processor such as a CPU. In addition, the processor a1 may also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 may also function as Figure 7 the multiple components of the encoding device 100 shown, excluding the components for storing information.

[0292] Memory a2 is a dedicated or general-purpose memory that stores information for the processor a1 to encode an image. Memory a2 can be either an electronic circuit or connected to the processor a1. Additionally, memory a2 can be included in the processor a1. Moreover, memory a2 can also be an aggregate of multiple electronic circuits. Furthermore, memory a2 can be a magnetic disk, an optical disc, etc., or can be represented as a storage or a recording medium, etc. Additionally, memory a2 can be either a non-volatile memory or a volatile memory.

[0293] For example, memory a2 can store the encoded image or the stream corresponding to the encoded image. Additionally, a program for the processor a1 to encode an image can also be stored in memory a2.

[0294] Moreover, for example, memory a2 can also serve as Figure 7 the component for storing information among the multiple components of the encoding device 100 shown. Specifically, memory a2 can serve as Figure 7 the block memory 118 and the frame memory 122 shown. More specifically, reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.) can be stored in memory a2.

[0295] Moreover, in the encoding device 100, not all of the multiple components shown in Figure 7 may be installed, and not all of the above-mentioned multiple processes may be performed. Figure 7 A part of the multiple components shown in

[0296] may be included in other devices, or a part of the above-mentioned multiple processes may be executed by other devices.

[0297] [Overall Process of Encoding Processing]

[0298] Figure 9 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0299] First, the segmentation unit 102 of the encoding device 100 divides the pictures included in the original image into multiple blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block with the fixed size into multiple blocks that constitute the selected segmentation pattern. Then, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 for each of the multiple blocks.

[0300] The prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126, and the prediction control unit 128 generate a prediction image of the current block (step Sa_3). In addition, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.

[0301] Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual is also referred to as a prediction error.

[0302] Next, the transform unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing transformation and quantization on the prediction image (step Sa_5).

[0303] Next, the entropy encoding unit 110 generates a stream by encoding the plurality of quantization coefficients and prediction parameters related to the generation of the prediction image (specifically, entropy encoding) (step Sa_6).

[0304] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore the prediction residual by performing inverse quantization and inverse transformation on the plurality of quantization coefficients (step Sa_7).

[0305] Next, the addition unit 116 reconstructs the current block by adding the restored prediction residual to the prediction image (step Sa_8). Thereby, a reconstructed image is generated. In addition, the reconstructed image is also referred to as a reconstructed block. In particular, the reconstructed image generated by the encoding device 100 is also referred to as a local decoded block or a local decoded image.

[0306] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).

[0307] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.

[0308] In addition, in the above example, the encoding device 100 selects one segmentation style for a block of a fixed size and encodes each block according to the segmentation style. However, each block may also be encoded according to each of a plurality of segmentation styles. In this case, the encoding device 100 can evaluate the cost for each of the plurality of segmentation styles, and for example, can select the stream obtained by encoding according to the segmentation style with the minimum cost as the finally output stream.

[0309] In addition, the processing of these steps Sa_1 to Sa_10 can be sequentially performed by the encoding device 100, a plurality of parts of these processes can be performed in parallel, or the order can be changed.

[0310] The encoding process of such an encoding device 100 uses hybrid encoding that combines predictive encoding and transform encoding. In addition, the predictive encoding is performed through an encoding loop, which is composed of a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. That is, the prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126 constitutes a part of the encoding loop.

[0311] [Segmentation Unit]

[0312] The segmentation unit 102 divides each picture included in the original image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first divides the picture into blocks of a fixed size (e.g., 128×128 pixels). Such blocks of the fixed size are sometimes referred to as coding tree units (CTUs). And, the segmentation unit 102 divides each block of the fixed size into blocks of variable size (e.g., 64×64 pixels or less) based on recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. Such blocks of variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture can also be used as the processing unit for CUs, PUs, or TUs.

[0313] Figure 10 is a diagram showing an example of block segmentation of an embodiment. In Figure 10 the solid lines represent block boundaries based on quadtree block segmentation, and the dashed lines represent block boundaries based on binary tree block segmentation.

[0314] Here, block 10 is a square block of 128×128 pixels. This block 10 is first divided into 4 square blocks of 64×64 pixels (quadtree block segmentation).

[0315] The upper left 64×64 pixel square block is further vertically divided into 2 rectangular blocks each composed of 32×64 pixels, and the left 32×64 pixel rectangular block is further vertically divided into 2 rectangular blocks each composed of 16×64 pixels (binary tree block segmentation). As a result, the upper left 64×64 pixel square block is divided into 2 rectangular blocks 11 and 12 of 16×64 pixels and a rectangular block 13 of 32×64 pixels.

[0316] The upper right 64×64 pixel square block is horizontally divided into 2 rectangular blocks 14 and 15 each composed of 64×32 pixels (binary tree block segmentation).

[0317] The 64×64 pixel square block in the lower left is divided into 4 square blocks (quad-tree block division), each consisting of 32×32 pixels. The upper left and lower right blocks among the 4 square blocks, each consisting of 32×32 pixels, are further divided. The upper left 32×32 pixel square block is vertically divided into 2 rectangular blocks, each consisting of 16×32 pixels, and the right rectangular block consisting of 16×32 pixels is further horizontally divided into 2 square blocks, each consisting of 16×16 pixels (binary tree block division). The lower right 32×32 pixel square block is horizontally divided into 2 rectangular blocks, each consisting of 32×16 pixels (binary tree block division). As a result, the 64×64 pixel square block in the lower left is divided into 16 rectangular blocks 16 of 16×32 pixels, 2 square blocks 17 and 18 of 16×16 pixels each, 2 square blocks 19 and 20 of 32×32 pixels each, and 2 rectangular blocks 21 and 22 of 32×16 pixels each.

[0318] The block 23 consisting of 64×64 pixels in the lower right is not divided.

[0319] As described above, in Figure 10 , the block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.

[0320] In addition, in Figure 10 , 1 block is divided into 4 or 2 blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, 1 block may be divided into 3 blocks (ternary tree division). Divisions including such ternary tree division are sometimes called MBT (multi type tree) division.

[0321] Figure 11 is a diagram showing an example of the functional structure of the division unit 102. As Figure 11 shown, the division unit 102 may also include a block division determination unit 102a. As an example, the block division determination unit 102a may perform the following processing.

[0322] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the above-described division pattern based on this block information. The division unit 102 divides the original image according to this division pattern and outputs one or more blocks obtained by this division to the subtraction unit 104.

[0323] In addition, the block segmentation determination unit 102a outputs, for example, parameters indicating the above-described segmentation pattern to the transformation unit 106, the inverse transformation unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy encoding unit 110. The transformation unit 106 can transform the prediction residual based on the parameters, and the intra prediction unit 124 and the inter prediction unit 126 can generate a prediction image based on the parameters. In addition, the entropy encoding unit 110 can also perform entropy encoding on the parameters.

[0324] As an example, parameters related to the segmentation pattern can also be written into the stream as follows.

[0325] Figure 12 FIG. is an example showing a segmentation pattern. Examples of the segmentation pattern include: four-way split (QT), in which the block is split into two in the horizontal and vertical directions respectively; three-way split (HT or VT), in which the block is split in the same direction at a ratio of 1:2:1; two-way split (HB or VB), in which the block is split in the same direction at a ratio of 1:1; and no split (NS).

[0326] In addition, in the case of four-way split and no split, the segmentation pattern does not have a block segmentation direction, and in the case of two-way split and three-way split, the segmentation pattern has segmentation direction information.

[0327] Figure 13A and Figure 13B FIG. is an example of a syntax tree showing a segmentation pattern. In the example of Figure 13A , first, there is information indicating whether to perform segmentation (S: Split flag), and then there is information indicating whether to perform four-way split (QT: QT flag). Next, there is information indicating whether to perform three-way split or two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the segmentation direction (Ver: Vertical flag or Hor: Horizontal flag). In addition, for each of one or more blocks obtained by such segmentation based on the segmentation pattern, the same process can be repeatedly applied for further segmentation. That is, as an example, it is also possible to recursively perform determination of whether to perform segmentation, whether to perform four-way split, whether the segmentation method is horizontal or vertical, and whether to perform three-way split or two-way split, and encode the determination results implemented into the stream in the encoding order disclosed in the syntax tree shown in Figure 13A .

[0328] In addition, in the syntax tree shown in Figure 13A , these pieces of information are arranged in the order of S, QT, TT, Ver, but they can also be arranged in the order of S, QT, Ver, BT. That is, in Figure 13BIn the example, first, there is information indicating whether to perform splitting (S: Split flag), then there is information indicating whether to perform four-way splitting (QT: QT flag). Next, there is information indicating the splitting direction (Ver: Vertical flag or Hor: Horizontal flag), and finally there is information indicating whether to perform two-way splitting or three-way splitting (BT: BT flag or TT: TT flag).

[0329] In addition, the splitting pattern described here is an example. A splitting pattern other than the described one can be used, or only a part of the described splitting pattern can be used.

[0330] [Subtraction unit]

[0331] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units input from and split by the splitting unit 102. That is, the subtraction unit 104 calculates the prediction residual of the current block. And the subtraction unit 104 outputs the calculated prediction residual to the transformation unit 106.

[0332] The original image is an input signal of the encoding device 100, for example, a signal representing an image of each picture constituting a moving image (for example, a luma signal and two chroma signals).

[0333] [Transformation unit]

[0334] The transformation unit 106 transforms the prediction residual in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction residual in the spatial domain.

[0335] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types, use a transform basis function corresponding to the selected transformation type, and transform the prediction residual into transform coefficients. Such a transformation is sometimes referred to as EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform). In addition, the transform basis function is sometimes simply referred to as a basis.

[0336] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. In addition, these transformation types can be respectively described as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14is a table representing transform basis functions corresponding to respective transform types. In Figure 14 , N represents the number of input pixels. The selection of a transform type from among these multiple transform types can depend, for example, on the type of prediction (intra prediction, inter prediction, etc.) or on the intra prediction mode.

[0337] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. In addition, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0338] In addition, the transform unit 106 may also perform a re - transform on the transform coefficients (i.e., the transform result). Such a re - transform may be a case of what is called AST (adaptive secondary transform) or NSST (non - separable secondary transform). For example, the transform unit 106 performs a re - transform on each sub - block (e.g., a 4×4 pixel sub - block) included in a block of transform coefficients corresponding to an intra - prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are generally signaled at the CU level. In addition, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0339] In the transform unit 106, a separable transform and a non - separable transform may also be applied. A separable transform is a method in which multiple transforms are performed separately in each direction according to the number of input dimensions, and a non - separable transform is a method in which when the input is multi - dimensional, two or more dimensions are regarded as one dimension and transformed together.

[0340] For example, as an example of a non - separable transform, when the input is a 4×4 pixel block, it can be regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.

[0341] In addition, in a further example of a non - separable transform, after regarding a 4×4 pixel input block as a permutation having 16 elements, a transform (Hypercube Givens Transform) in which Givens rotations are performed on this permutation multiple times may also be performed.

[0342] In the transformation in the transformation unit 106, the type of transformation of the transformation basis function to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform).

[0343] Figure 15 It is a diagram showing an example of SVT.

[0344] In SVT, as Figure 15 shown, the CU is bisected in the horizontal direction or the vertical direction, and only one of the regions is transformed into the frequency domain. The transformation type can be set for each region. For example, DST7 and DCT8 are used. For example, for the region at position 0 among the two regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used. Or, for the region at position 1 among the two regions, DST7 is used. Similarly, for the region at position 0 among the two regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 are used. Or, for the region at position 1 among the two regions, DST7 is used. In such an Figure 15 example shown, only one of the two regions within the CU is transformed, and the other is not transformed, but it is also possible to transform the two regions separately. In addition, the splitting method is not limited to bisection, but can also be quartering. Furthermore, it can be more flexible, encoding the information indicating the splitting method and performing signaling in the same way as CU splitting. In addition, SVT is sometimes also referred to as SBT (Sub-block Transform).

[0345] The aforementioned AMT and EMT can also be referred to as MTS (Multiple Transform Selection). In the case of applying MTS, transform types such as DST7 or DCT8 can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. On the other hand, as a process of selecting the transform type used in the orthogonal transform based on the shape of the CU without encoding the index information, there is a process called IMTS (Implicit MTS). In the case of applying IMTS, for example, if the shape of the CU is rectangular, DST7 is used on the short side of the rectangle and DCT2 is used on the long side, and orthogonal transforms are performed respectively. Additionally, for example, in the case where the shape of the CU is square, if MTS is effective in the sequence, DCT2 is used for orthogonal transform, and if MTS is ineffective, DST7 is used for orthogonal transform. DCT2 and DST7 are just examples, and other transform types can be used, or different combinations of the used transform types can be set. IMTS can be used only in blocks for intra prediction, or can be used together in blocks for intra prediction and blocks for inter prediction.

[0346] As described above, as selection processes for selectively switching the transform type used in the orthogonal transform, the three processes of MTS, SBT, and IMTS have been described. However, all three selection processes can be effective, or only some of the selection processes can be selectively made effective. Regarding whether each selection process is effective, it can be identified by flag information in headers such as SPS. For example, if all three selection processes are effective, one is selected from the three selection processes in units of CUs for orthogonal transform. Additionally, as long as the selection process for selectively switching the transform type can achieve at least one of the following four functions [1] to [4], a selection process different from the above three selection processes can be used, or the above three selection processes can be replaced with other processes respectively. Function [1] is the function of performing an orthogonal transform on the entire range within the CU and encoding the information indicating the transform type used in the transform. Function [2] is the function of performing an orthogonal transform on the entire range of the CU, determining the transform type based on a predetermined rule without encoding the information indicating the transform type. Function [3] is the function of performing an orthogonal transform on a partial region of the CU and encoding the information indicating the transform type used in the transform. Function [4] is the function of performing an orthogonal transform on a partial region of the CU, not encoding the information indicating the transform type used in the transform, and determining the transform type based on a predetermined rule, etc.

[0347] In addition, the presence or absence of the application of MTS, IMTS, and SBT respectively can also be determined for each processing unit. For example, it can be determined for each sequence unit, picture unit, tile unit, slice unit, CTU unit, or CU unit.

[0348] In addition, the tool for selectively switching the transformation type in the present disclosure may also be referred to as a method for adaptively selecting the basis used in the transformation process, a selection process, or a process for selecting the basis. In addition, the tool for selectively switching the transformation type may also be referred to as a mode for adaptively selecting the transformation type.

[0349] Figure 16 It is a flowchart showing an example of the process performed by the transformation unit 106.

[0350] For example, the transformation unit 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined to perform an orthogonal transformation (Yes in step St_1), the transformation unit 106 selects the transformation type used for the orthogonal transformation from among multiple transformation types (step St_2). Next, the transformation unit 106 performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Then, the transformation unit 106 outputs the information indicating the selected transformation type to the entropy coding unit 110, thereby encoding this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transformation (No in step St_1), the transformation unit 106 outputs the information indicating that no orthogonal transformation is performed to the entropy coding unit 110, thereby encoding this information (step St_5). In addition, the determination of whether to perform an orthogonal transformation in step St_1 can be determined based on, for example, the size of the transformation block, the prediction mode applied to the CU, etc. In addition, the information indicating the transformation type used for the orthogonal transformation may not be encoded, and an orthogonal transformation may be performed using a pre-specified transformation type.

[0351] Figure 17 It is a flowchart showing another example of the process performed by the transformation unit 106. In addition, Figure 17 The example shown is the same as the example shown in Figure 16 and is an example of an orthogonal transformation in the case of applying a method for selectively switching the transformation type used for the orthogonal transformation.

[0352] As an example, the first transformation type group may include DCT2, DST7, and DCT8. In addition, as an example, the second transformation type group may include DCT2. In addition, the transformation types included in the first transformation type group and the second transformation type group may be partially repeated or may be all different transformation types.

[0353] Specifically, the transform unit 106 determines whether the transform size is equal to or less than a specified value (step Su_1). Here, when it is determined that the size is equal to or less than the specified value (yes in step Su_1), the transform unit 106 orthogonally transforms the prediction residual of the current block using the transform types included in the first transform type group (step Su_2). Next, the transform unit 106 encodes the information indicating which transform type among one or more transform types included in the first transform type group is used by outputting the information to the entropy encoding unit 110 (step Su_3). On the other hand, when it is determined that the transform size is not equal to or less than the specified value (no in step Su_1), the transform unit 106 orthogonally transforms the prediction residual of the current block using the second transform type group (step Su_4).

[0354] In step Su_3, the information indicating the transform type used for the orthogonal transform may be information representing a combination of the transform type applied to the vertical direction of the current block and the transform type applied to the horizontal direction. Additionally, the first transform type group may include only one transform type, and the information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and the information indicating the transform type used in the orthogonal transform among one or more transform types included in the second transform type group may also be encoded.

[0355] Alternatively, the transform type may be determined based only on the transform size. Furthermore, if the process of determining the transform type used for the orthogonal transform is based on the transform size, it is not limited to the determination of whether the transform size is equal to or less than the specified value.

[0356] [Quantization Unit]

[0357] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the multiple transform coefficients of the current block in a specified scan order and quantizes the transform coefficients based on the quantization parameter (QP) corresponding to the scanned transform coefficients. Then, the quantization unit 108 outputs the quantized multiple transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0358] The specified scan order is the order used for quantization / inverse quantization of the transform coefficients. For example, the specified scan order is defined by ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).

[0359] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the error (quantization error) of the quantization coefficients increases.

[0360] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transformation sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing the values sampled at a predetermined interval by associating them with predetermined levels, and in this technical field, expressions such as rounding, truncation, or scaling are sometimes used.

[0361] As methods of using the quantization matrix, there are a method of using the quantization matrix directly set on the encoding device 100 side and a method of using the default quantization matrix (default matrix). On the encoding device 100 side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix. In addition, instead of directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix used in the quantization of the current block can be generated based on the default quantization matrix or the encoded quantization matrix.

[0362] On the other hand, there is also a method of performing quantization in such a way that the coefficients of the high-frequency components and the coefficients of the low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix (flat matrix) in which all the coefficients are the same value.

[0363] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level.

[0364] When using the quantization matrix, the quantization unit 108 scales, for example, the quantization width obtained according to the quantization parameter, etc. for each transform coefficient using the value of the quantization matrix. The quantization process without using the quantization matrix can also be a process of quantizing the transform coefficient based on the quantization width obtained according to the quantization parameter, etc. In addition, in the quantization process without using the quantization matrix, a predetermined value common to all the transform coefficients in the block can also be multiplied by the quantization width.

[0365] Figure 18 It is a block diagram showing an example of the functional structure of the quantization unit 108.

[0366] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108a, a prediction quantization parameter generation unit 108b, a quantization parameter generation unit 108c, a quantization parameter storage unit 108d, and a quantization processing unit 108e.

[0367] Figure 19 It is a flowchart showing an example of the quantization performed by the quantization unit 108.

[0368] As an example, the quantization unit 108 can be based onFigure 19 The flowchart shown performs quantization for each CU. Specifically, the quantization parameter generation unit 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generation unit 108c generates quantization parameters for the current block (step Sv_2) and saves the quantization parameters in the quantization parameter storage unit 108d (step Sv_3).

[0369] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). Then, the predicted quantization parameter generation unit 108b obtains quantization parameters of a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). The predicted quantization parameter generation unit 108b generates predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generation unit 108c and the predicted quantization parameters of the current block generated by the predicted quantization parameter generation unit 108b (step Sv_7). By calculating this difference, differential quantization parameters are generated. The differential quantization parameter generation unit 108a outputs the differential quantization parameters to the entropy encoding unit 110, thereby encoding the differential quantization parameters (step Sv_8).

[0370] In addition, the differential quantization parameters can also be encoded at the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameters can be generated using the initial values of the quantization parameters and the differential quantization parameters.

[0371] In addition, the quantization unit 108 can include multiple quantizers, and dependent quantization that quantizes the transform coefficients using a quantization method selected from multiple quantization methods can also be applied.

[0372] [Entropy Encoding Unit]

[0373] Figure 20 is a block diagram showing an example of the functional structure of the entropy encoding unit 110.

[0374] The entropy encoding unit 110 performs entropy encoding on the quantized coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130, thereby generating a stream. In this entropy encoding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy encoding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization that transforms multi-value signals such as quantized coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, etc. The context control unit 110b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. In this method of deriving the context value, for example, there are bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic coding unit 110c performs arithmetic coding on the binarized signal using the derived context value.

[0375] Figure 21 It is a diagram showing the process of CABAC in the entropy encoding unit 110.

[0376] First, in the CABAC in the entropy encoding unit 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding unit 110c and setting of the initial context value are performed. Then, the binarization unit 110a and the binary arithmetic coding unit 110c sequentially perform binarization and arithmetic coding on the multiple quantized coefficients of the CTU, for example. At this time, the context control unit 110b updates the context value each time arithmetic coding is performed. Then, the context control unit 110b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.

[0377] [Inverse Quantization Unit]

[0378] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0379] [Inverse Transform Unit]

[0380] The inverse transform unit 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Further, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.

[0381] In addition, since information is usually lost due to quantization in the restored prediction residual, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual usually includes a quantization error.

[0382] [Addition unit]

[0383] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 and the predicted image input from the prediction control unit 128. As a result, a reconstructed image is generated. Further, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.

[0384] [Block memory]

[0385] The block memory 118 is a storage unit for storing blocks referred to in intra prediction and is a block within the current picture. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.

[0386] [Frame memory]

[0387] The frame memory 122 is a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.

[0388] [Loop filter unit]

[0389] The loop filter unit 120 performs a loop filter process on the reconstructed image output from the addition unit 116 and outputs the reconstructed image after the filter process to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).

[0390] Figure 22 is a block diagram showing an example of the functional configuration of the loop filter unit 120.

[0391] For example Figure 22As shown, the loop filter unit 120 includes a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 120b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 120c applies the above-described ALF processing to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is a process of improving the image quality by reducing ringing (a phenomenon in which pixel values around the edge are deformed in a fluctuating manner) and correcting the deviation of pixel values. In this SAO processing, for example, there are edge offset processing and band offset processing. In addition, the loop filter unit 120 may not include Figure 22 all the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 120 may also be structured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 22 .

[0392] [Loop Filter Unit > Adaptive Loop Filter]

[0393] In the ALF, a least squares error filter used to remove coding distortion is employed. For example, for each 2×2 pixel sub-block within the current block, one filter selected from multiple filters based on the direction and activity of the gradient based on locality is used.

[0394] Specifically, first, sub-blocks (for example, 2×2 pixel sub-blocks) are classified into multiple classes (for example, 15 or 25 classes). The classification of sub-blocks is performed, for example, based on the direction and activity of the gradient. In a specific example, using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4), the classification value C (for example, C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes.

[0395] The gradient direction value D is derived, for example, by comparing the gradients in multiple directions (for example, horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantifying the addition result.

[0396] Based on the result of such classification, the filter to be used for the sub-block is determined from among multiple filters.

[0397] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 23A to 23C It is a diagram showing multiple examples of the shape of the filter used in the ALF. Figure 23A It shows a 5×5 rhombus-shaped filter,Figure 23B Represents a 7×7 rhombus-shaped filter, Figure 23C Represents a 9×9 rhombus-shaped filter. Information indicating the shape of the filter is typically signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0398] The on / off of ALF can also be determined, for example, at the picture level or CU level. For example, for luminance, it can be determined at the CU level whether to employ ALF, and for chrominance difference, it can be determined at the picture level whether to employ ALF. Information indicating the on / off of ALF is typically signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of ALF need not be limited to the picture level or CU level and can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0399] Additionally, as described above, one filter is selected from multiple filters and ALF processing is applied to the sub-block. For each of these multiple filters (e.g., up to 15 or 25 filters), the set of coefficients composed of the multiple coefficients used in the filter is typically signaled at the picture level. Additionally, the signaling of the set of coefficients need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0400] [Loop Filter > Cross Component Adaptive Loop Filter (or Cross Component Adaptive Loop Filter)]

[0401] Figure 23D Is a diagram showing an example where the Y sample (the first component) is used for the CCALF of Cb and the CCALF of Cr (multiple components different from the first component). Figure 23E Is a diagram showing a rhombus-shaped filter.

[0402] One example of CC-ALF is by using a linear rhombus-shaped filter ( Figure 23D , Figure 23E) It operates on the luminance channels of each color difference component. For example, filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled to have variable block sizes and is notified by a context-encoded flag received for each block of samples. The block size and the CC-ALF enable flag are received at the slice level of each color difference component. The syntax and semantics of CC-ALF are provided in the Appendix. In this document, block sizes of 16x16, 32x32, 64x64, and 128x128 are supported (for color difference samples).

[0403] [Loop Filter>Joint Chroma Cross Component Adaptive Loop Filter]

[0404] Figure 23F is a diagram showing an example of JC-CCALF. Figure 23G is a diagram showing an example of the weight_index candidates of JC-CCALF.

[0405] One example of JC-CCALF uses only one CCALF filter, generates one CCALF filter output as the color difference adjustment signal for only one color component, and applies a properly weighted version of the same color difference adjustment signal to other color components. In this way, the complexity of the existing CCALF is approximately halved.

[0406] The weight value is encoded into a sign flag and a weight index. The weight index (denoted as weight_index) is encoded as 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight. It cannot be the same as 0. The magnitude of JcCcWeight is determined as follows.

[0407] · When weight_index is 4 or less, JcCcWeight is equal to weight_index >> 2.

[0408] · In other cases, JcCcWeight is equal to 4 / (weight_index - 4).

[0409] The on / off control at the block level for ALF filtering of Cb and Cr is separate. This is the same as CCALF, and two individual sets of on / off control flags at the block level are encoded. Here, different from CCALF, the on / off control block sizes for Cb and Cr are the same, so only one block size variable is encoded.

[0410] [Loop Filter Section > Deblocking Filter]

[0411] In the deblocking filter process, the loop filter section 120 reduces the distortion generated at the block boundary by filtering the block boundary of the reconstructed image.

[0412] Figure 24 It is a block diagram showing an example of the detailed structure of the deblocking filter processing section 120a.

[0413] The deblocking filter processing section 120a includes, for example, a boundary determination section 1201, a filtering determination section 1203, a filtering processing section 1205, a processing determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0414] The boundary determination section 1201 determines whether there are pixels to be deblocked (i.e., target pixels) near the block boundary. Then, the boundary determination section 1201 outputs its determination result to the switch 1202 and the processing determination section 1208.

[0415] When it is determined by the boundary determination section 1201 that target pixels exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1204. On the contrary, when it is determined by the boundary determination section 1201 that target pixels do not exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1206. In addition, the image before filtering processing is an image composed of target pixels and at least one surrounding pixel located around the target pixel.

[0416] The filtering determination section 1203 determines whether to perform deblocking filter processing on the target pixels based on the pixel values of at least one surrounding pixel located around the target pixel. Then, the filtering determination section 1203 outputs the determination result to the switch 1204 and the processing determination section 1208.

[0417] When it is determined by the filtering determination section 1203 that deblocking filter processing is to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the filtering processing section 1205. On the contrary, when it is determined by the filtering determination section 1203 that deblocking filter processing is not to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the switch 1206.

[0418] When the image before filtering processing is obtained via the switches 1202 and 1204, the filtering processing section 1205 performs deblocking filter processing with the filtering characteristics determined by the filtering characteristic determination section 1207 on the target pixels. Then, the filtering processing section 1205 outputs the pixels after the filtering processing to the switch 1206.

[0419] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs pixels that have not been deblocking-filtered and pixels that have been deblocking-filtered by the filtering processing unit 1205.

[0420] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processing determination unit 1208 outputs the deblocking-filtered pixels from the switch 1206. In addition, in cases other than the above, the processing determination unit 1208 outputs the pixels that have not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206. In addition, Figure 24 The structure shown is an example of the structure in the deblocking filtering unit 120a, and the deblocking filtering unit 120a may have other structures.

[0421] Figure 25 is a diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to the block boundary.

[0422] In deblocking filtering, for example, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Figure 25 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.

[0423] q’0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0424] q’1 = (p0 + q0 + q1 + q2 + 2) / 4

[0425] q’2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0426] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering are filtering coefficients.

[0427] Furthermore, in the deblocking filter process, clipping processing may also be performed in such a way that the pixel value after operation does not change when it exceeds a threshold value. In this clipping processing, a threshold value determined according to the quantization parameter is used to clip the pixel value after operation based on the above formula to "the pixel value before operation ± 2 × the threshold value". Thereby, excessive smoothing can be prevented.

[0428] Figure 26 FIG. is an example of a block boundary for explaining the deblocking filter process. Figure 27 FIG. is an example showing the BS value.

[0429] The block boundary for the deblocking filter process is, for example, Figure 26 the boundary of a CU, PU, or TU of an 8×8 pixel block shown in the figure. The deblocking filter process is performed, for example, in units of 4 rows or 4 columns. First, for Figure 26 the blocks P and Q shown in the figure, the Bs (Boundary Strength) value is determined as Figure 27 shown.

[0430] Based on the Figure 27 Bs value, it can be determined whether to perform deblocking filter processing with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filter processing for the chrominance signal is performed. When the Bs value is 1 or more and a specified condition is satisfied, deblocking filter processing for the luminance signal is performed. In addition, the determination condition of the Bs value is not limited to the Figure 27 condition shown in the figure, and it can also be determined based on other parameters.

[0431] [Prediction unit (intra prediction unit / inter prediction unit / prediction control unit)]

[0432] Figure 28 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the encoding device 100. In addition, as an example, the prediction unit is composed of all or part of the constituent elements of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.

[0433] The prediction unit generates a prediction image of the current block (step Sb_1). In addition, in the prediction image, there are, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). Specifically, the prediction unit uses the reconstructed image that has already been obtained by generating a prediction image of other blocks, generating a prediction residual, generating quantization coefficients, restoring the prediction residual, and adding the prediction images to generate a prediction image of the current block.

[0434] The reconstructed image can be, for example, an image referring to a reference picture, or an image including the current block, i.e., the encoded blocks within the current picture (i.e., the other blocks described above). The encoded blocks within the current picture are, for example, the neighboring blocks of the current block.

[0435] Figure 29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding apparatus 100.

[0436] The prediction unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such prediction methods, the above-described reconstructed image can also be used.

[0437] Next, the prediction unit evaluates the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit evaluates these prediction images by calculating a cost C for each of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c and comparing the costs C of these prediction images. In addition, the cost C is calculated by an equation of an R-D optimization model, such as C = D + λ × R. In this equation, D is the encoding distortion of the prediction image, and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. Further, R is the bit rate of the stream. In addition, λ is, for example, the undetermined multiplier of Lagrange.

[0438] Next, the prediction unit selects one of the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects a method or mode for obtaining the final prediction image. For example, the prediction unit selects the prediction image with the minimum cost C based on the cost C calculated for these prediction images. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 can also be performed based on the parameters used in the encoding process. The encoding apparatus 100 can signal an information signal for determining the selected prediction image, method, or mode as a stream. This information can be, for example, a flag or the like. Thus, the decoding apparatus 200 can generate a prediction image based on this information in accordance with the method or mode selected in the encoding apparatus 100. In addition, in the Figure 29 example shown, after generating the prediction images by each method, the prediction unit selects any one of the prediction images. However, before generating these prediction images, the prediction unit can select a method or mode based on the parameters used in the above-described encoding process, and can generate a prediction image according to this method or mode.

[0439] For example, the first mode and the second mode are intra prediction and inter prediction, respectively, and the prediction unit may select a final prediction image for the current block from the prediction images generated according to these prediction modes.

[0440] Figure 30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device 100.

[0441] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a) and generates a prediction image by inter prediction (step Sd_1b). In addition, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0442] Next, the prediction unit evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above-mentioned cost C may also be used in this evaluation. Then, the prediction unit may select the prediction image that calculates the minimum cost C from the intra prediction image and the inter prediction image as the final prediction image for the current block (step Sd_3). That is, the prediction mode or pattern used to generate the prediction image for the current block is selected.

[0443] [Intra Prediction Unit]

[0444] The intra prediction unit 124 performs intra prediction (also referred to as intra-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction image (i.e., an intra prediction image) of the current block. Specifically, the intra prediction unit 124 generates an intra prediction image by referring to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control unit 128.

[0445] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0446] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC standard.

[0447] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may also include 32-direction prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 31This is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent 33 directions specified by the H.265 / HEVC standard, and the dashed arrows represent 32 additional directions (the 2 non-directional prediction modes are not shown in Figure 31 .

[0448] In various installation examples, in the intra prediction of chrominance blocks, the luminance blocks can also be referred to. That is, the chrominance components of the current block can also be predicted based on the luminance component of the current block. Such intra prediction is sometimes referred to as CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block in this way (for example, called the CCLM mode) can also be added as one of the intra prediction modes of the chrominance block.

[0449] The intra prediction unit 124 can also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intraprediction combination). The information indicating whether PDPC is used (for example, called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0450] Figure 32 This is a flowchart showing an example of the processing performed by the intra prediction unit 124.

[0451] The intra prediction unit 124 selects one intra prediction mode from multiple intra prediction modes (step Sw_1). Then, the intra prediction unit 124 generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM consists of, for example, 6 intra prediction modes. Two of the 6 intra prediction modes can be the Planar prediction mode and the DC prediction mode, and the remaining 4 modes can be directional prediction modes. Then, the intra prediction unit 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0452] Here, when it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). Further, the MPM flag set to 1 and the information indicating the intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110.

[0453] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction unit 124 does not set the MPM flag. Then, the intra prediction unit 124 generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). Further, the MPM flag set to 0 and the information indicating the intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110. The information indicating the intra prediction mode represents, for example, any value from 0 to 60.

[0454] [Inter - prediction unit]

[0455] The inter - prediction unit 126 performs inter - prediction (also referred to as inter - picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a predicted image (inter - prediction image). The inter - prediction is performed in units of the current block or a current sub - block within the current block. A sub - block is included in a block and is a unit smaller than the block. The size of the sub - block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub - block can also be switched in units of slices, bricks, or pictures, etc.

[0456] For example, the inter - prediction unit 126 performs motion search (motion estimation) within the reference picture for the current block or current sub - block to find the reference block or sub - block that most matches the current block or current sub - block. And, the inter - prediction unit 126 obtains motion information (such as a motion vector) for compensating the motion or change from the reference block or sub - block to the current block or sub - block. The inter - prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, thereby generating an inter - prediction image of the current block or sub - block. And, the inter - prediction unit 126 outputs the generated inter - prediction image to the prediction control unit 128.

[0457] The motion information used in motion compensation is signaled as an inter - prediction image in various forms. For example, a motion vector can also be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.

[0458] [List of reference pictures]

[0459] Figure 33 is a diagram showing an example of each reference picture, Figure 34 is a conceptual diagram showing an example of the list of reference pictures. The list of reference pictures is a list showing one or more reference pictures stored in the frame memory 122. In addition, in Figure 33 , the rectangle represents a picture, the arrow represents the reference relationship of the pictures, the horizontal axis represents time, I, P, and B in the rectangle respectively represent an intra-predicted picture, a single-predicted picture, and a bi-predicted picture, and the numbers in the rectangle represent the decoding order. As Figure 33 shown, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As Figure 34 shown, the list of reference pictures is a list showing candidates for reference pictures. For example, one picture (or slice) can have more than one list of reference pictures. For example, if the current picture is a single-predicted picture, one list of reference pictures is used, and if the current picture is a bi-predicted picture, two lists of reference pictures are used. In Figure 33 and Figure 34 's example, picture B3 as the current picture currPic has two lists of reference pictures, namely the L0 list and the L1 list. When the current picture currPic is picture B3, the candidates for the reference pictures of this current picture currPic are I0, P1, and B2, and each list of reference pictures (that is, the L0 list and the L1 list) represents these pictures. The inter-frame prediction unit 126 or the prediction control unit 128 specifies whether to actually refer to which picture in each list of reference pictures through the reference picture index refidxLx. In Figure 34 , the reference pictures P1 and B2 are specified through the reference picture indexes refIdxL0 and refIdxL1.

[0460] Such a list of reference pictures can be generated in units of sequence, picture, slice, block, CTU, or CU. In addition, the reference picture indexes of the reference pictures shown in the list of reference pictures that are referred to in inter-frame prediction can be encoded at the sequence level, picture level, slice level, block level, CTU level, or CU level. In addition, in multiple inter-frame prediction modes, a common list of reference pictures can also be used.

[0461] [Basic process of inter-frame prediction]

[0462] Figure 35 is a flowchart showing the basic process of inter-frame prediction.

[0463] The inter-frame prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the prediction image as a prediction residual (step Se_4).

[0464] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter-frame prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. In addition, in the candidate MV list, the previously derived MV can be added as a candidate MV. Further, in the derivation of the MV, the inter-frame prediction unit 126 may also determine the MV of the current block by further selecting at least one candidate MV from at least one candidate MV and determining the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. In addition, the action of searching the region of the reference picture may also be referred to as motion estimation.

[0465] In addition, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126. However, for example, the processing of step Se_1 or step Se_2 may also be performed by other components included in the encoding device 100.

[0466] In addition, a candidate MV list may be created for each process in each inter-frame prediction mode, or a common candidate MV list may be used in multiple inter-frame prediction modes. In addition, the processes of steps Se_3 and Se_4 respectively correspond to Figure 9 the processes of steps Sa_3 and Sa_4 shown. In addition, the process of step Se_3 corresponds to Figure 30 the process of step Sd_1b.

[0467] [Flow of MV Derivation]

[0468] Figure 36 is a flowchart showing an example of MV derivation.

[0469] The inter-frame prediction unit 126 may derive the MV of the current block in a mode of encoding motion information (such as MV). In this case, for example, the motion information may be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.

[0470] Alternatively, the inter-frame prediction unit 126 may derive an MV in a mode where motion information is not encoded. In this case, the motion information is not included in the stream.

[0471] Here, the modes of MV derivation include the following normal inter-frame mode, normal merge mode, FRUC mode, and affine mode. Among these modes, the modes in which motion information is encoded include the normal inter-frame mode, normal merge mode, and affine mode (specifically, affine inter-frame mode and affine merge mode). In addition, the motion information may include not only the MV but also the predicted MV selection information described later. In addition, the modes in which motion information is not encoded include the FRUC mode. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and derives the MV of the current block using the selected mode.

[0472] Figure 37 is a flowchart showing another example of MV derivation.

[0473] The inter-frame prediction unit 126 may derive the MV of the current block in a mode where the differential MV is encoded. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. The differential MV is the difference between the MV of the current block and its predicted MV. In addition, the predicted MV is a predicted motion vector.

[0474] Alternatively, the inter-frame prediction unit 126 may derive an MV in a mode where the differential MV is not encoded. In this case, the encoded differential MV is not included in the stream.

[0475] Here, as described above, the modes of MV derivation include the following normal inter-frame mode, normal merge mode, FRUC mode, and affine mode. Among these modes, the modes in which the differential MV is encoded include the normal inter-frame mode and the affine mode (specifically, the affine inter-frame mode). In addition, the modes in which the differential MV is not encoded include the FRUC mode, normal merge mode, and affine mode (specifically, the affine merge mode). The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes and derives the MV of the current block using the selected mode.

[0476] [Modes of MV Derivation]

[0477] Figure 38A and Figure 38B is an example of a diagram showing the classification of each mode of MV derivation. For example, as Figure 38AAs shown, according to whether motion information is encoded and whether differential MVs are encoded, the MV derivation modes are classified into three major modes. The three modes are the inter-frame mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter-frame mode is the mode for performing motion search, and it is the mode for encoding motion information and differential MVs. For example, as Figure 38B shown, the inter-frame mode includes the affine inter-frame mode and the normal inter-frame mode. The merge mode is the mode that does not perform motion search, and it is the mode for selecting an MV from the surrounding encoded blocks and using that MV to derive the MV of the current block. This merge mode is basically the mode for encoding motion information and not encoding differential MVs. For example, as Figure 38B shown, the merge mode includes the normal merge mode (sometimes also referred to as the usual merge mode or the regular merge mode), the MMVD (Merge with Motion Vector Difference) mode, the CIIP (Combined inter merge / intraprediction) mode, the triangular mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the various modes included in the merge mode, differential MVs are encoded exception. In addition, the above-mentioned affine merge mode and affine inter-frame mode are the modes included in the affine mode. The affine mode is the mode for assuming an affine transformation and deriving the MV of the current block by using the MVs of the multiple sub-blocks that make up the current block. The FRUC mode is the mode for deriving the MV of the current block by searching between encoded regions, and it is the mode for not encoding both motion information and differential MVs. In addition, the details of these various modes will be described later.

[0478] In addition, Figure 38A and Figure 38B shown, the classification of each mode is an example and is not limited thereto. For example, when differential MVs are encoded in the CIIP mode, this CIIP mode is classified into the inter-frame mode.

[0479] [MV Derivation > Normal Inter-frame Mode]

[0480] The normal inter-frame mode is an inter-frame prediction mode for deriving the MV of the current block by finding a block similar to the image of the current block from the region of the reference picture represented by the candidate MVs. In addition, in this normal inter-frame mode, differential MVs are encoded.

[0481] Figure 39 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.

[0482] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0483] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction MV candidates, respectively, in a predetermined order of priority (step Sg_2). In addition, this order of priority is predetermined for each of the N candidate MVs.

[0484] Next, the inter-frame prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the selected prediction MV into the stream. That is, the inter-frame prediction unit 126 outputs the prediction MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0485] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0486] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are executed for each block. For example, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. In addition, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. Furthermore, it may be that when the processes of steps Sg_1 to Sg_5 are not executed for all the blocks included in a slice but for some blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, it may be that when the processes of steps Sg_1 to Sg_5 are executed for some blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends.

[0487] In addition, the predicted image is the above-mentioned inter-frame prediction signal. Further, information indicating an inter-frame prediction mode (in the above example, the normal inter-frame mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.

[0488] In addition, the candidate MV list can also be commonly used with lists used in other modes. Further, processing related to the candidate MV list can be applied to processing related to lists used in other modes. Processing related to this candidate MV list is, for example, extracting or selecting a candidate MV from the candidate MV list, rearranging the candidate MVs, or deleting a candidate MV, etc.

[0489] [MV derivation > Normal merge mode]

[0490] The normal merge mode is an inter-frame prediction mode in which a candidate MV is selected from the candidate MV list as the MV of the current block to derive the MV. In addition, the normal merge mode is a narrow sense of the merge mode and is sometimes simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.

[0491] Figure 40 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode.

[0492] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of encoded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0493] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 encodes MV selection information for identifying the selected candidate MV into the stream. That is, the inter-frame prediction unit 126 outputs the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0494] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processes of steps Sh_1 to Sh_3 are performed for each block, for example. For example, when the processes of steps Sh_1 to Sh_3 are respectively performed for all the blocks included in a slice, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_1 to Sh_3 are respectively performed for all the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends. In addition, the processes of steps Sh_1 to Sh_3 may be such that when they are not performed for all the blocks included in a slice but for a part of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processes of steps Sh_1 to Sh_3 may be such that when they are performed for a part of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.

[0495] In addition, information indicating the inter-frame prediction mode (in the above example, the normal merge mode) used in the generation of the predicted image included in the stream is encoded as, for example, prediction parameters.

[0496] Figure 41 It is a diagram for explaining an example of the MV derivation process of the current picture based on the normal merge mode.

[0497] First, the inter-frame prediction unit 126 generates a candidate MV list in which candidate MVs are registered. As candidate MVs, there are: spatially adjacent candidate MVs, which are MVs possessed by a plurality of encoded blocks located in the spatial vicinity of the current block; temporally adjacent candidate MVs, which are MVs possessed by blocks in the vicinity where the position of the current block in the encoded reference picture is projected; combined candidate MVs, which are MVs generated by combining the MV values of the spatially adjacent candidate MVs and the temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero, etc.

[0498] Next, the inter-frame prediction unit 126 determines one candidate MV as the MV of the current block by selecting one candidate MV from the plurality of candidate MVs registered in the candidate MV list.

[0499] Moreover, in the entropy encoding unit 110, a signal indicating which candidate MV is selected, i.e., merge_idx, is described in the stream and encoded.

[0500] In addition, Figure 41 The candidate MVs registered in the candidate MV list described in are an example, and the number may be different from the number in the figure, or the structure may not include some types of candidate MVs in the figure, or the structure may be appended with candidate MVs other than the types of candidate MVs in the figure.

[0501] The MV of the current block exported through the normal merge mode can also be used to perform the subsequent DMVR (dynamic motion vector refreshing) to determine the final MV. In addition, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same way as the normal merge mode, but encodes the differential MV. As Figure 38B shown, such an MMVD can also be classified as a merge mode together with the normal merge mode. In addition, the differential MV in the MMVD mode may not be the same as the differential MV used in the inter-frame mode. For example, the derivation of the differential MV in the MMVD mode may also be a process with a smaller processing amount than the derivation of the differential MV in the inter-frame mode.

[0502] In addition, it is also possible to overlap the predicted image generated in the inter-frame prediction with the predicted image generated in the intra-frame prediction to perform the CIIP (Combined inter merge / intra prediction) mode for generating the predicted image of the current block.

[0503] In addition, the candidate MV list may also be referred to as a candidate list. In addition, merge_idx is MV selection information.

[0504] [MV Derivation > HMVP Mode]

[0505] Figure 42 is a diagram for explaining an example of the MV derivation process of the current picture based on the HMVP mode.

[0506] In the normal merge mode, one candidate MV is selected from the candidate MV list generated by referring to the encoded block (e.g., CU), thereby determining the MV of, for example, the CU of the current block. Here, other candidate MVs can also be registered in the candidate MV list. The mode of registering such other candidate MVs is called the HMVP mode.

[0507] In the HMVP mode, separately from the candidate MV list of the normal merge mode, a FIFO (First-In First-Out) buffer for HMVP is used to manage the candidate MVs.

[0508] In the FIFO buffer, motion information such as the MV of the block processed in the past is sequentially stored from the new FIFO buffer. In the management of this FIFO buffer, whenever one block is processed, the MV of the latest block (i.e., the immediately preceding processed CU) is stored in the FIFO buffer, and instead, the MV of the earliest CU (i.e., the CU processed first) in the FIFO buffer is deleted from the FIFO buffer. In Figure 42In the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the earliest block.

[0509] Then, for example, the inter-frame prediction unit 126 sequentially checks each MV managed in the FIFO buffer starting from HMVP1 to see if the MV is different from all the candidate MVs already registered in the candidate MV list in the ordinary merge mode. Also, when it is determined that the MV is different from all the candidate MVs, the inter-frame prediction unit 126 can add the MV managed in the FIFO buffer as a candidate MV to the candidate MV list in the ordinary merge mode. At this time, the candidate MV registered in the FIFO buffer can be one or more.

[0510] In this way, by using the HMVP mode, not only can the MVs of spatially or temporally adjacent blocks of the current block be added to the candidates, but also the MVs of the blocks processed in the past can be added to the candidates. As a result, by expanding the variation of the candidate MVs in the ordinary merge mode, the possibility of improving the coding efficiency becomes higher.

[0511] In addition, the above MV can also be motion information. That is, the information stored in the candidate MV list and the FIFO buffer can include not only the value of the MV, but also information such as the information indicating the reference picture, the reference direction, and the number of pictures. In addition, the above block is, for example, a CU.

[0512] In addition, Figure 42 The candidate MV list and the FIFO buffer are an example, and the candidate MV list and the FIFO buffer can also be lists or buffers of different sizes from Figure 42 or a structure in which the candidate MVs are registered in a different order from Figure 42 Here, the processing described here is common to both the encoding device 100 and the decoding device 200.

[0513] In addition, the HMVP mode can also be applied to modes other than the ordinary merge mode. For example, motion information such as the MVs of the blocks processed in the affine mode in the past can be sequentially saved from a new FIFO buffer and used as candidate MVs. The mode in which the HMVP mode is applied in the affine mode can also be called the historical affine mode.

[0514] [MV derivation>FRUC mode]

[0515] Motion information may also not be signaled from the encoding device 100 side, but derived on the decoding device 200 side. For example, motion information may also be derived by performing a motion search on the decoding device 200 side. In such a case, a motion search is performed on the decoding device 200 side without using the pixel values of the current block. Such a mode of performing a motion search on the decoding device 200 side includes a FRUC (frame rate up-conversion) mode or a PMMVD (pattern matched motion vector derivation) mode, etc.

[0516] Figure 43 An example of FRUC processing is shown. First, referring to the MVs of each encoded block adjacent to the current block in space or time, a list representing these MVs as candidate MVs is generated (i.e., it is a candidate MV list and may also be common to the candidate MV list in the normal merge mode) (step Si_1). Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected as the best candidate MV based on this evaluation value. And, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. In addition, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral area of the position in the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and evaluation values in the reference picture may also be performed on the peripheral area of the best candidate MV. In the case where there is an MV with a better evaluation value, the best candidate MV is updated to this MV and used as the final MV for the current block. The update to an MV with a better evaluation value may also not be implemented.

[0517] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_1 to Si_5 are performed on each block, for example. For example, when the processes of steps Si_1 to Si_5 are performed on all the blocks included in a slice respectively, the inter-frame prediction using the FRUC mode for that slice ends. Also, when the processes of steps Si_1 to Si_5 are performed on all the blocks included in a picture respectively, the inter-frame prediction using the FRUC mode for that picture ends. In addition, the processes of steps Si_1 to Si_5 may be such that when they are not performed on all the blocks included in a slice but on some of the blocks, the inter-frame prediction using the FRUC mode for that slice ends. Similarly, the processes of steps Si_1 to Si_5 may be such that when they are performed on some of the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends.

[0518] The same processing as that in the above block unit may also be performed in the case of processing in sub-block units.

[0519] The evaluation value may also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the MV is compared with the reconstructed image of a specified region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). Then, the difference in pixel values of the two reconstructed images may be calculated and used as the evaluation value for the MV. In addition, it may be that other information is used in addition to the difference value to calculate the evaluation value.

[0520] Next, the pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (also called the merge list) is selected as the starting point for the search based on pattern matching. As the pattern matching, the first pattern matching or the second pattern matching may be used. The first pattern matching and the second pattern matching may be cases called bilateral matching and template matching respectively.

[0521] [MV Derivation>FRUC>Bilateral Matching]

[0522] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate MV, the region in another reference picture along the motion trajectory of the current block is used.

[0523] Figure 44This is a diagram illustrating an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along a motion trajectory. As Figure 44 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two MVs (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.

[0524] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional MVs are derived.

[0525] [MV Derivation > FRUC > Template Matching]

[0526] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (the block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and the block in the reference picture. Thus, in the second pattern matching, the block adjacent to the current block in the current picture is used as the specified region for calculating the evaluation value of the above candidate MV.

[0527] Figure 45 This is a diagram illustrating an example of the pattern matching (template matching) between the template in the current picture and the block in the reference picture. As Figure 45 shown, in the second pattern matching, the MV of the current block is derived by searching for the block in the reference picture (Ref0) that most matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or one of the left adjacent and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.

[0528] Such information indicating whether the FRUC mode is adopted (for example, called a FRUC flag) is signaled at the CU level. In addition, when the FRUC mode is adopted (for example, when the FRUC flag is true), information indicating the method of style matching that can be adopted (first style matching or second style matching) is signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and can also be other levels (for example, sequence level, picture level, slice level, brick level, CTU level or sub-block level).

[0529] [MV Export > Affine Mode]

[0530] The affine mode is a mode for generating MV using affine transform, and for example, MV can be derived in sub-block units based on MVs of a plurality of adjacent blocks. This mode is sometimes referred to as affine motion compensation prediction mode.

[0531] Figure 46A FIG. 1 is a diagram for explaining an example of deriving an MV in sub-block units based on MVs of a plurality of adjacent blocks. Figure 46A In the example, the current block includes 16 sub-blocks composed of 4×4 pixels. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the MV of the adjacent sub-block. Then, according to the following formula (1A), the two motion vectors v0 and v1 are projected, and the motion vectors (v x , v y ).

[0532]

Formula 1

[0533]

[0534] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predetermined weight coefficient.

[0535] Information indicating such an affine mode (e.g., called an affine flag) may be signaled at the CU level. Furthermore, the signaling of the information indicating the affine mode need not be limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).

[0536] In addition, such an affine mode may include several modes in which the MVs of the upper left and upper right control points are derived in different ways. For example, the affine mode includes two modes: an affine inter (also called an affine normal inter) mode and an affine merge mode.

[0537] Figure 46B This is a diagram for explaining an example of the derivation of the MV of a sub-block unit in an affine mode using three control points. In Figure 46B , the current block includes 16 sub-blocks of 4×4 pixels. Here, the motion vector v0 of the upper-left control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper-right control point of the current block is derived based on the MV of the adjacent block, and the motion vector v2 of the lower-left control point of the current block is derived based on the MV of the adjacent block. Then, according to the following equation (1B), the three motion vectors v0, v1, and v2 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.

[0538]

Equation 2

[0539]

[0540] Here, x and y respectively represent the horizontal position and the vertical position of the sub-block center, and w and h represent pre-determined weight coefficients. Alternatively, w can represent the width of the current block, and h can represent the height of the current block.

[0541] Affine modes using different numbers of control points (e.g., two and three) can also be switched at the CU level and signaled. Additionally, information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0542] Furthermore, in such an affine mode with three control points, several modes with different derivation methods for the MVs of the upper-left, upper-right, and lower-left control points can be included. For example, in the affine mode with three control points, similar to the affine mode with two control points, there are two modes: the affine inter prediction mode and the affine merge mode.

[0543] Moreover, in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and can be other sizes. For example, the size of each sub-block can also be 8×8 pixels.

[0544] [MV Derivation>Affine Mode>Control Points]

[0545] Figure 47A , Figure 47B and Figure 47C are conceptual diagrams for explaining an example of the MV derivation of the control points in the affine mode.

[0546] In the affine mode, as Figure 47AAs shown, for example, based on a plurality of MVs corresponding to blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) that are coded in an affine mode and are adjacent to the current block, a predicted MV for each control point of the current block is calculated. Specifically, these blocks are checked in the order of coded block A (left), block B (above), block C (upper right), block D (lower left), and block E (upper left) to determine the first valid block coded in the affine mode. The MV of the control point of the current block is calculated based on the plurality of MVs corresponding to the determined block.

[0547] For example, as Figure 47B shown, in the case where block A adjacent to the left side of the current block is coded in an affine mode with 2 control points, motion vectors v3 and v4 projected onto positions of the upper left corner and the upper right corner of the coded block including block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are calculated.

[0548] For example, as Figure 47C shown, when block A adjacent to the left side of the current block is coded in an affine mode with 3 control points, motion vectors v3, v4, and v5 projected onto positions of the upper left corner, the upper right corner, and the lower left corner of the coded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block are calculated.

[0549] In addition, Figures 47A to 47C the method for deriving the MVs shown can be used for deriving the MVs of each control point of the current block in step Sk_1 described later, and can also be used for deriving the predicted MVs of each control point of the current block in step Sj_1 described later. Figure 50 shown, and can also be used for deriving the predicted MVs of each control point of the current block in step Sj_1 shown later. Figure 51

[0550] Figure 48A And Figure 48B are conceptual diagrams for explaining another example of the derivation of the control point MVs in the affine mode.

[0551] Figure 48A is a diagram for explaining the affine mode with 2 control points.

[0552] In this affine mode, as Figure 48AAs shown, the MV selected from the MVs of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block.

[0553] Figure 48B FIG. is for explaining an affine mode with three control points.

[0554] In this affine mode, as Figure 48B shown, the MV selected from the MVs of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block. In addition, the MV selected from the MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 of the lower left control point of the current block.

[0555] In addition, Figure 48A and Figure 48B the MV derivation method shown can be used for the derivation of the MV of each control point of the current block in step Sk_1 described later, and can also be used for the derivation of the predicted MV of each control point of the current block in step Sj_1 of [[ID=276 shown later. ​

[0556] Here, for example, in the case where an affine mode with different numbers of control points (e.g., 2 and 3) is signaled at the CU level, etc., the number of control points may be different depending on the encoded block and the current block.

[0557] ​ and ​ are conceptual diagrams for explaining an example of the MV derivation method of control points in the case where the number of control points is different between the encoded block and the current block.

[0558] For example, as ​ shown, the current block has three control points: upper left, upper right, and lower left, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points. In this case, the motion vectors v3 and v4 projected to the positions of the upper left and upper right corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block are calculated. Further, based on the derived motion vectors v0 and v1, the motion vector v2 of the lower left control point is calculated.

[0559] For example, as​ As shown in FIG. 1 , the current block has two control points, the upper left corner and the upper right corner, and the block A adjacent to the left of the current block is encoded in an affine mode with three control points. In this case, motion vectors v3, v4, and v5 projected to the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block including the block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are calculated.

[0560] in addition, ​ and ​ The MV derivation method shown can be used in the following ​ The derivation of the MV of each control point of the current block in step Sk_1 shown in FIG. 1 can also be used for the following ​ The predicted MV of each control point of the current block is derived in step Sj_1.

[0561] [MV Export > Affine Mode > Affine Merge Mode]

[0562] ​ This is a flowchart showing an example of the affine merge mode.

[0563] In the affine merge mode, first, the inter-frame prediction unit 126 derives the MV of each control point of the current block (step Sk_1). ​ As shown, it is the upper left and upper right corner points of the current block, or as ​ As shown, these are the points at the upper left corner, upper right corner, and lower left corner of the current block. At this time, the inter prediction unit 126 may also encode MV selection information for identifying the derived two or three MVs into the stream.

[0564] For example, when using ​ In the case of the MV derivation method shown, ​ As shown, the inter-frame prediction unit 126 checks the encoded blocks A (left), block B (top), block C (top right), block D (bottom left) and block E (top left) in this order, and determines the initial valid block encoded in the affine mode.

[0565] The inter-frame prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control point. For example, when block A is determined and block A has two control points, ​As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block containing block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block.

[0566] Alternatively, when block A is determined and block A has three control points, as ​ shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block.

[0567] In addition, as described above ​ shown, block A can be determined, and when block A has two control points, the MVs of three control points can be calculated. Also, as described above ​ shown, block A can be determined, and when block A has three control points, the MVs of two control points can be calculated.

[0568] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two motion vectors v0 and v1 and the above formula (1A), or uses three motion vectors v0, v1, and v2 and the above formula (1B) to calculate the MV of the sub-block as an affine MV (step Sk_2). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-block (step Sk_3). When the processes of steps Sk_2 and Sk_3 are respectively executed for all the sub-blocks included in the current block, the process of generating the prediction image using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and the prediction image of the current block is generated.

[0569] In addition, in step Sk_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can, for example, also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be ​ the MV derivation method shown ​ and​ The method for deriving the MV shown ​ and ​ the method for deriving the MV shown, and any combination of other methods for deriving MVs.

[0570] In addition, the candidate MV list may also include candidate MVs of a mode that performs prediction in units of sub-blocks other than the affine mode.

[0571] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode having two control points and candidate MVs of an affine merge mode having three control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode having two control points may be generated separately, and a candidate MV list including candidate MVs of an affine merge mode having three control points may be generated. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes having two control points and the affine merge mode having three control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or the MV of a valid block among these blocks.

[0572] In addition, as the MV selection information, an index indicating which candidate MV in the candidate MV list may be sent.

[0573] [MV Derivation>Affine Mode>Affine Inter-Frame Mode]

[0574] ​ is a flowchart showing an example of the affine inter-frame mode.

[0575] In the affine inter-frame mode, first, the inter-frame prediction unit 126 derives the predicted MVs (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_1). As ​ or ​ shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0576] For example, in the case of using the method for deriving the MV shown in Figure 48A and Figure 48B the inter-frame prediction unit 126 derives the predicted MVs (v0, v1) or (v0, v1, v2) of the control points of the current block by selecting the MV of a certain block among the encoded blocks near each control point of the current block shown in Figure 48A or Figure 48B At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the two or three selected predicted MVs into the stream.

[0577] For example, the inter-frame prediction unit 126 can determine which block's MV among the encoded blocks adjacent to the current block is to be used as the predicted MV for the control point by using cost evaluation or the like, and can describe a flag indicating which predicted MV is selected in the bitstream. That is, the inter-frame prediction unit 126 outputs prediction MV selection information such as a flag as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0578] Next, while updating the predicted MVs selected or derived in step Sj_1 respectively (step Sj_2), the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4). That is, the inter-frame prediction unit 126 uses the MV of each sub-block corresponding to the predicted MV to be updated as the affine MV, and calculates it using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference pictures (step Sj_4). Whenever the predicted MV is updated in step Sj_2, the processes of steps Sj_3 and Sj_4 are executed for all blocks within the current block. As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted MV that can obtain the minimum cost as the MV of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0579] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference pictures (step Sj_6).

[0580] In addition, in step Sj_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can be, for example, a list containing candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be Figures 47A - 47C the MV derivation methods shown, Figure 48A and Figure 48B the MV derivation methods shown, Figure 49A and Figure 49B the MV derivation methods shown, and any combination of other MV derivation methods.

[0581] In addition, the candidate MV list can also contain candidate MVs of prediction modes that are performed in units of sub-blocks other than the affine mode.

[0582] In addition, as a candidate MV list, a candidate MV list including candidate MVs of an affine inter-frame mode with two control points and candidate MVs of an affine inter-frame mode with three control points may also be generated. Alternatively, a candidate MV list including candidate MVs of an affine inter-frame mode with two control points and a candidate MV list including candidate MVs of an affine inter-frame mode with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of a mode of either an affine inter-frame mode with two control points or an affine inter-frame mode with three control points may be generated. The candidate MV may be, for example, the MVs of the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), or may be the MVs of the valid blocks among these blocks.

[0583] In addition, as prediction MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent out.

[0584] [MV Derivation > Triangle Mode]

[0585] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter-frame prediction unit 126 may generate a plurality of prediction images with shapes different from the rectangle for the rectangular current block, and generate a final rectangular prediction image by combining these plurality of prediction images. The shape different from the rectangle may also be a triangle, for example.

[0586] Figure 52A is a diagram for explaining the generation of two triangular prediction images.

[0587] The inter-frame prediction unit 126 performs motion compensation for the first partition of the triangle within the current block using the first MV of the first partition, thereby generating a triangular prediction image. Similarly, the inter-frame prediction unit 126 performs motion compensation for the second partition of the triangle within the current block using the second MV of the second partition, thereby generating a triangular prediction image. Then, the inter-frame prediction unit 126 combines these prediction images, thereby generating a rectangular prediction image identical to the current block.

[0588] In addition, as the prediction image of the first partition, a first rectangular prediction image corresponding to the current block may also be generated using the first MV. In addition, as the prediction image of the second partition, a second rectangular prediction image corresponding to the current block may also be generated using the second MV. The prediction image of the current block may also be generated by weighted addition of the first prediction image and the second prediction image. In addition, the region for weighted addition may also be only a partial region sandwiching the boundary between the first partition and the second partition.

[0589] Figure 52BIt is a conceptual diagram showing an example of a first part of a first partition that overlaps with a second partition, and a first sample set and a second sample set that can be weighted as part of a correction process. The first part can be, for example, one-fourth of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero. For example, N can be the integer 2. Figure 52B The example on the left shows a rectangular partition with a rectangular part having a width that is one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example in the center shows a rectangular partition with a rectangular part having a height that is one-fourth of the height of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example on the right shows a triangular partition with a polygonal part having a height corresponding to 2 samples. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.

[0590] The first part can be a part of the first partition that overlaps with an adjacent partition. Figure 52C It is a conceptual diagram showing a first part of a first partition that is a part of the first partition overlapping with a part of an adjacent partition. For simplicity of explanation, a rectangular partition having a part overlapping with a spatially adjacent rectangular partition is shown. Partitions having other shapes such as triangular partitions can be used, and the overlapping part can also overlap with a spatially or temporally adjacent partition.

[0591] In addition, an example of generating a prediction image for two partitions respectively using inter-frame prediction is shown, but an intra-frame prediction can also be used to generate a prediction image for at least one partition.

[0592] Figure 53 It is a flowchart showing an example of a triangular mode.

[0593] In the triangular mode, first, the inter-frame prediction unit 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame prediction unit 126 can encode information related to the division into each partition, that is, partition information, as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 can output the partition information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0594] Next, the inter-frame prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially located around the current block (step Sx_2). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0595] Then, the inter-frame prediction unit 126 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_1 (step Sx_3). At this time, the inter-frame prediction unit 126 may also encode the MV selection information for identifying the selected candidate MV as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 may output the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0596] Next, the inter-frame prediction unit 126 uses the selected first MV and the encoded reference picture to perform motion compensation, thereby generating a first prediction image (step Sx_4). Similarly, the inter-frame prediction unit 126 uses the selected second MV and the encoded reference picture to perform motion compensation, thereby generating a second prediction image (step Sx_5).

[0597] Finally, the inter-frame prediction unit 126 performs weighted addition on the first prediction image and the second prediction image, thereby generating a prediction image of the current block (step Sx_6).

[0598] In addition, in Figure 52A the example shown, the first partition and the second partition are respectively triangles, but they may also be trapezoids, or they may be respectively different shapes. Moreover, in Figure 52A the example shown, the current block is composed of 2 partitions, but it may also be composed of 3 or more partitions.

[0599] In addition, the first partition and the second partition may also be repeated. That is, the first partition and the second partition may also include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition may also be used to generate the prediction image of the current block.

[0600] In addition, in this example, an example in which prediction images are generated by inter-frame prediction in both 2 partitions is shown, but prediction images may also be generated by intra-frame prediction for at least 1 partition.

[0601] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different, or they may be the same candidate MV list.

[0602] In addition, the partition information may include an index indicating a splitting direction for splitting at least the current block into a plurality of partitions. The MV selection information may also include an index indicating the selected first MV and an index indicating the selected second MV. One index may also represent a plurality of pieces of information. For example, one index summarizing a part or the whole of the partition information and a part or the whole of the MV selection information may be encoded.

[0603] [MV Derivation > ATMVP Mode]

[0604] Figure 54 FIG. is an example of the ATMVP mode for deriving MVs in units of sub-blocks.

[0605] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, candidate MVs in units of sub-blocks are registered in the candidate MV list for the normal merge mode.

[0606] Specifically, in the ATMVP mode, first, as Figure 54 shown, in the encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, a temporal MV reference block corresponding to the current block is determined. Then, for each sub-block within the current block, the MV used for encoding the region corresponding to the sub-block within the temporal MV reference block is determined. The MVs thus determined are included in the candidate MV list as candidate MVs for the sub-blocks of the current block. In the case of selecting the candidate MVs for each sub-block from the candidate MV list, motion compensation is performed on the sub-block using the candidate MV as the MV of the sub-block. Thereby, a predicted image for each sub-block is generated.

[0607] In addition, in Figure 54 the example shown, as the peripheral MV reference block, the block adjacent to the lower left of the current block is used, but other blocks may also be used. In addition, the size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may also be switched in units such as slices, bricks, or pictures.

[0608] [Motion Search > DMVR]

[0609] Figure 55 FIG. is a diagram showing the relationship between the merge mode and DMVR.

[0610] The inter-frame prediction unit 126 derives the MV of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform an MV search, i.e., a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in the merge mode.

[0611] On the other hand, when it is determined in step Sl_1 to perform a motion search (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by the DMVR.

[0612] Figure 56 It is a conceptual diagram for explaining an example of the DMVR for determining the MV.

[0613] First, for example, in the merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are determined based on the encoded picture in the L0 list, i.e., the first reference picture (L0). Similarly, according to the candidate MV (L1), reference pixels are determined based on the encoded picture in the L1 list, i.e., the second reference picture (L1). A template is generated by taking the average of these reference pixels.

[0614] Next, using this template, the peripheral areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV of the current block. In addition, the cost can be calculated, for example, using the difference values between the pixel values of the template and the pixel values of the search area and the candidate MV values, etc.

[0615] Even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV to derive the final MV, any processing can be used.

[0616] Figure 57 It is a conceptual diagram for explaining another example of the DMVR for determining the MV. Figure 57 The example shown is different from Figure 56 an example of the DMVR shown, and does not generate a template but calculates the cost.

[0617] First, the inter-frame prediction unit 126 searches the periphery of the reference blocks included in the reference pictures of the L0 list and the L1 list based on the candidate MV, i.e., the initial MV, obtained from the candidate MV list. For example, as Figure 57As shown, the initial MV corresponding to the reference block of the L0 list is InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is InitMV_L1. In motion search, the inter-frame prediction unit 126 first sets the search position for the reference picture in the L0 list. The differential vector representing this set search position, specifically, the differential vector from the position represented by the initial MV (i.e., InitMV_L0) to this search position is MVd_L0. Then, the inter-frame prediction unit 126 determines the search position in the reference picture of the L1 list. This search position is represented by the differential vector from the position indicated by the initial MV (i.e., InitMV_L1) to this search position. Specifically, the inter-frame prediction unit 126 determines this differential vector as MVd_L1 by mirroring MVd_L0. That is, the inter-frame prediction unit 126 sets the position symmetric to the position represented by the initial MV as the search position in the reference pictures of the L0 list and the L1 list respectively. The inter-frame prediction unit 126 calculates, for each search position, the sum of the absolute differences (SAD) of the pixel values within the block at this search position as the cost, and finds the search position with the minimum cost.

[0618] Figure 58A FIG. is an example diagram showing the motion search in DMVR, Figure 58B is a flowchart showing an example of this motion search.

[0619] First, in Step1, the inter-frame prediction unit 126 calculates the cost of the search position (also called the start point) represented by the initial MV and the 8 search positions around it. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than the start point is the minimum. Here, when it is determined that the cost of the search positions other than the start point is the minimum, the inter-frame prediction unit 126 moves to the search position with the minimum cost and performs the processing of Step2. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 126 skips the processing of Step2 and performs the processing of Step3.

[0620] In Step2, the inter-frame prediction unit 126 uses the search position moved according to the processing result of Step1 as the new start point and performs the same search as the processing of Step1. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than this start point is the minimum. Here, if the cost of the search positions other than the start point is the minimum, the inter-frame prediction unit 126 performs the processing of Step4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 126 performs the processing of Step3.

[0621] In Step4, the inter-frame prediction unit 126 processes the search position of this start point as the final search position, and determines the difference between the position represented by the initial MV and this final search position as the differential vector.

[0622] In Step 3, the inter-frame prediction unit 126 determines the pixel position with sub-pixel precision having the minimum cost based on the costs at four points above, below, left, and right of the start point in Step 1 or Step 2, and sets this pixel position as the final search position. This pixel position with sub-pixel precision is determined by weighted summing the vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) at the four points above, below, left, and right, with the costs at the respective search positions of the four points as weights. Then, the inter-frame prediction unit 126 determines the difference between the position indicated by the initial MV and this final search position as the difference vector.

[0623] [Motion Compensation > BIO / OBMC / LIC]

[0624] In motion compensation, there is a mode of generating a predicted image and correcting this predicted image. This mode is, for example, BIO, OBMC, and LIC described later.

[0625] Figure 59 It is a flowchart showing an example of the generation of a predicted image.

[0626] The inter-frame prediction unit 126 generates a predicted image (step Sm_1), and corrects this predicted image by any of the above modes (step Sm_2).

[0627] Figure 60 It is a flowchart showing another example of the generation of a predicted image.

[0628] The inter-frame prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image using this MV (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform a correction process (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting this predicted image (step Sn_4). Also, in LIC described later, it is also possible to correct luminance and chrominance in step Sn_4. On the other hand, when it is determined not to perform a correction process (No in step Sn_3), the inter-frame prediction unit 126 outputs this predicted image as the final predicted image without correcting this predicted image (step Sn_5).

[0629] [Motion Compensation > OBMC]

[0630] Not only the motion information of the current block obtained through motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame predicted image. Specifically, an inter-frame predicted image can also be generated in units of sub-blocks within the current block by performing weighted addition of a predicted image based on the motion information obtained through motion search (within the reference picture) and a predicted image based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or the OBMC mode.

[0631] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Additionally, the level at which these pieces of information are signaled need not be limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0632] A more specific description of the OBMC mode will be given. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC.

[0633] First, as Figure 62 shown, using the MV assigned to the current block, a predicted image (Pred) based on normal motion compensation is obtained. In Figure 62 , the arrow "MV" points to the reference picture and indicates which block in the reference picture the current block in the current picture refers to for obtaining the predicted image.

[0634] Next, the MV (MV_L) that has been derived for the already encoded left adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two predicted images Pred and Pred_L, the first correction of the predicted image is performed. This has the effect of blending the boundaries between adjacent blocks.

[0635] Similarly, the Motion Vector (MV_U) that has been derived for the already - coded upper - adjacent block is applied (re - used) to the current block to obtain a predicted image (Pred_U). The MV (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second - stage correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted images that have undergone the first - stage correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second - stage correction is the final predicted image of the current block where the boundaries with adjacent blocks are blended (smoothed).

[0636] In addition, in the above example, a two - path correction method using the left - adjacent and upper - adjacent blocks is used, but this correction method can also be a three - path or more - path correction method that also uses the right - adjacent and / or lower - adjacent blocks.

[0637] In addition, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.

[0638] In addition, the prediction - image correction process of OBMC has been described here. The prediction - image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can be applied to each of the multiple reference pictures. In this case, after obtaining the corrected predicted images from each reference picture through the OBMC - based image correction for multiple reference pictures, the final predicted image is obtained by further overlapping the multiple obtained corrected predicted images.

[0639] In addition, in OBMC, the unit of the current block can be the PU unit or a sub - block unit obtained by further dividing the PU.

[0640] As a method for determining whether to apply OBMC, for example, there is a method of using a signal indicating whether to apply OBMC, namely obmc_flag. As a specific example, the encoding device 100 can also determine whether the current block belongs to a region with complex motion. When the encoding device 100 determines that the current block belongs to a region with complex motion, it sets the obmc_flag value to 1 and applies OBMC for encoding. When it does not belong to a region with complex motion, it sets the obmc_flag value to 0 and encodes the block without applying OBMC. On the other hand, in the decoding device 200, by decoding the obmc_flag described in the bitstream, it switches whether to apply OBMC for decoding according to this value.

[0641] [Motion Compensation > BIO]

[0642] Next, a method for deriving an MV will be described. First, a mode for deriving an MV based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi - directional optical flow) mode. Additionally, the bi - directional optical flow can also be expressed as BDOF instead of BIO.

[0643] Figure 63 is a diagram for explaining a model assuming uniform linear motion. In Figure 63 , (v x , v y ) represents the velocity vector, τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.

[0644] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) holds.

[0645]

Equation 3

[0646]

[0647] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a candidate MV list or the like can be corrected in pixel units.

[0648] In addition, an MV can also be derived on the decoder device 200 side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, a motion vector can also be derived in sub - block units based on the MVs of multiple adjacent blocks.

[0649] Figure 64 is a flowchart showing an example of inter - frame prediction according to BIO. Additionally, Figure 65This is a diagram showing an example of the functional structure of the inter-frame prediction unit 126 that performs inter-frame prediction in accordance with BIO.

[0650] As Figure 65 shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a predicted image correction unit 126f. Additionally, the memory 126a may be the frame memory 122.

[0651] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 126 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_1). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0652] Next, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . Additionally, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_2). Here, the interpolation image I 0 is the image included in the reference picture Ref0 derived for the current block, and the interpolation image I 1 is the image included in the reference picture Ref1 derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Alternatively, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each be an image larger than the current block. Additionally, the interpolation image I 0 and I 1 can include the predicted image derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1), as well as the motion compensation filter.

[0653] Additionally, the gradient image derivation unit 126c derives the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 , Iy 0 , Iy 1)(Step Sy_3). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivation unit 126c can also derive the gradient image by applying a gradient filter to the interpolation image, for example. The gradient image only needs to represent the spatial change amount of the pixel values along the horizontal direction or the vertical direction.

[0654] Next, the optical flow derivation unit 126d derives the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block, using the interpolation image (I 0 , I 1 ) and the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (Step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and can also be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can be a 4x4 pixel sub-CU. In addition, the derivation of the optical flow can also be performed in other units such as pixel units instead of sub-block units.

[0655] Next, the inter-frame prediction unit 126 corrects the prediction image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives the correction value of the values of the pixels included in the current block using the optical flow (vx, vy) (Step Sy_5). Moreover, the prediction image correction unit 126f can also correct the prediction image of the current block using the correction value (Step Sy_6). In addition, the correction value can be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.

[0656] In addition, the processing flow of BIO is not limited to Figure 64 the processing disclosed. It can perform only a part of the processing disclosed in Figure 64 , or can add or replace different processing, or can execute in a different processing order.

[0657] [Motion Compensation > LIC]

[0658] Next, an example of a mode for generating a prediction image (prediction) using LIC (local illumination compensation) will be described.

[0659] Figure 66A is a diagram for explaining an example of a prediction image generation method using a luminance correction process based on LIC. In addition, Figure 66BIt is a flowchart showing an example of a prediction image generation method using this LIC.

[0660] First, the inter-frame prediction unit 126 derives the MV from the encoded reference picture and obtains the reference image corresponding to the current block (step Sz_1).

[0661] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values change in the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. Then, the inter-frame prediction unit 126 calculates the luminance correction parameter using the information indicating how the luminance values change (step Sz_3).

[0662] The inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image in the reference picture specified by the MV by applying this luminance correction parameter (step Sz_4). That is, the prediction image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the color difference can be corrected. That is, information indicating how the color difference changes can also be used to calculate the correction parameter for the color difference and perform the color difference correction process.

[0663] In addition, Figure 66A The shape of the peripheral reference region in is an example, and shapes other than this can also be used.

[0664] Furthermore, the process of generating a prediction image based on one reference picture is described here, but the same applies when generating a prediction image based on multiple reference pictures. A prediction image can also be generated after performing a luminance correction process on the reference images obtained from each reference picture in the same manner as described above.

[0665] As a method for determining whether to apply the LIC, for example, there is a method of using the lic_flag as a signal indicating whether to apply the LIC. As a specific example, in the encoding device 100, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and encoding is performed by applying the LIC. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC. On the other hand, in the decoding device 200, it is also possible to decode the lic_flag described in the stream and switch whether to apply the LIC according to its value for decoding.

[0666] As other methods for determining whether to apply LIC, for example, there is also a method of determining based on whether LIC is applied to neighboring blocks. As a specific example, when the current block is processed in the merge mode, the inter-frame prediction unit 126 determines whether the neighboring encoded blocks selected when deriving the MV in the merge mode are encoded with LIC applied. Based on the result, the inter-frame prediction unit 126 switches whether to apply LIC for encoding. Additionally, in the case of this example, the same processing also applies to the decoding device 200 side.

[0667] Use Figure 66A And Figure 66B LIC (luminance correction processing) has been described. Hereinafter, its detailed content will be described.

[0668] First, the inter-frame prediction unit 126 derives an MV for obtaining a reference image corresponding to the current block from the reference picture that is an encoded picture.

[0669] Next, for the current block, the inter-frame prediction unit 126 uses the luminance pixel values of the left neighboring and upper neighboring encoded peripheral reference regions and the luminance pixel values at the same positions in the reference picture specified by the MV to extract information indicating how the luminance values change between the reference picture and the current picture, and calculates a luminance correction parameter. For example, let the luminance pixel value of a certain pixel in the peripheral reference region within the current picture be p0, and let the luminance pixel value of the pixel at the same position in the peripheral reference region within the reference picture be p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1 + B = p0 as the luminance correction parameter for multiple pixels in the peripheral reference region.

[0670] Next, the inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image in the reference picture specified by the MV using the luminance correction parameter. For example, let the luminance pixel value in the reference image be p2, and let the luminance pixel value of the prediction image after the luminance correction process be p3. The inter-frame prediction unit 126 generates a prediction image after the luminance correction process by calculating A×p2 + B = p3 for each pixel in the reference image.

[0671] In addition, it is also possible to use Figure 66A a part of the peripheral reference region shown. For example, a region including a specified number of pixels respectively spaced and removed from the upper neighboring pixel and the left neighboring pixel can be used as the peripheral reference region. Additionally, the peripheral reference region is not limited to the region adjacent to the current block, and can also be a region not adjacent to the current block. Additionally, in Figure 66AIn the example shown, the peripheral reference region in the reference picture is the region specified by the MV of the current picture from the peripheral reference region in the current picture, but it can also be the region specified by other MVs. For example, the other MV can also be the MV of the peripheral reference region in the current picture.

[0672] In addition, here, the operation in the encoding device 100 has been described, but the operation in the decoding device 200 is the same.

[0673] Furthermore, LIC can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0674] In addition, LIC can also be applied in units of sub-blocks. For example, correction parameters can be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0675] [Prediction control unit]

[0676] The prediction control unit 128 selects one of the intra-predicted image (pixels or signals output from the intra-prediction unit 124) and the inter-predicted image (pixels or signals output from the inter-prediction unit 126), and outputs the selected predicted image to the subtraction unit 104 and the addition unit 116.

[0677] [Prediction parameter generation unit]

[0678] The prediction parameter generation unit 130 can output information related to the selection of the predicted image in intra-prediction, inter-prediction, and the prediction control unit 128 as prediction parameters to the entropy encoding unit 110. The entropy encoding unit 110 can generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream, and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selected prediction signal (e.g., MV, prediction type, or the prediction mode used by the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0679] [Decoding device]

[0680] Next, a decoding device 200 that can decode the stream output from the above encoding device 100 will be described. Figure 67It is a block diagram showing an example of the functional structure of the decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes an encoded image, i.e., a stream, in units of blocks.

[0681] As Figure 67 shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. In addition, the intra prediction unit 216 and the inter prediction unit 218 are each configured as a part of the prediction processing unit.

[0682] [Installation example of decoding device]

[0683] Figure 68 It is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 67 as shown, the multiple components of the decoding device 200 are implemented by Figure 68 the shown processor b1 and memory b2.

[0684] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor b1 can also play the role of Figure 67 multiple components of the decoding device 200 as shown, excluding the components for storing information.

[0685] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode a stream. The memory b2 can be either an electronic circuit or connected to the processor b1. In addition, the memory b2 can also be included in the processor b1. In addition, the memory b2 can also be an aggregate of multiple electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disk, etc., and can also be represented as a storage or a recording medium, etc. In addition, the memory b2 can be either a non-volatile memory or a volatile memory.

[0686] For example, the memory b2 can store an image or a stream. In addition, a program for the processor b1 to decode a stream can also be stored in the memory b2.

[0687] In addition, for example, the memory b2 can also play the role of Figure 67The functions of the components for storing information among the multiple components of the decoding device 200 shown, etc. Specifically, the memory b2 can function as Figure 67 the block memory 210 and the frame memory 214 shown. More specifically, reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.) can be stored in the memory b2.

[0688] In addition, in the decoding device 200, all of the multiple components shown, etc. may not be installed, Figure 67 nor may all of the above-mentioned multiple processes be performed. Figure 67 A part of the multiple components shown, etc. may be included in other devices, or a part of the above-mentioned multiple processes may be performed by other devices.

[0689] Hereinafter, after explaining the overall processing flow of the decoding device 200, each component included in the decoding device 200 will be described. In addition, for the components that perform the same processing as the components included in the encoding device 100 among the components included in the decoding device 200, detailed descriptions will be omitted. For example, the inverse quantization unit 204, inverse transformation unit 206, addition unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 included in the decoding device 200 respectively perform the same processing as the inverse quantization unit 112, inverse transformation unit 114, addition unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120 included in the encoding device 100.

[0690] [Overall Flow of Decoding Process]

[0691] Figure 69 is a flowchart showing an example of the overall decoding process performed by the decoding device 200.

[0692] First, the segmentation determination unit 224 of the decoding device 200 determines the segmentation style of each of the multiple fixed-size blocks (128×128 pixels) included in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This segmentation style is the segmentation style selected by the encoding device 100. Then, the decoding device 200 performs the processes of steps Sp_2 to Sp_6 on each of the multiple blocks constituting this segmentation style.

[0693] The entropy decoding unit 202 decodes the encoded quantization coefficients and prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).

[0694] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore the prediction residual of the current block by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (step Sp_3).

[0695] Next, the prediction processing unit including the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction image of the current block (step Sp_4).

[0696] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction residual to the prediction image (step Sp_5).

[0697] Moreover, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).

[0698] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeated.

[0699] In addition, the processing of these steps Sp_1 to Sp_7 can be sequentially performed by the decoding device 200. Multiple processes of a part of these processes can be performed in parallel, or the order can be changed.

[0700] [Segmentation Decision Unit]

[0701] Figure 70 It is a diagram showing the relationship between the segmentation decision unit 224 and other components. As an example, the segmentation decision unit 224 can also perform the following processing.

[0702] The segmentation decision unit 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoding unit 202. Moreover, the segmentation decision unit 224 can determine the segmentation pattern of a fixed-size block based on the block information and the parameters. And the segmentation decision unit 224 can also output the information indicating the determined segmentation pattern to the inverse transformation unit 206, the intra prediction unit 216, and the inter prediction unit 218. The inverse transformation unit 206 can perform inverse transformation on the transformation coefficients based on the segmentation pattern indicated by the information from the segmentation decision unit 224. The intra prediction unit 216 and the inter prediction unit 218 can generate a prediction image based on the segmentation pattern indicated by the information from the segmentation decision unit 224.

[0703] [Entropy Decoding Unit]

[0704] Figure 71 It is a block diagram showing an example of the functional structure of the entropy decoding unit 202.

[0705] The entropy decoding unit 202 performs entropy decoding on the stream to generate quantized coefficients, prediction parameters, and parameters related to the segmentation pattern, etc. For example, CABAC is used in this entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a de - binarization unit 202c. The binary arithmetic decoding unit 202a uses the context value derived by the context control unit 202b to perform arithmetic decoding on the stream as a binary signal. Similar to the context control unit 110b of the encoding device 100, the context control unit 202b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. The de - binarization unit 202c performs de - binarization to transform the binary signal output from the binary arithmetic decoding unit 202a into a multi - value signal representing the above - mentioned quantized coefficients, etc. This de - binarization is performed in the same manner as the above - mentioned binarization.

[0706] The entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the prediction parameters included in the stream (refer to Figure 1 ) to the intra - prediction unit 216, the inter - prediction unit 218, and the prediction control unit 220. The intra - prediction unit 216, the inter - prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra - prediction unit 124, the inter - prediction unit 126, and the prediction control unit 128 on the encoding device 100 side.

[0707] [Entropy Decoding Unit]

[0708] Figure 72 is a diagram showing the process of CABAC in the entropy decoding unit 202.

[0709] First, in the CABAC in the entropy decoding unit 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding unit 202c and setting of the initial context value are performed. Then, the binary arithmetic decoding unit 202c and the de - binarization unit 202c perform arithmetic decoding and de - binarization on the encoded data of the CTU, for example. At this time, the context control unit 202b updates the context value each time arithmetic decoding is performed. Then, the context control unit 202b saves the context value as post - processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.

[0710] [Inverse Quantization Unit]

[0711] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the current block that are input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Further, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transformation unit 206.

[0712] Figure 73 FIG. is a block diagram showing an example of the functional configuration of the inverse quantization unit 204.

[0713] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204a, a predicted quantization parameter generation unit 204b, a quantization parameter storage unit 204d, and an inverse quantization processing unit 204e.

[0714] Figure 74 FIG. is a flowchart showing an example of the inverse quantization performed by the inverse quantization unit 204.

[0715] As an example, the inverse quantization unit 204 may perform inverse quantization processing for each CU based on the Figure 74 procedure shown. Specifically, the quantization parameter generation unit 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation unit 204a obtains the differential quantization parameter of the current block from the entropy decoding unit 202 (step Sv_12).

[0716] Next, the predicted quantization parameter generation unit 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). The predicted quantization parameter generation unit 204b generates a predicted quantization parameter for the current block based on the obtained quantization parameter (step Sv_14).

[0717] Then, the quantization parameter generation unit 204a adds the differential quantization parameter of the current block obtained from the entropy decoding unit 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generation unit 204b (step Sv_15). By this addition, the quantization parameter of the current block is generated. Further, the quantization parameter generation unit 204a stores the quantization parameter of the current block in the quantization parameter storage unit 204d (step Sv_16).

[0718] Next, the inverse quantization processing unit 204e inverse-quantizes the quantization coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).

[0719] In addition, the differential quantization parameter can also be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can also be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.

[0720] In addition, the inverse quantization unit 204 may include a plurality of inverse quantizers, and may also inverse-quantize the quantization coefficients using an inverse quantization method selected from a plurality of inverse quantization methods.

[0721] [Inverse transform unit]

[0722] The inverse transform unit 206 restores the prediction residual by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.

[0723] For example, when the information read from the stream indicates the application of EMT or AMT (for example, the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the transform type read.

[0724] In addition, for example, when the information read from the stream indicates the application of NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients.

[0725] Figure 75 It is a flowchart showing an example of the process performed by the inverse transform unit 206.

[0726] For example, the inverse transform unit 206 determines whether there is information indicating that orthogonal transformation is not performed in the stream (step St_11). Here, when it is determined that there is no such information (No in step St_11), the inverse transform unit 206 obtains the information indicating the transform type that has been decoded by the entropy decoding unit 202 (step St_12). Then, the inverse transform unit 206 determines the transform type used in the orthogonal transformation of the encoding device 100 based on this information (step St_13). And the inverse transform unit 206 performs inverse orthogonal transformation using the determined transform type (step St_14).

[0727] Figure 76 It is a flowchart showing another example of the process performed by the inverse transform unit 206.

[0728] For example, the inverse transform unit 206 determines whether the transform size is equal to or less than a specified value (step Su_11). Here, when it is determined that the size is equal to or less than the specified value (Yes in step Su_11), the inverse transform unit 206 obtains from the entropy decoding unit 202 information indicating which one of one or more transform types included in the first transform type group is used by the encoding device 100 (step Su_12). Further, such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.

[0729] Based on this information, the inverse transform unit 206 determines the transform type to be used in the orthogonal transform in the encoding device 100 (step Su_13). Then, the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, when it is determined in step Su_11 that the transform size is not equal to or less than the specified value (No in step Su_11), the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the second transform type group (step Su_15).

[0730] In addition, as an example, the inverse orthogonal transform performed by the inverse transform unit 206 can be implemented for each TU according to the Figure 75 or Figure 76 shown process. Further, instead of decoding the information indicating the transform type used in the orthogonal transform, an inverse orthogonal transform may be performed using a predefined transform type. Specifically, the transform type is DST7 or DCT8, etc., and in the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.

[0731] [Addition unit]

[0732] The addition unit 208 reconstructs the current block by adding the prediction residual, which is an input from the inverse transform unit 206, to the prediction image, which is an input from the prediction control unit 220. That is, a reconstructed image of the current block is generated. Then, the addition unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.

[0733] [Block memory]

[0734] The block memory 210 is a storage unit for storing blocks within the current picture that are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the addition unit 208.

[0735] [Loop filter unit]

[0736] The loop filter unit 212 applies loop filtering to the reconstructed image generated by the addition unit 208, and outputs the filtered reconstructed image to the frame memory 214, the display device, and the like.

[0737] When the information indicating the ON / OFF of ALF read from the stream indicates that ALF is ON, one filter is selected from among a plurality of filters based on the direction and activity of the locality-based gradient, and the selected filter is applied to the reconstructed image.

[0738] Figure 77 FIG. is a block diagram showing an example of the functional configuration of the loop filter unit 212. In addition, the loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding apparatus 100.

[0739] The loop filter unit 212 is, for example, as Figure 77 shown, includes a deblocking filter processing unit 212a, an SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 212b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. In addition, the loop filter unit 212 may not include Figure 77 all of the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 212 may also be configured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 77 .

[0740] [Frame Memory]

[0741] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed image that has been filtered by the loop filter unit 212.

[0742] [Prediction Unit (Intra-Frame Prediction Unit / Inter-Frame Prediction Unit / Prediction Control Unit)]

[0743] Figure 78 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the decoding apparatus 200. In addition, as an example, the prediction unit is composed of all or a part of the constituent elements of an intra-frame prediction unit 216, an inter-frame prediction unit 218, and a prediction control unit 220. The prediction processing unit includes, for example, an intra-frame prediction unit 216 and an inter-frame prediction unit 218.

[0744] The prediction unit generates a prediction image of the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. Additionally, among the prediction signals, there are, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction unit generates a prediction image of the current block using the reconstructed image that has already been obtained by generating a prediction image of other blocks, restoring the prediction residual, and adding the prediction images. The prediction unit of the decoding device 200 generates the same prediction image as the prediction image generated by the prediction unit of the encoding device 100. That is, the methods for generating the prediction images used in these prediction units are mutually common or corresponding.

[0745] The reconstructed image can be, for example, an image of a reference picture, or can be an image of a decoded block within the picture containing the current block, that is, the current picture (i.e., the above-mentioned other blocks). The decoded blocks within the current picture are, for example, adjacent blocks of the current block.

[0746] Figure 79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device 200.

[0747] The prediction unit determines the method or mode for generating the prediction image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.

[0748] When it is determined that the first method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the first method (step Sr_2a). Additionally, when it is determined that the second method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the second method (step Sr_2b). Additionally, when it is determined that the third method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the third method (step Sr_2c).

[0749] The first method, the second method, and the third method are different methods for generating the prediction image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.

[0750] Figure 80A and Figure 80B It is a flowchart showing another example of the processing performed by the prediction unit in the decoding device 200.

[0751] As an example, the prediction unit can also perform the prediction process according to the Figure 80A and Figure 80B shown flow. Additionally, Figure 80A and Figure 80BThe intra block copy shown belongs to one mode of inter prediction, which is a mode in which a block included in the current picture is referred to as a reference picture or a reference block. That is, in the intra block copy, a picture different from the current picture is not referred to. In addition, Figure 80A The PCM mode shown belongs to one mode of intra prediction, which is a mode without transformation and quantization.

[0752] [Intra prediction unit]

[0753] The intra prediction unit 216 performs intra prediction by referring to a block in the current picture stored in the block memory 210 based on the intra prediction mode read from the stream, thereby generating a predicted picture (i.e., an intra prediction picture) of the current block. Specifically, the intra prediction unit 216 generates an intra prediction picture by referring to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction picture to the prediction control unit 220.

[0754] In addition, when an intra prediction mode referring to a luminance block is selected in the intra prediction of a chrominance difference block, the intra prediction unit 216 may also predict the chrominance difference component of the current block based on the luminance component of the current block.

[0755] Furthermore, when the information read from the stream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel values after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0756] Figure 81 It is a diagram showing an example of the processing performed by the intra prediction unit 216 of the decoding device 200.

[0757] The intra prediction unit 216 first determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 exists (Yes in step Sw_11), the intra prediction unit 216 obtains information indicating the intra prediction mode selected in the encoding device 100 from the entropy decoding unit 202 (step Sw_12). In addition, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Next, the intra prediction unit 216 determines the MPM (step Sw_13). The MPM is composed of, for example, six intra prediction modes. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information obtained in step Sw_12 from among the multiple intra prediction modes included in this MPM (step Sw_14).

[0758] On the other hand, when it is determined in step Sw_11 that there is no MPM flag indicating 1 in the stream (No in step Sw_11), the intra prediction unit 216 acquires information indicating the intra prediction mode selected in the encoding device 100 (step Sw_15). That is, the intra prediction unit 216 acquires from the entropy decoding unit 202 information indicating the intra prediction mode selected in the encoding device 100 among one or more intra prediction modes not included in the MPM. In addition, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information acquired in step Sw_15 from among the one or more intra prediction modes not included in the MPM (step Sw_17).

[0759] The intra prediction unit 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).

[0760] [Inter prediction unit]

[0761] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks within the current block. In addition, a sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can also be switched in units of slices, bricks, or pictures, etc.

[0762] For example, the inter prediction unit 218 performs motion compensation using motion information (e.g., MV) read from the stream (e.g., prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction image of the current block or sub-block, and outputting the inter prediction image to the prediction control unit 220.

[0763] When the information read from the stream indicates that the OBMC mode is applied, the inter prediction unit 218 uses not only the motion information of the current block obtained through motion search but also the motion information of adjacent blocks to generate an inter prediction image.

[0764] In addition, when the information read from the stream indicates that the FRUC mode is applied, the inter prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) read from the stream, thereby deriving motion information. And the inter prediction unit 218 uses the derived motion information for motion compensation (prediction).

[0765] In addition, when the inter-frame prediction unit 218 applies the BIO mode, it derives the MV based on a model assuming uniform linear motion. In addition, when the information decoded from the stream indicates the application of the affine mode, the inter-frame prediction unit 218 derives the MV in sub-block units based on the MVs of multiple adjacent blocks.

[0766] [Flow of MV derivation]

[0767] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device 200.

[0768] The inter-frame prediction unit 218 determines, for example, whether to decode motion information (e.g., MV). For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the motion information, the inter-frame prediction unit 218 derives the MV of the current block in the mode of decoding the motion information. On the other hand, when it is determined not to decode the motion information, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the motion information.

[0769] Here, the MV derivation modes include the ordinary inter-frame mode, ordinary merge mode, FRUC mode, and affine mode, etc., which will be described later. Among these modes, the modes of decoding motion information include the ordinary inter-frame mode, ordinary merge mode, and affine mode (specifically, affine inter-frame mode and affine merge mode), etc. In addition, the motion information can include not only the MV, but also the predicted MV selection information described later. In addition, the mode of not decoding motion information includes the FRUC mode, etc. The inter-frame prediction unit 218 selects the mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.

[0770] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device 200.

[0771] The inter-frame prediction unit 218 determines, for example, whether to decode the differential MV. For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the differential MV, the inter-frame prediction unit 218 can derive the MV of the current block in the mode of decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.

[0772] On the other hand, when it is determined not to decode the differential MV, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the differential MV. In this case, the encoded differential MV is not included in the stream.

[0773] Here, as described above, the export modes of the MV include the following normal inter-frame mode, normal merge mode, FRUC mode, and affine mode. Among these modes, the modes for encoding the differential MV include the normal inter-frame mode and the affine mode (specifically, the affine inter-frame mode). In addition, the modes for not encoding the differential MV include the FRUC mode, the normal merge mode, and the affine mode (specifically, the affine merge mode). The inter-frame prediction unit 218 selects a mode for exporting the MV of the current block from these multiple modes and uses the selected mode to export the MV of the current block.

[0774] [MV Export > Normal Inter-Frame Mode]

[0775] For example, when the information decoded from the stream indicates the application of the normal inter-frame mode, the inter-frame prediction unit 218 exports the MV in the normal merge mode based on the information decoded from the stream and uses the MV for motion compensation (prediction).

[0776] Figure 84 It is a flowchart showing an example of inter-frame prediction by the normal inter-frame mode in the decoding device 200.

[0777] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sg_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0778] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_11 as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined order of priority (step Sg_12). In addition, the order of priority may also be determined in advance for each of the N prediction MV candidates.

[0779] Next, the inter-frame prediction unit 218 decodes the prediction MV selection information from the input stream and uses the decoded prediction MV selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction MV of the current block (step Sg_13).

[0780] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction MV (step Sg_14).

[0781] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are performed for each block. For example, when the processes of steps Sg_11 to Sg_15 are performed separately for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. Also, when the processes of steps Sg_11 to Sg_15 are performed separately for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. In addition, the processes of steps Sg_11 to Sg_15 may be such that when they are not performed for all the blocks included in a slice but for some of the blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, when the processes of steps Sg_11 to Sg_15 are performed for some of the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture may end.

[0782] [MV Derivation>Normal Merge Mode]

[0783] For example, in the case where the information read from the stream indicates the application of the normal merge mode, the inter-frame prediction unit 218 derives an MV in the normal merge mode and uses this MV for motion compensation (prediction).

[0784] Figure 85 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode in the decoding device 200.

[0785] The inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sh_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0786] Next, the inter-frame prediction unit 218 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_11 (step Sh_12). Specifically, the inter-frame prediction unit 218 obtains, for example, MV selection information included as a prediction parameter in the stream and selects the candidate MV identified by this MV selection information as the MV of the current block.

[0787] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processes of steps Sh_11 to Sh_13 are performed on each block, for example. For example, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a slice respectively, the inter-frame prediction using the ordinary merge mode for that slice ends. Also, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a picture respectively, the inter-frame prediction using the ordinary merge mode for that picture ends. In addition, the processes of steps Sh_11 to Sh_13 may be such that when they are not performed on all the blocks included in a slice but on some of the blocks, the inter-frame prediction using the ordinary merge mode for that slice ends. Similarly, the processes of steps Sh_11 to Sh_13 may be such that when they are performed on some of the blocks included in a picture, the inter-frame prediction using the ordinary merge mode for that picture ends.

[0788] [MV Derivation > FRUC Mode]

[0789] For example, in the case where the information read from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding device 100 side but is derived on the decoding device 200 side. For example, the decoding device 200 may also derive the motion information by performing a motion search. In this case, the decoding device 200 does not use the pixel values of the current block for the motion search.

[0790] Figure 86 It is a flowchart showing an example of the inter-frame prediction based on the FRUC mode in the decoding device 200.

[0791] First, the inter-frame prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block, and generates a list representing these MVs as candidate MVs (i.e., it is a candidate MV list and can also be common with the candidate MV list of the normal merge mode) (step Si_11). Next, the inter-frame prediction unit 218 selects the best candidate MV from among the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter-frame prediction unit 218 calculates the evaluation value of each candidate MV included in the candidate MV list, and selects one candidate MV as the best candidate MV based on this evaluation value. Then, the inter-frame prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Additionally, for example, the MV for the current block can also be derived by performing pattern matching in the peripheral region of the position within the reference picture corresponding to the selected best candidate MV. That is, for the region around the best candidate MV, a search using pattern matching and evaluation values in the reference picture is performed. Furthermore, in the case where there is an MV with a good evaluation value, the best candidate MV can also be updated to this MV and used as the final MV for the current block. The update to an MV with a better evaluation value may not be performed.

[0792] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processing of steps Si_11 to Si_15 is performed on each block, for example. For example, when the processing of steps Si_11 to Si_15 is performed on all the blocks included in a slice respectively, the inter-frame prediction using the FRUC mode for that slice ends. Additionally, when the processing of steps Si_11 to Si_15 is performed on all the blocks included in a picture respectively, the inter-frame prediction using the FRUC mode for that picture ends. The processing can also be performed in the same manner as the above block unit in units of sub-blocks.

[0793] [MV Derivation > Affine Merge Mode]

[0794] For example, in the case where the information read from the stream indicates the application of the affine merge mode, the inter-frame prediction unit 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using this MV.

[0795] Figure 87 It is a flowchart showing an example of inter-frame prediction based on the affine merge mode in the decoding device 200.

[0796] In the affine merge mode, the inter-frame prediction unit 218 first derives the MV of each control point of the current block (step Sk_11). As Figure 46AAs shown, the control points are the points at the upper left and upper right corners of the current block, or as Figure 46B shown, are the points at the upper left, upper right, and lower left corners of the current block.

[0797] For example, in the case of using the Figures 47A - 47C MV export method shown, as Figure 47A shown, the inter-frame prediction unit 218 checks these blocks in the order of the decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and determines the first valid block decoded in the affine mode.

[0798] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to export the MV of the control points. For example, when it is determined that block A has 2 control points, as Figure 47B shown, the inter-frame prediction unit 218 projects the motion vectors v3 and v4 of the upper left and upper right corners of the decoded block containing block A onto the current block to calculate the motion vector v0 of the upper left control point of the current block and the motion vector v1 of the upper right control point. Thus, the MV of each control point is exported.

[0799] In addition, as Figure 49A shown, when it is determined that block A has 2 control points, it is also possible to calculate the MV of 3 control points, or as Figure 49B shown, determine block A, and when block A has 3 control points, calculate the MV of 2 control points.

[0800] In addition, when the stream includes MV selection information as a prediction parameter, the inter-frame prediction unit 218 can also use this MV selection information to export the MV of each control point of the current block.

[0801] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 218 calculates the MV of each of the multiple sub-blocks as an affine MV using 2 motion vectors v0 and v1 and the above formula (1A), or using 3 motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_12). Then, the inter-frame prediction unit 218 uses these affine MVs and the decoded reference picture to perform motion compensation on the sub-block (step Sk_13). When the processes of steps Sk_12 and Sk_13 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a prediction image of the current block is generated.

[0802] In addition, in step Sk_11, the above-mentioned candidate MV list may also be generated. The candidate MV list may also be, for example, a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods may be Figures 47A - 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in, and any combination of other MV derivation methods.

[0803] In addition, the candidate MV list may also include candidate MVs of a prediction mode that performs prediction in units of sub-blocks other than the affine mode.

[0804] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with two control points and candidate MVs of an affine merge mode with three control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode with two control points and a candidate MV list including candidate MVs of an affine merge mode with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the modes of an affine merge mode with two control points and an affine merge mode with three control points may be generated.

[0805] [MV Derivation > Affine Inter-Frame Mode]

[0806] For example, when the information read from the stream indicates the application of the affine inter-frame mode, the inter-frame prediction unit 218 derives an MV in the affine inter-frame mode and performs motion compensation (prediction) using the MV.

[0807] Figure 88 is a flowchart showing an example of inter-frame prediction based on the affine inter-frame mode in the decoding device 200.

[0808] In the affine inter-frame mode, first, the inter-frame prediction unit 218 derives the predicted MVs (v0, v1) or (v0, v1, v2) of each of the two or three control points of the current block (step Sj_11). The control points are, for example, as Figure 46A or Figure 46B shown, the upper left corner, upper right corner, or lower left corner of the current block.

[0809] The inter-frame prediction unit 218 obtains the predicted MV selection information included as a prediction parameter in the stream, and uses the MV identified by the predicted MV selection information to derive the predicted MVs of the respective control points of the current block. For example, when using Figure 48A and Figure 48B the MV derivation methods shown in, the inter-frame prediction unit 218 selects Figure 48Aor Figure 48B The motion vectors of the blocks identified by the prediction MV selection information in the decoded blocks near the respective control points of the current block shown in Figure 48B are used to derive the predicted motion vectors (v0, v1) or (v0, v1, v2) of the control points of the current block.

[0810] Next, the inter-frame prediction unit 218 obtains, for example, each differential motion vector included as a prediction parameter in the stream, and adds the predicted motion vector of each control point of the current block and the differential motion vector corresponding to the predicted motion vector (step Sj_12). Thereby, the motion vectors of each control point of the current block are derived.

[0811] Next, the inter-frame prediction unit 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. That is, for each of the plurality of sub-blocks, the inter-frame prediction unit 218 calculates the motion vector of the sub-block as an affine motion vector using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sj_13). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine motion vectors and the decoded reference picture (step Sj_14). When the processes of steps Sj_13 and Sj_14 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0812] In addition, in step Sj_11, the above candidate motion vector list may be generated in the same manner as in step Sk_11.

[0813] [MV Derivation>Triangle Mode]

[0814] For example, in the case where the information read from the stream indicates the application of the triangle mode, the inter-frame prediction unit 218 derives a motion vector in the triangle mode and performs motion compensation (prediction) using the motion vector.

[0815] Figure 89 is a flowchart showing an example of inter-frame prediction based on the triangle mode in the decoding apparatus 200.

[0816] In the triangle mode, first, the inter-frame prediction unit 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter-frame prediction unit 218 may obtain, from the stream, partition information, which is information related to the division into each partition, as a prediction parameter. Moreover, the inter-frame prediction unit 218 may divide the current block into a first partition and a second partition according to the partition information.

[0817] Next, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0818] Then, the inter-frame prediction unit 218 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter-frame prediction unit 218 may also obtain MV selection information for identifying the selected candidate MVs from the stream as a prediction parameter. Then, the inter-frame prediction unit 218 may select the first MV and the second MV according to the MV selection information.

[0819] Next, the inter-frame prediction unit 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame prediction unit 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).

[0820] Finally, the inter-frame prediction unit 218 generates a prediction image of the current block by performing weighted addition on the first prediction image and the second prediction image (step Sx_16).

[0821] [Motion Search > DMVR]

[0822] For example, in the case where the information read from the stream indicates the application of DMVR, the inter-frame prediction unit 218 performs motion search by DMVR.

[0823] Figure 90 It is a flowchart showing an example of motion search based on DMVR in the decoding device 200.

[0824] The inter-frame prediction unit 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter-frame prediction unit 218 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is determined by DMVR.

[0825] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device 200.

[0826] First, the inter-frame prediction unit 218 is in Figure 58AIn Step 1 shown above, the search positions (also referred to as start points) of the initial MV representation and the costs of 8 search positions around it are calculated. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, when it is determined that the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 moves to the search position with the minimum cost and performs Figure 58A the process of Step 2 shown above. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 skips Figure 58A the process of Step 2 shown above and performs the process of Step 3.

[0827] In Figure 58A Step 2 shown above, the inter-frame prediction unit 218 uses the search position moved according to the processing result of Step 1 as a new start point and performs the same search as the processing of Step 1. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, if the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 performs the process of Step 4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 performs the process of Step 3.

[0828] In Step 4, the inter-frame prediction unit 218 treats the search position of the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0829] In Figure 58A Step 3 shown above, the inter-frame prediction unit 218 determines the pixel position with the minimum cost in decimal precision based on the costs of 4 points above, below, left, and right of the start point in Step 1 or Step 2, and uses this pixel position as the final search position. This pixel position in decimal precision is determined by weighted addition of the vectors (0, 1), (0, -1), (-1, 0), (1, 0) of the 4 points above, below, left, and right with the costs of the search positions of the respective 4 points as weights. Then, the inter-frame prediction unit 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0830] [Motion Compensation > BIO / OBMC / LIC]

[0831] For example, in the case where the information read from the stream represents the application of correction of the predicted image, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to the correction mode. This mode is, for example, the above-mentioned BIO, OBMC, and LIC, etc.

[0832] Figure 92 is a flowchart showing an example of generation of the predicted image in the decoding device 200.

[0833] The inter-frame prediction unit 218 generates a prediction image (step Sm_11), and corrects the prediction image by any of the above-described modes (step Sm_12).

[0834] Figure 93 It is a flowchart showing another example of the generation of the prediction image in the decoding device 200.

[0835] The inter-frame prediction unit 218 derives the MV of the current block (step Sn_11). Next, the inter-frame prediction unit 218 generates a prediction image using the MV (step Sn_12), and determines whether to perform a correction process (step Sn_13). For example, the inter-frame prediction unit 218 acquires the prediction parameters included in the stream, and determines whether to perform a correction process based on the prediction parameters. The prediction parameter is, for example, a flag indicating whether to apply each of the above-described modes. Here, when it is determined to perform a correction process (Yes in step Sn_13), the inter-frame prediction unit 218 generates a final prediction image by correcting the prediction image (step Sn_14). Further, in the LIC, the luminance and color difference of the prediction image can be corrected in step Sn_14. On the other hand, when it is determined not to perform a correction process (No in step Sn_13), the inter-frame prediction unit 218 outputs the prediction image without correcting it as the final prediction image (step Sn_15).

[0836] [Motion Compensation>OBMC]

[0837] For example, in the case where the information read from the stream indicates the application of OBMC, when generating a prediction image, the inter-frame prediction unit 218 corrects the prediction image according to OBMC.

[0838] Figure 94 It is a flowchart showing an example of the correction of the prediction image based on OBMC in the decoding device 200. In addition, Figure 94 The flowchart of Figure 62 shows the process of correcting the prediction image using the current picture and the reference picture shown in

[0839] First, as Figure 62 shown, the inter-frame prediction unit 218 acquires a prediction image (Pred) based on normal motion compensation using the MV assigned to the current block.

[0840] Next, the inter-frame prediction unit 218 applies (reuses) the MV (MV_L) that has been derived for the decoded left adjacent block to the current block, and acquires a prediction image (Pred_L). Then, the inter-frame prediction unit 218 performs the first correction of the prediction image by overlapping the two prediction images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0841] Similarly, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_U) that has been derived for the decoded upper adjacent block to the current block to obtain a predicted image (Pred_U). Then, the inter-frame prediction unit 218 performs a second correction of the predicted image by overlapping the predicted image Pred_U with the predicted images that have undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block that is blended (smoothed) with the boundaries of the adjacent blocks.

[0842] [Motion Compensation > BIO]

[0843] For example, in the case where the information read from the stream indicates the application of BIO, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to BIO.

[0844] Figure 95 It is a flowchart showing an example of the correction of the predicted image based on BIO in the decoding device 200.

[0845] As Figure 63 shown, the inter-frame prediction unit 218 uses two reference pictures (Ref0, Ref1) different from the picture (CurPic) containing the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 218 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_11). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0846] Next, the inter-frame prediction unit 218 uses the motion vector M0 and the reference picture L0 to derive the interpolation image I 0 of the current block. Additionally, the inter-frame prediction unit 218 uses the motion vector M1 and the reference picture L1 to derive the interpolation image I 1 of the current block (step Sy_12). Here, the interpolation image I 0 is the image contained in the reference picture Ref0 derived for the current block, and the interpolation image I 1 is the image contained in the reference picture Ref1 derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Or, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each also be an image larger than the current block. Furthermore, the interpolation image I 0 and I 1It may include a predicted image derived by applying motion vectors (M0, M1) and reference pictures (L0, L1), as well as a motion compensation filter.

[0847] In addition, the inter-frame prediction unit 218 derives the gradient image (Ix 0 and Ix 1 of the current block from the interpolation images I 0 and Ix 1 Iy 0 and Iy 1 )(step Sy_13). In addition, the gradient image in the horizontal direction is (Ix 0 and Ix 1 ), and the gradient image in the vertical direction is (Ix 0 and Ix 1 ). The inter-frame prediction unit 218 can also derive the gradient image by applying a gradient filter to the interpolation image, for example. The gradient image only needs to be an image representing the spatial variation amount of pixel values along the horizontal direction or the vertical direction.

[0848] Next, the inter-frame prediction unit 218 uses the interpolation images (I 0 and I 1 ) and the gradient image (Ix 0 and Ix 1 Iy 0 and Iy 1 ) to derive the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block (step Sy_14). As an example, the sub-block may be a 4x4 pixel sub-CU.

[0849] Next, the inter-frame prediction unit 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame prediction unit 218 derives a correction value of the pixel values included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter-frame prediction unit 218 can correct the predicted image of the current block using the correction value (step Sy_16). In addition, the correction value can be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.

[0850] In addition, the processing flow of BIO is not limited to Figure 95 the disclosed processing. It can either only implement Figure 95 a part of the disclosed processing, or append or replace different processing, or execute according to a different processing order.

[0851] [Motion Compensation>LIC]

[0852] For example, in the case where the information read from the stream represents the application of LIC, when generating a predicted image, the inter-frame prediction unit 218 corrects the predicted image according to LIC.

[0853] Figure 96 It is a flowchart showing an example of the correction of the predicted image based on LIC in the decoding device 200.

[0854] First, the inter-frame prediction unit 218 uses the MV to obtain a reference image corresponding to the current block from the decoded reference pictures (step Sz_11).

[0855] Next, the inter-frame prediction unit 218 extracts information indicating how the luminance values change in the reference picture and the current picture for the current block (step Sz_12). As Figure 66A shown, this extraction is based on the luminance pixel values of the decoded left adjacent reference region (peripheral reference region) and the decoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. Then, the inter-frame prediction unit 218 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_13).

[0856] The inter-frame prediction unit 218 generates a predicted image for the current block by performing a luminance correction process of applying its luminance correction parameter to the reference image in the reference picture specified by the MV (step Sz_14). That is, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the color difference can be corrected.

[0857] [Prediction control unit]

[0858] The prediction control unit 220 selects one of the intra-frame predicted image and the inter-frame predicted image, and outputs the selected predicted image to the adder 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device 200 side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device 100 side.

[0859] [First form]

[0860] Figure 97 It is a diagram showing an example of the syntax in which the encoding device 100 notifies one or more PTL parameters and one or more HRD parameters in the VPS.

[0861] The PTL parameter is a parameter representing information including profile information, level information, and tier information. Here, the profile information is information including information representing a specific subset in the moving picture coding technology. The level information is information including information related to a restriction group related to values of parameters and syntax elements in the moving picture coding technology, etc. The tier information is information including information related to the category of level restrictions forming a nested structure in a certain level restriction group.

[0862] The HRD parameter is a parameter representing information including information such as a bit rate and the size of a CPB (Coded Picture Buffer).

[0863] In addition, the OLS (Output Layer Set) is a set of layers including at least one output layer. The encoding device 100 and the decoding device 200 can process the OLS as a combination of layers capable of decoding and displaying images simultaneously.

[0864] As Figure 97 shown, the encoding device 100 can generate a bitstream according to information related to the number of each parameter (for example, the value of vps_num_ptls_hrds) that is shared in one VPS, to notify one or more PTL parameters (for example, profile_tier_level()) and one or more HRD parameters (for example, ols_hrd_parameters()).

[0865] In addition, the encoding device 100 can also generate a bitstream that omits the notification of an HRD parameter group including general_hrd_parameters() and ols_hrd_parameters(), etc. by notifying a specified flag (for example, vps_general_hrd_params_present_flag).

[0866] Furthermore, the encoding device 100 can also change the number of pieces of information related to time sub-layers included in the i-th PTL parameter and the i-th HRD parameter according to information related to the maximum value of the time layer ID shared in one VPS (for example, the value of ptl_hrd_max_temporal_id[i]).

[0867] In addition, the encoding device 100 can also use an index (for example, ols_ptl_hrd_idx[i]) shared in one VPS to generate a bitstream for the encoding device 100 to notify the following information: information about one or more PTL parameters notified by the VPS, and the PTL parameter and HRD parameter information corresponding to the i-th OLS among one or more HRD parameters.

[0868] In addition, if the PTL parameters and the HRD parameter group are notified after the information related to the number of each parameter (e.g., vps_num_ptls_hrds) shared in a VPS and the maximum value information of the temporal layer ID (e.g., ptl_hrd_max_temporal_id) shared in a VPS, they may also be notified in an order different from the Figure 97 example shown.

[0869] Figure 98 FIG. is an example of the syntax representing the PTL parameters.

[0870] When the encoding device 100 notifies the i-th PTL parameter in the Figure 97 syntax shown, it sets the value of the i-th ptl_hrd_max_temporal_id for the argument maxNumSubLayersMinus1 and calls profile_tier_level(). The encoding device 100 changes the number of sublayer_level_present_flag or sublayer_level_idc notified in the bitstream according to the value of the argument maxNumSubLayersMinus1.

[0871] Figure 99 FIG. is an example of the syntax representing the HRD parameters.

[0872] When the encoding device 100 notifies the i-th HRD parameter in the Figure 97 syntax shown, it sets the i-th ptl_hrd_max_temporal_id for the argument maxSubLayers and calls ols_hrd_parameters(). The encoding device 100 changes the number of fixed_pic_rate_general_flag or sublayer_hrd_parameters() notified in the bitstream according to the value of the argument maxSubLayers.

[0873] Figure 100 FIG. is a flowchart showing the process of the decoding device 200 analyzing the PTL parameters and HRD parameters notified in the VPS.

[0874] First, the encoding device 100 obtains the information N related to the number of both in the PTL parameters and the HRD parameters (step S100). The information N related to the number of both is, for example, Figure 97 vps_num_ptls_hrds.

[0875] Next, the encoding device 100 performs loop processing in a VPS based on the shared information N related to the number of each parameter. The encoding device 100 obtains the information related to the maximum value of the temporal layer ID shared by the i-th PTL parameter and the HRD parameter in the loop processing (step S101). The information related to the maximum value of the temporal layer ID shared by the i-th PTL parameter and the HRD parameter is, for example, Figure 97 The ptl_hrd_max_temporal_id[i] shown. In addition, i is an integer between 0 and N - 1.

[0876] In addition, the encoding device 100 performs loop processing in a VPS based on the shared information N related to the number of each parameter. The encoding device 100 obtains N PTL parameters (step S102).

[0877] Next, the encoding device 100 performs loop processing according to the number of OLSs ( Figure 97 Shown as TotalNumOlss). The encoding device 100 obtains the index shared by the PTL parameter or the HRD parameter corresponding to each OLS, which is the number below the value of TotalNumOlss (step S103). The index shared by the PTL parameter or the HRD parameter corresponding to each OLS is, for example, Figure 97 Of ols_hrd_parameters().

[0878] Then, the encoding device 100 starts N loops. Then, the encoding device 100 obtains the HRD parameter according to the maximum value of the corresponding temporal layer ID (step S104). Here, the encoding device 100 ends N loops.

[0879] As described above, according to the structure of the first form, the encoding device 100 may be able to reduce the number of bits of the VPS parameter by using the parameters shared between the PTL parameter and the HRD parameter. Here, the parameters shared between the PTL parameter and the HRD parameter refer to vps_num_ptls_hrds, ptl_hrd_max_temporal_id[i], or ols_ptl_hrd_idx[i].

[0880] In addition, the encoding device 100 may be able to simplify the correspondence between the OLS and the PTL parameter or the HRD parameter, and may be able to simplify the structure of the decoding device 200.

[0881] [Combination with Other Forms]

[0882] This form can be combined with at least a part of other forms in the embodiments of the present disclosure. In addition, a part of the processing described in the flowchart of this form, a part of the structure of the device, a part of the syntax, etc. can also be combined and implemented with other forms.

[0883] In addition, the encoding device 100 can output parameters in the same process as the above VPS analysis process in the decoding device 200. In addition, the encoding device 100 does not necessarily need all the components described in this form, and may only have a part of the components of the first form.

[0884] [Second form]

[0885] Figure 101 It is a diagram showing an example of the syntax for notifying one or more PTL parameters in the VPS.

[0886] As Figure 101 In the example shown, the encoding device 100 can generate a bitstream according to information indicating the number of each parameter such as the PTL parameter (for example, the value indicating the number shown by vps_num_ptls, etc.) to notify one or more PTL parameters, etc. (for example, profile_tier_level()).

[0887] In addition, the encoding device 100 can generate a bitstream to notify the following value: the value of the index (for example, ols_ptl_idx[i]) for each OLS in order to indicate which PTL among the one or more PTLs notified by the VPS corresponds to the i-th OLS.

[0888] In addition, the encoding device 100 can also generate a bitstream for the following notification: for an OLS with the number of layers included in the OLS being 1, the same PTL parameter as the PTL parameter notified in the SPS referred to during decoding of the layer of this OLS is used as the PTL parameter corresponding to this OLS and is notified by the VPS. Here, for example, the number of layers included in the i-th OLS is represented by NumLayersInOls[i].

[0889] In addition, the layer described here refers to the layer identified by nuh_layer_id notified by the NAL unit layer, which is different from the temporal layer identified by nuh_temporal_id_plus1 notified by the NAL unit layer. Figure 101 The example shown is an example of the semantics of ols_ptl_idx[i].

[0890] The following shows an example of the semantics of ols_ptl_idx[i].

[0891] ols_ptl_idx[i] identifies the index into the list of profile_tier_level() syntax structures applicable to the i-th OLS. The value of ols_hrd_idx[i] is in the range of 0 to num_ols_hrd_params_minus1.

[0892] When NumLayersInOls[i] is equal to 1, the profile_tier_level() syntax structure applied to the i-th OLS exists in both the VPS and SPS referenced by the layer in the i-th OLS. Then, when the profile_tier_level() syntax structure exists in both the VPS and the SPS, it is a condition for satisfying the standard conformance of the bitstream that the profile_tier_level() syntax structure encoded into the VPS and the SPS for the i-th OLS is the same.

[0893] Alternatively, the encoding device 100 may omit encoding of ols_ptl_idx[i] by predetermining a method for determining which index to use when ols_ptl_idx[i] does not exist.

[0894] Furthermore, the encoding apparatus 100 may generate a bit stream in such a manner that, for an OLS including one layer, the PTL parameter corresponding to the OLS is notified by the SPS referred to when decoding the layers of the OLS but by the VPS.

[0895] Furthermore, when the OLS is the 0th OLS, the encoding apparatus 100 may generate a bit stream in such a manner that the corresponding PTL parameter is not notified by the VPS but by the SPS referred to when decoding the layer of the OLS.

[0896] Figure 102 This is a flowchart showing a process in which the encoding device 100 notifies the PTL parameters in the VPS and the SPS.

[0897] First, the encoding device 100 starts the first loop as a loop for each layer. The encoding device 100 determines the PTL parameter for each layer in the first loop (step S200). Then, the encoding device 100 ends the first loop as a loop for each layer.

[0898] Next, the coding device 100 starts the second loop which is a loop for each OLS. The coding device 100 determines a PTL parameter for each OLS, and describes the PTL parameter for each OLS in the VPS.

[0899] Next, the encoding device 100 determines whether the number of layers in the OLS is equal to 1 (step S201). When the encoding device 100 determines that the number of layers in the OLS is equal to 1 (Yes in step S201), the encoding device 100 obtains the PTL parameters corresponding to the layer included in the OLS from the PTL parameters determined for each layer by the encoding device 100 in the first cycle (step S203). In step S203, the encoding device 100 may copy the PTL parameters corresponding to the layer included in the OLS.

[0900] Next, the encoding device 100 writes the obtained PTL parameters into the VPS (step S204).

[0901] When the encoding device 100 does not determine that the number of layers in the OLS is equal to 1 (No in step S201), the encoding device 100 determines the PTL parameters of the OLS (step S202).

[0902] Then, the encoding device 100 writes the determined PTL parameters into the VPS (step S204). Then, the encoding device 100 ends the second cycle as the cycle for each OLS.

[0903] In addition, the encoding device 100 starts the third cycle as the cycle for each layer. The encoding device 100 writes the PTL parameters determined for each layer in the first cycle into the SPS corresponding to each layer (step S205).

[0904] Then, the encoding device 100 ends the third cycle as the cycle for each layer.

[0905] As described above, according to the structure of the second form, the encoding device 100 can generate a bitstream that can obtain the PTL parameters corresponding to all OLSs by parsing the VPS, and it is possible to simplify the bitstream parsing process in a system or a decoding device 200 that processes bitstreams including multiple layers.

[0906] [Combination with other forms]

[0907] This form can be combined with at least a part of other forms in the embodiments of the present disclosure. In addition, a part of the processes described in the flowchart of this form, a part of the structure of the device, a part of the syntax, etc. can also be implemented in combination with other forms.

[0908] In addition, the encoding device 100 can also output parameters according to the same process as the above VPS parsing process in the decoding device 200. In addition, it is not necessary to have all the constituent elements described in this form, and it is also possible to have only a part of the constituent elements of the second form.

[0909] [Third form]

[0910] Figure 103 This is a figure showing an example of the syntax for notifying one or more HRD parameters in a VPS. As Figure 103 shown in the example, the encoding device 100 generates a bitstream in a manner that notifies one or more HRD parameters according to the information related to the number of HRD parameters. Here, the information related to the number of HRD parameters can be, for example, a value represented by an expression such as num_ols_hrd_params_minus1 + 1. Additionally, the HRD parameter can also refer to the value indicated by ols_hrd_parameters().

[0911] Furthermore, the encoding device 100 can also generate a bitstream that omits the notification of the HRD parameter group by writing the information represented by a flag into the bitstream. Here, the flag refers to, for example, vps_general_hrd_params_present_flag. Additionally, the HRD parameters included in the HRD parameter group are, for example, general_hrd_parameters() or ols_hrd_parameters().

[0912] Moreover, the encoding device 100 can also change the number ...

Claims

1. An encoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, in operation, the circuit for a set of layers including at least one output layer, even if the number of layers included in the set of layers is 1, encodes performance requirement information representing the performance requirements of a decoding device into a header shared by the layers included in the set of layers, encodes a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled into the shared header, generates a bitstream without encoding the HRD parameters into the shared header according to the flag, the bitstream includes the header shared by the layers included in the set of layers and data obtained by encoding at least one image included in the output layer.

2. An encoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, in operation, the circuit for the layers included in an OLS output layer set with 1 layer, even if the number of layers in the OLS is 1, encodes information representing the performance requirements of a decoding device into a VPS (Video Parameter Set), encodes a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled into the VPS, generates a bitstream without encoding the HRD parameters into the VPS according to the flag, the bitstream includes the VPS and data obtained by encoding at least one image included in the OLS.

3. A decoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, in the operation of decoding a bitstream, the circuit for a set of layers including at least one output layer, even if the number of layers included in the set of layers is 1, decodes performance requirement information representing the performance requirements of a decoding device from the header shared by the layers included in the set of layers, decodes a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled from the shared header, generates at least one image without decoding the HRD parameters from the shared header according to the flag, the bitstream includes the header shared by the layers included in the set of layers and data obtained by encoding the at least one image included in the output layer.

4. A decoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, in the operation of decoding a bitstream, the circuit for the layers included in an OLS output layer set with 1 layer, even if the number of layers in the OLS is 1, decodes information representing the performance requirements of a decoding device from a VPS (Video Parameter Set), decodes a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled from the VPS, generates at least one image without decoding the HRD parameters from the VPS according to the flag, the bitstream includes the VPS and data obtained by encoding the at least one image included in the OLS.

5. An encoding method, wherein for a set of layers including at least one output layer, even if the number of layers included in the set of layers is 1, encodes performance requirement information representing the performance requirements of a decoding device into the header shared by the layers included in the set of layers, encodes a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled into the shared header, Generate a bitstream without encoding the HRD parameters into the common header according to the flag. The bitstream includes a common header shared by the layers included in the layer set and data obtained by encoding at least one image included in the output layer.

6. A decoding method, wherein For a layer set including at least one output layer, even if the number of layers included in the layer set is 1, decode performance requirement information representing performance requirements of a decoding device from a common header shared by the layers included in the layer set. Decode a flag indicating whether HRD parameters, i.e., assumed reference decoder parameters, are signaled from the common header. According to the flag, do not decode the HRD parameters from the common header. Generate at least one image based on the performance requirement information. The bitstream includes a common header shared by the layers included in the layer set and data obtained by encoding the at least one image included in the output layer.

Citation Information

Patent Citations

  • Hypothetical reference decoder parameter syntax structure

    CN104704842A

  • Image decoding device

    CN106170981A