Encoding device, decoding device, encoding method, and decoding method

By using syntactic elements and deblocking filtering processing associated with chroma tool offset in the video encoding device and decoding device, dynamically control parameters, optimize the video encoding process, solve the problems of low encoding efficiency, poor picture quality and large processing volume, and achieve more efficient video encoding and decoding.

CN115176479BActive Publication Date: 2025-07-04PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180015527.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-25
Filing Date
2021-02-25
Publication Date
2025-07-04
Estimated Expiration
2041-02-25

AI Technical Summary

Technical Problem

When processing video data, existing video encoding technologies have problems such as low encoding efficiency, poor picture quality, large processing volume, excessive circuit scale and improper filter selection, making it difficult to achieve effective optimization and improvement.

Method used

By using syntactic elements associated with chroma tool offset and deblocking filtering processing in the encoding device and the decoding device, the storage and use of the first and second parameters are dynamically controlled, the encoding and decoding process is optimized, the bitstream size and processing volume are reduced, and the image quality is improved.

Benefits of technology

The improvement of encoding efficiency, image quality, reduction of processing volume, reduction of circuit scale and acceleration of processing speed have been achieved. Filters and other encoding elements have been appropriately selected, solving the shortcomings in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115176479B_ABST
    Figure CN115176479B_ABST
Patent Text Reader

Abstract

The encoding device (100) includes a circuit and a memory connected to the circuit, and saves a first parameter in a bitstream, where the first parameter indicates whether a syntax element associated with a chrominance tool offset exists in the bitstream (S201). When the first parameter indicates that the syntax element exists in the bitstream (Yes in S202), (i) in addition to the syntax element, one or more second parameters used in the deblocking filtering process of the chrominance samples of the image to be processed are saved in the bitstream (S203), (ii) the deblocking filtering process is performed using one or more second parameters, and (iii) the image to be processed is encoded (S204). When the first parameter indicates that the syntax element does not exist in the bitstream (No in S202), the image to be processed is encoded without saving one or more second parameters in the bitstream (S205).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to video coding, and particularly to systems, components, and methods in the encoding and decoding of moving images, etc. Background Art

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to handle the continuously increasing amount of digital video data in various applications, there has always been a need to provide improvements and optimizations to video coding technology. The present disclosure relates to further progress, improvements, and optimizations in video coding.

[0003] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above video coding technology.

[0004] Prior Art Documents

[0005] Non-Patent Documents

[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention

[0007] Problems to be Solved by the Invention

[0008] Regarding the encoding method as described above, for the improvement of encoding efficiency, the improvement of image quality, the reduction of processing volume, the reduction of circuit scale, or the appropriate selection of elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, etc., it is desired to propose a new method.

[0009] The present disclosure provides a structure or method that can contribute to one or more of, for example, the improvement of encoding efficiency, the improvement of image quality, the reduction of processing volume, the reduction of circuit scale, the improvement of processing speed, and the appropriate selection of elements or operations. In addition, the present disclosure may include a structure or method that can contribute to benefits other than the above.

[0010] Means for Solving the Problems

[0011] For example, an encoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit. During operation, the circuit stores a first parameter in the bitstream, where the first parameter indicates whether a syntax element associated with a chroma tool offset exists in the bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in deblocking filtering of chroma samples of the image to be processed are stored in the bitstream, (ii) the deblocking filtering is performed using the one or more second parameters, and (iii) the image to be processed is encoded. When the first parameter indicates that the syntax element does not exist in the bitstream, the image to be processed is encoded without storing the one or more second parameters in the bitstream.

[0012] Each embodiment or a part of the structure or method in the present disclosure can respectively achieve at least any one of improvements in encoding efficiency, improvements in image quality, reduction in the amount of encoding / decoding processing, reduction in circuit scale, or improvement in encoding / decoding processing speed, for example. Or, each embodiment or a part of the structure or method in the present disclosure can respectively make an appropriate selection of elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, etc. during encoding and decoding. In addition, the present disclosure also includes the disclosure of structures or methods that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in the amount of processing.

[0013] Based on the description and the drawings, further advantages and effects in one aspect of the present disclosure have been clarified. These advantages and / or effects are respectively obtained through several embodiments and the features described in the description and the drawings, but it is not necessary to provide all of them in order to obtain one or more of the advantages and / or effects.

[0014] In addition, these general or specific aspects can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can also be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0015] Advantages of the Invention

[0016] The structure or method according to one aspect of the present disclosure can contribute to one or more of improvements in encoding efficiency, improvements in image quality, reduction in the amount of processing, reduction in circuit scale, improvement in processing speed, and appropriate selection of elements or actions, for example. In addition, the structure or method according to one aspect of the present disclosure can also contribute to benefits other than the above. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram showing an example of the structure of a transmission system according to an embodiment.

[0018] Figure 2 It is a diagram showing an example of the hierarchical structure of data in a stream.

[0019] Figure 3 It is a diagram showing an example of the structure of a slice.

[0020] Figure 4 It is a diagram showing an example of the structure of a tile.

[0021] Figure 5 It is a diagram showing an example of the coding structure in scalable coding.

[0022] Figure 6 It is a diagram showing an example of the coding structure in scalable coding.

[0023] Figure 7 It is a block diagram showing an example of the functional structure of an encoding device according to an embodiment.

[0024] Figure 8 It is a block diagram showing an example of the installation of an encoding device.

[0025] Figure 9 It is a flowchart showing an example of the overall encoding process performed by an encoding device.

[0026] Figure 10 It is a diagram showing an example of block segmentation.

[0027] Figure 11 It is a diagram showing an example of the functional structure of a segmentation unit.

[0028] Figure 12 It is a diagram showing an example of a segmentation pattern.

[0029] Figure 13A It is a diagram showing an example of the syntax tree of a segmentation pattern.

[0030] Figure 13B It is a diagram showing another example of the syntax tree of a segmentation pattern.

[0031] Figure 14 It is a table showing transform basis functions corresponding to each transform type.

[0032] Figure 15 It is a diagram showing an example of SVT.

[0033] Figure 16 It is a flowchart showing an example of the process performed by a transform unit.

[0034] Figure 17 It is a flowchart showing another example of the process performed by a transform unit.

[0035] Figure 18 It is a block diagram showing an example of the functional structure of the quantization unit.

[0036] Figure 19 It is a flowchart showing an example of the quantization performed by the quantization unit.

[0037] Figure 20 It is a block diagram showing an example of the functional structure of the entropy encoding unit.

[0038] Figure 21 It is a diagram showing the process of CABAC in the entropy encoding unit.

[0039] Figure 22 It is a block diagram showing an example of the functional structure of the loop filtering unit.

[0040] Figure 23A It is a diagram showing an example of the shape of the filter used in the ALF (adaptive loop filter).

[0041] Figure 23B It is a diagram showing another example of the shape of the filter used in the ALF.

[0042] Figure 23C It is a diagram showing another example of the shape of the filter used in the ALF.

[0043] Figure 23D It is a diagram showing an example where the Y sample (first component) is used for the CCALF of Cb and the CCALF of Cr (multiple components different from the first component).

[0044] Figure 23E It is a diagram showing a diamond-shaped filter.

[0045] Figure 23F It is a diagram showing an example of JC-CCALF.

[0046] Figure 23G It is a diagram showing an example of the weight_index candidate of JC-CCALF.

[0047] Figure 24 It is a block diagram showing an example of the detailed structure of the loop filtering unit that functions as a DBF.

[0048] Figure 25 It is a diagram showing an example of deblocking filtering with filtering characteristics symmetric with respect to the block boundary.

[0049] Figure 26 It is a diagram for explaining an example of the block boundary where deblocking filtering processing is performed.

[0050] Figure 27 This is a diagram showing an example of the Bs value.

[0051] Figure 28 This is a flowchart showing an example of the processing performed by the prediction unit of the encoding device.

[0052] Figure 29 This is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.

[0053] Figure 30 This is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.

[0054] Figure 31 This is a diagram showing an example of 67 intra prediction modes in intra prediction.

[0055] Figure 32 This is a flowchart showing an example of the processing performed by the intra prediction unit.

[0056] Figure 33 This is a diagram showing an example of each reference picture.

[0057] Figure 34 This is a conceptual diagram showing an example of the reference picture list.

[0058] Figure 35 This is a flowchart showing the flow of the basic processing of inter prediction.

[0059] Figure 36 This is a flowchart showing an example of MV derivation.

[0060] Figure 37 This is a flowchart showing another example of MV derivation.

[0061] Figure 38A This is a diagram showing an example of the classification of each mode of MV derivation.

[0062] Figure 38B This is a diagram showing an example of the classification of each mode of MV derivation.

[0063] Figure 39 This is a flowchart showing an example of inter prediction based on the normal inter mode.

[0064] Figure 40 This is a flowchart showing an example of inter prediction based on the normal merge mode.

[0065] Figure 41 This is a diagram for explaining an example of the MV derivation process based on the normal merge mode.

[0066] Figure 42 This is a diagram for explaining an example of the MV derivation process based on the HMVP mode.

[0067] Figure 43 It is a flowchart showing an example of FRUC (frame rate up conversion).

[0068] Figure 44 It is a diagram showing an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.

[0069] Figure 45 It is a diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture.

[0070] Figure 46A It is a diagram showing an example of the derivation of MV in units of sub - blocks in the affine mode using two control points.

[0071] Figure 46B It is a diagram showing an example of the derivation of MV in units of sub - blocks in the affine mode using three control points.

[0072] Figure 47A It is a conceptual diagram showing an example of the derivation of MV of control points in the affine mode.

[0073] Figure 47B It is a conceptual diagram showing an example of the derivation of MV of control points in the affine mode.

[0074] Figure 47C It is a conceptual diagram showing an example of the derivation of MV of control points in the affine mode.

[0075] Figure 48A It is a diagram showing an affine mode with two control points.

[0076] Figure 48B It is a diagram showing an affine mode with three control points.

[0077] Figure 49A It is a conceptual diagram showing an example of a method for deriving MV of control points when the number of control points in the encoded block and the current block is different.

[0078] Figure 49B It is a conceptual diagram showing another example of a method for deriving MV of control points when the number of control points in the encoded block and the current block is different.

[0079] Figure 50 It is a flowchart showing an example of the process of the affine merge mode.

[0080] Figure 51 It is a flowchart showing an example of the process of the affine inter - frame mode.

[0081] Figure 52A It is a diagram for explaining the generation of predicted images of two triangles.

[0082] Figure 52B It is a conceptual diagram showing an example of the first part of the first partition, the first sample set, and the second sample set.

[0083] Figure 52C It is a conceptual diagram showing the first part of the first partition.

[0084] Figure 53 It is a flowchart showing an example of a triangular pattern.

[0085] Figure 54 It is a diagram showing an example of the ATMVP mode for deriving MV in sub - block units.

[0086] Figure 55 It is a diagram showing the relationship between the merge mode and DMVR (dynamic motion vector refreshing).

[0087] Figure 56 It is a conceptual diagram for explaining an example of DMVR.

[0088] Figure 57 It is a conceptual diagram for explaining another example of DMVR for determining MV.

[0089] Figure 58A It is a diagram showing an example of motion search in DMVR.

[0090] Figure 58B It is a flowchart showing an example of motion search in DMVR.

[0091] Figure 59 It is a flowchart showing an example of the generation of a predicted image.

[0092] Figure 60 It is a flowchart showing another example of the generation of a predicted image.

[0093] Figure 61 It is a flowchart for explaining an example of the predicted image correction process based on OBMC (overlapped block motion compensation).

[0094] Figure 62 It is a conceptual diagram for explaining an example of the predicted image correction process based on OBMC.

[0095] Figure 63 It is a diagram for explaining a model assuming uniform linear motion.

[0096] Figure 64 It is a flowchart showing an example of inter-frame prediction according to BIO.

[0097] Figure 65 It is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.

[0098] Figure 66A It is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC (local illumination compensation).

[0099] Figure 66B It is a flowchart showing an example of a method for generating a predicted image using a luminance correction process based on LIC.

[0100] Figure 67 It is a block diagram showing the functional structure of a decoding device according to an embodiment.

[0101] Figure 68 It is a block diagram showing an example of the installation of a decoding device.

[0102] Figure 69 It is a flowchart showing an example of the overall decoding process performed by a decoding device.

[0103] Figure 70 It is a diagram showing the relationship between a segmentation determination unit and other components.

[0104] Figure 71 It is a block diagram showing an example of the functional structure of an entropy decoding unit.

[0105] Figure 72 It is a diagram showing the process of CABAC in an entropy decoding unit.

[0106] Figure 73 It is a block diagram showing an example of the functional structure of an inverse quantization unit.

[0107] Figure 74 It is a flowchart showing an example of the inverse quantization performed by an inverse quantization unit.

[0108] Figure 75 It is a flowchart showing an example of the process performed by an inverse transform unit.

[0109] Figure 76 It is a flowchart showing another example of the process performed by an inverse transform unit.

[0110] Figure 77 It is a block diagram showing an example of the functional structure of a loop filter unit.

[0111] Figure 78 This is a flowchart showing an example of the processing performed by the prediction unit of the decoding device.

[0112] Figure 79 This is a flowchart showing another example of the processing performed by the prediction unit of the decoding device.

[0113] Figure 80A This is a flowchart showing a part of another example of the processing performed by the prediction unit of the decoding device.

[0114] Figure 80B This is a flowchart showing the remaining part of another example of the processing performed by the prediction unit of the decoding device.

[0115] Figure 81 This is a diagram showing an example of the processing performed by the intra prediction unit of the decoding device.

[0116] Figure 82 This is a flowchart showing an example of MV derivation in the decoding device.

[0117] Figure 83 This is a flowchart showing another example of MV derivation in the decoding device.

[0118] Figure 84 This is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding device.

[0119] Figure 85 This is a flowchart showing an example of inter prediction based on the normal merge mode in the decoding device.

[0120] Figure 86 This is a flowchart showing an example of inter prediction based on the FRUC mode in the decoding device.

[0121] Figure 87 This is a flowchart showing an example of inter prediction based on the affine merge mode in the decoding device.

[0122] Figure 88 This is a flowchart showing an example of inter prediction based on the affine inter mode in the decoding device.

[0123] Figure 89 This is a flowchart showing an example of inter prediction based on the triangular mode in the decoding device.

[0124] Figure 90 This is a flowchart showing an example of motion search based on DMVR in the decoding device.

[0125] Figure 91 This is a flowchart showing a detailed example of motion search based on DMVR in the decoding device.

[0126] Figure 92 This is a flowchart showing an example of generating a predicted image in a decoding device.

[0127] Figure 93 This is a flowchart showing another example of generating a predicted image in a decoding device.

[0128] Figure 94 This is a flowchart showing an example of correcting a predicted image based on OBMC in a decoding device.

[0129] Figure 95 This is a flowchart showing an example of correcting a predicted image based on BIO in a decoding device.

[0130] Figure 96 This is a flowchart showing an example of correcting a predicted image based on LIC in a decoding device.

[0131] Figure 97 This is a flowchart showing an example of the internal configuration of a loop filter in a decoding device of the first form.

[0132] Figure 98 This is a diagram showing an example of the position of the first parameter in a bitstream.

[0133] Figure 99 This is a diagram showing an example of the position of the second parameter in a bitstream.

[0134] Figure 100 This is a diagram showing an example of the syntax structure in a PPS (Picture Parameter Set).

[0135] Figure 101 This is a diagram showing an example of the syntax structure in a PH (Picture Header).

[0136] Figure 102 This is a diagram showing an example of the syntax structure in a SH (Slice Header).

[0137] Figure 103 This is a flowchart showing the operations performed by an encoding device.

[0138] Figure 104 This is a flowchart showing the operations performed by a decoding device.

[0139] Figure 105 This is an overall structure diagram of a content supply system for implementing a content distribution service.

[0140] Figure 106 This is a diagram showing an example of a display screen of a web page.

[0141] Figure 107 This is a diagram showing an example of a display screen of a web page.

[0142] Figure 108 This is a diagram showing an example of a smartphone.

[0143] Figure 109 This is a block diagram showing a structural example of a smartphone. Detailed implementation

[0144] [Introduction]

[0145] An encoding device according to an aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, the circuit stores a first parameter in the bitstream, the first parameter indicating whether a syntax element associated with chroma tool offset exists in the bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking filtering process of the chroma samples of the image to be processed are stored in the bitstream, (ii) the deblocking filtering process is performed using the one or more second parameters, and (iii) the image to be processed is encoded. When the first parameter indicates that the syntax element does not exist in the bitstream, the one or more second parameters are not stored in the bitstream and the image to be processed is encoded.

[0146] Thus, by sharing the first parameter between the process using the syntax element associated with chroma tool offset and the deblocking filtering process, the encoding device can reduce the size of the bitstream and the processing amount at the same time.

[0147] For example, the first parameter may be pps_chroma_tool_offsets_present_flag set in the picture parameter set.

[0148] Thus, the encoding device can determine whether one or more second parameters exist in the bitstream, for example, without parsing a slice header more than the picture parameter set.

[0149] For example, the chroma samples may represent the Cb component and the Cr component in the YCbCr color space.

[0150] Thus, since the image to be processed has information on luminance (Y component) and two chroma components (Cb component and Cr component), the encoding device can reduce information related to chroma while suppressing a decrease in subjective image quality. Therefore, since the amount of information included in the bitstream can be reduced, the encoding amount can be reduced.

[0151] For example, the one or more second parameters may be used to control the amount of change in the values of the chroma samples in the deblocking filtering process.

[0152] Thus, the encoding device can use more than one second parameter to limit the amount of change in the values of chrominance samples in the deblocking process.

[0153] For example, it may also be that each of the more than one second parameter represents an offset for at least one of β and tC which are parameters for controlling the amount of change.

[0154] Thus, the encoding device can control the deblocking process of the chrominance of the image to be processed based on the values of these deblocking parameter offsets.

[0155] For example, it may also be that when the first parameter indicates that the syntax element exists in the bitstream, the circuit performs processing using the syntax element. In this case, it may also be that the syntax element is a quantization offset for controlling the quantization parameter of the chrominance component, and the circuit uses the quantization offset to perform quantization processing.

[0156] Thus, when the syntax element exists in the bitstream, the encoding device can perform quantization processing using the quantization offset.

[0157] In addition, a decoding device according to an aspect of the present disclosure includes a circuit and a memory connected to the circuit. During operation, the circuit obtains a first parameter that indicates whether a syntax element associated with a chrominance tool offset exists in the bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking process of chrominance samples of the image to be processed are read from the bitstream, (ii) the deblocking process is performed using the one or more second parameters, and (iii) the image to be processed is decoded. When the first parameter indicates that the syntax element does not exist in the bitstream, the one or more second parameters are not read from the bitstream and the image to be processed is decoded.

[0158] Thus, by commonly using the first parameter between the processing using the syntax element associated with the chrominance tool offset and the deblocking process, the decoding device can reduce the processing related to the determination of whether one or more second parameters should be read, and reduce the storage amount.

[0159] For example, it may also be that the first parameter is pps_chroma_tool_offsets_present_flag set in the picture parameter set.

[0160] Thus, the decoding device can, for example, determine whether one or more second parameters exist in the bitstream without parsing a slice header more than the picture parameter set.

[0161] For example, it may also be that the chrominance sample represents the Cb component and the Cr component in the YCbCr color space.

[0162] Thus, since the decoding device has information on the luminance (Y component) and two chrominances (Cb component and Cr component) of the processing target image, it can reduce the information related to chrominance while suppressing the degradation of subjective image quality. Therefore, since the amount of information to be decoded from the bitstream can be reduced, the decoding amount can be reduced.

[0163] For example, it may also be that the one or more second parameters are used to control the change amount of the value of the chrominance sample in the deblocking filter processing.

[0164] Thus, the decoding device can use one or more second parameters to limit the change amount of the value of the chrominance sample in the deblocking filter processing.

[0165] For example, it may also be that the one or more second parameters respectively represent an offset for at least one of β and tC which are parameters for controlling the change amount.

[0166] Thus, the decoding device can control the deblocking filter processing of the chrominance of the processing target image based on the offset values of these parameters.

[0167] For example, it may also be that the circuit performs processing using the syntax element when the first parameter indicates that the syntax element exists in the bitstream. In this case, it may also be that the syntax element is a quantization offset for controlling the quantization parameter of the chrominance component, and the circuit uses the quantization offset to perform inverse quantization processing.

[0168] Thus, when a syntax element exists in the bitstream, the decoding device can perform inverse quantization processing using the quantization offset.

[0169] In addition, in an encoding method according to an aspect of the present disclosure, a first parameter is stored in the bitstream, the first parameter indicating whether a syntax element associated with a chrominance tool offset exists in the bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking filter processing of the chrominance samples of the processing target image are stored in the bitstream, (ii) the deblocking filter processing is performed using the one or more second parameters, and (iii) the processing target image is encoded. When the first parameter indicates that the syntax element does not exist in the bitstream, the one or more second parameters are not stored in the bitstream and the processing target image is encoded.

[0170] Thus, the apparatus for performing the encoding method can reduce the processing amount while reducing the size of the bitstream by commonly using a first parameter between the process using the syntax element associated with the chrominance tool offset and the deblocking filter process.

[0171] In addition, a decoding method according to an aspect of the present disclosure obtains a first parameter indicating whether a syntax element associated with a chrominance tool offset exists in a bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking filter process for chrominance samples of a processing target image are read out from the bitstream, (ii) the deblocking filter process is performed using the one or more second parameters, and (iii) the processing target image is decoded. When the first parameter indicates that the syntax element does not exist in the bitstream, the one or more second parameters are not read out from the bitstream and the processing target image is decoded.

[0172] Thus, the apparatus for performing the decoding method can reduce the process related to the determination of whether one or more second parameters should be read out and reduce the storage amount by commonly using the first parameter between the process using the syntax element associated with the chrominance tool offset and the deblocking filter process.

[0173] [Definition of Terms]

[0174] As an example, each term can be defined as follows.

[0175] (1) Image

[0176] Is a unit of data composed of a set of pixels, composed of a picture or a block smaller than a picture, and includes a still image in addition to a moving image.

[0177] (2) Picture

[0178] Is a processing unit of an image composed of a set of pixels, and is sometimes referred to as a frame or a field.

[0179] (3) Block

[0180] Is a processing unit including a set of a specific number of pixels. As listed in the following examples, the name is not limited. In addition, the shape is not limited. For example, of course, it includes a rectangle composed of M×N pixels, a square composed of M×M pixels, and also includes a triangle, a circle, and other shapes.

[0181] (Examples of Blocks)

[0182] · Slice / Tile / Brick

[0183] · CTU / Superblock / Base Partition Unit

[0184] ·Processing division unit of VPDU / Hardware

[0185] ·CU / Processing block unit / Prediction block unit (PU) / Orthogonal transform block unit (TU) / Cell

[0186] ·Sub-block

[0187] (4) Pixel / Sample

[0188] It is the point that is the smallest unit constituting an image, and includes not only pixels at integer positions but also pixels at fractional positions generated based on pixels at integer positions.

[0189] (5) Pixel value / Sample value

[0190] It is the inherent value that a pixel has, and of course includes luminance value, chrominance difference value, grayscale of RGB, and also includes depth value, or binary values of 0 and 1.

[0191] (6) Flag

[0192] In addition to 1 bit, it also includes cases of multiple bits. For example, it can also be a parameter or index of 2 bits or more. In addition, not only binary-valued two values are used, but also multi-values using other number systems can be used.

[0193] (7) Signal

[0194] It is symbolized and encoded for transmitting information, and in addition to discrete digital signals, it also includes analog signals taking continuous values.

[0195] (8) Stream / Bitstream

[0196] It refers to a data string of digital data or a stream of digital data. In addition to 1 stream, the stream / bitstream can also be divided into multiple layers and composed of multiple streams. In addition, in addition to the case of transmission on a single transmission path through serial communication, it also includes the case of transmission through packet communication in multiple transmission paths.

[0197] (9) Difference / Differential

[0198] In the case of a scalar, in addition to the simple difference (x - y), as long as it includes the operation of difference, it includes the absolute value of the difference (|x - y|), the square difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a and b are constants), and the offset difference (x - y + a: a is the offset).

[0199] (10) Sum

[0200] In the case of a scalar, in addition to the simple sum (x + y), any operation involving a sum is acceptable, including the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted sum (ax + by: a and b are constants), and the offset sum (x + y + a: a is the offset).

[0201] (11) Based on

[0202] This also includes cases other than adding elements that are the basis. In addition, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0203] (12) Used, using

[0204] This also includes cases other than adding elements that are the object being used. In addition, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via an intermediate result.

[0205] (13) Prohibit, forbid

[0206] This can also be referred to as not allowing. In addition, not prohibiting or allowing does not necessarily imply an obligation.

[0207] (14) Limit, restriction / restrict / restricted

[0208] This can also be referred to as not allowing. In addition, not prohibiting or allowing does not necessarily imply an obligation. Also, as long as a part is prohibited in terms of quantity or quality, it includes cases where it is completely prohibited.

[0209] (15) Chroma

[0210] It is an adjective represented by the notations Cb and Cr that specifies one of two color difference signals associated with the primary colors for a sample arrangement or a single sample representation. Also, the term chrominance can be used instead of the term chroma.

[0211] (16) Luma

[0212] It is an adjective represented by the notation or subscript Y or L that specifies a monochrome signal associated with the primary colors for a sample arrangement or a single sample representation. The term luminance can also be used instead of the term luma.

[0213] [Regarding the explanations in the description]

[0214] In the accompanying drawings, the same reference numerals denote the same or similar components. In addition, the dimensions and relative positions of the components in the drawings are not necessarily drawn to a certain scale.

[0215] Hereinafter, embodiments will be specifically described with reference to the accompanying drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and do not limit the meaning of the claims.

[0216] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processes and / or structures described in various aspects of the present disclosure. The processes and / or structures can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processes and / or structures applied to the embodiments, any of the following can be performed.

[0217] (1) For a certain one of the multiple components of the encoding device or decoding device of the embodiment described in each aspect of the present disclosure, it can be replaced with other components described in a certain one of the aspects of the present disclosure, or they can be combined;

[0218] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can be made to the functions or processes performed by a part of the multiple components of the encoding device or decoding device. For example, replace any function or process with other functions or processes described in a certain one of the aspects of the present disclosure, or combine them;

[0219] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can be made to a part of the multiple processes included in the method. For example, replace any process in the method with other processes described in a certain one of the aspects of the present disclosure, or combine them;

[0220] (4) A part of the multiple components constituting the encoding device or decoding device of the embodiment can be combined with the components described in a certain one of the aspects of the present disclosure, or can be combined with components having a part of the functions described in a certain one of the aspects of the present disclosure, or can be combined with components that perform a part of the processes performed by the components described in the aspects of the present disclosure;

[0221] (5) A component that is part of the function of the encoding device or decoding device of the embodiment, or a component that processes part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced by a component described in one of the various forms of the present disclosure, a component that is part of the function described in one of the various forms of the present disclosure, or a component that processes part of the processing described in one of the various forms of the present disclosure;

[0222] (6) In the method implemented by the encoding device or decoding device of the embodiment, one of the multiple processes included in the method is replaced by a process described in one of the various forms of the present disclosure or a similar process, or they are combined;

[0223] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various forms of the present disclosure.

[0224] (8) The manner of implementing the processes and / or structures described in the various forms of the present disclosure is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the motion image encoding or motion image decoding disclosed in the embodiment.

[0225] [System Structure]

[0226] Figure 1 It is a schematic diagram showing an example of the structure of the transmission system of the present embodiment.

[0227] The transmission system Trs is a system that transmits the stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs is, for example, as Figure 1 shown, and includes an encoding device 100, a network Nw, and a decoding device 200.

[0228] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding.

[0229] In addition, the original image input to the encoding device 100 before being encoded is also referred to as the original image, original signal, or original sample. Additionally, the image can be a moving image or a still image. Furthermore, an image is a superordinate concept of sequences, pictures, blocks, etc., and is not restricted by spatial and temporal regions unless otherwise specified. Also, an image is composed of an arrangement of pixels or pixel values, and the signal or pixel value representing the image is also called a sample. Moreover, a stream can be referred to as a bitstream, encoded bitstream, compressed bitstream, or encoded signal. Further, the encoding device can also be called an image encoding device or a moving image encoding device, and the encoding method of the encoding device 100 can also be called an encoding method, image encoding method, or moving image encoding method.

[0230] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Additionally, the network Nw can also be replaced by a storage medium that records the stream such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0231] The decoding device 200 generates a decoded image, such as an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method of the encoding device 100.

[0232] Moreover, the decoding device can also be called an image decoding device or a moving image decoding device, and the decoding method of the decoding device 200 can also be called a decoding method, image decoding method, or moving image decoding method.

[0233] [Data Structure]

[0234] Figure 2 is a diagram showing an example of the hierarchical structure of the data in the stream. The stream includes, for example, a video sequence. This video sequence, for example, as shown in (a) of Figure 2 contains a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and multiple pictures.

[0235] The VPS includes coding parameters common to multiple layers in a moving image composed of multiple layers, and coding parameters associated with the multiple layers or each layer included in the moving image.

[0236] The SPS includes parameters used for a sequence, that is, coding parameters that the decoding device 200 refers to for decoding the sequence. For example, the coding parameters can also represent the width or height of a picture. Additionally, there can be multiple SPSs.

[0237] The PPS includes parameters used for a picture, that is, coding parameters that the decoding device 200 refers to for decoding each picture in the sequence. For example, the coding parameters can also include a reference value of the quantization width used in decoding the picture and a flag indicating the application of weighted prediction. In addition, there can be multiple PPSs. Furthermore, the SPS and PPS are sometimes simply referred to as parameter sets.

[0238] As Figure 2 shown in (b) of [], a picture can include a picture header and one or more slices. The picture header includes coding parameters that the decoding device 200 refers to for decoding the one or more slices.

[0239] As Figure 2 shown in (c) of [], a slice includes a slice header and one or more tiles. The slice header includes coding parameters that the decoding device 200 refers to for decoding the one or more tiles.

[0240] As Figure 2 shown in (d) of [], a tile includes one or more CTUs (Coding Tree Units).

[0241] In addition, a picture may not include slices and include a tile group instead of the slices. In this case, the tile group includes one or more tiles. Additionally, slices can be included in a tile.

[0242] The CTU is also referred to as a super block or a basic segmentation unit. As Figure 2 shown in (e) of [], such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes coding parameters that the decoding device 200 refers to for decoding the one or more CUs.

[0243] A CU can also be divided into multiple smaller CUs. Additionally, as Figure 2As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual described later. In addition, the CU is basically the same as the PU (Prediction Unit) and the TU (Transform Unit). However, for example, in the SBT described later, it may also include a plurality of TUs smaller than the CU. In addition, the CU may also process each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. The VPDU is, for example, a fixed unit that can be processed in one stage during pipeline processing in hardware.

[0244] In addition, the stream may not have Figure 2 any part of the hierarchies shown. In addition, the order of these hierarchies can be swapped, and any one hierarchy can be replaced with another hierarchy. In addition, the picture that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is called the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded, and if the processing is decoding, the current picture is synonymous with the picture to be decoded. In addition, a block such as a CU or a CU that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded, and if the processing is decoding, the current block is synonymous with the block to be decoded.

[0245] [Structural Slices / Tiles of Picture]

[0246] For parallel decoding of a picture, the picture may sometimes be composed of slices or tiles.

[0247] A slice is the basic encoding unit that constitutes a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs.

[0248] Figure 3This is a diagram showing an example of the structure of a slice. For example, the picture contains 11×8 CTUs and is divided into 4 slices (Slice 1 - 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU within the picture belongs to a certain slice. The shape of the slice is the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can be anywhere among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. Additionally, the slice contains a slice header and encoded data. In the slice header, features of the slice such as the CTU address at the start of the slice and the slice type can also be described.

[0249] A tile is the unit of the rectangular area that makes up the picture. Numbers called TileIds can also be assigned to each tile in the raster scan order.

[0250] Figure 4 This is a diagram showing an example of the structure of a tile. For example, the picture contains 11×8 CTUs and is divided into 4 rectangular area tiles (Tile 1 - 4). When using tiles, the processing order of the CTUs is changed compared to the case without using tiles. Without using tiles, multiple CTUs within the picture are processed in the raster scan order, for example. When using tiles, in each of the multiple tiles, at least 1 CTU is processed in the raster scan order, for example. For example, as Figure 4 shown, the processing order of the multiple CTUs contained in Tile 1 is the order from the left end of the first column of Tile 1 towards the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 towards the right end of the second column of Tile 1.

[0251] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.

[0252] Furthermore, the picture can also be composed of tile set units. A tile set can contain more than one tile group and can also contain more than one tile. The picture can be composed of only one of the tile set, tile group, and tile. For example, the order of scanning multiple tiles in the raster order for each tile set is set as the basic encoding order of the tiles. A set of one or more tiles that are consecutive in the basic encoding order within each tile set is set as a tile group. Such a picture can also be composed of the later-described division unit 102 (refer to Figure 7 ).

[0253] [Scalable Coding]

[0254] Figure 5 and Figure 6This is a diagram showing an example of the structure of a scalable stream.

[0255] As Figure 5 shown, the encoding device 100 can perform encoding by dividing multiple pictures into a certain layer among multiple layers, thereby generating a temporally / spatially scalable stream. For example, the encoding device 100 realizes the scalability where the enhancement layer exists above the base layer by encoding pictures for each layer. The encoding of each such picture is called scalable encoding. Thus, the decoding device 200 can switch the picture quality of the image displayed by decoding this stream. That is, the decoding device 200 determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication band. As a result, the decoding device 200 can freely switch the same content to a low-resolution content and a high-resolution content for decoding. For example, a user of this stream, while on the move, uses a smartphone to view and listen to a moving image of this stream halfway, and after returning home, uses a device such as an Internet TV to view and listen to the subsequent part of this moving image. In addition, decoding devices 200 with the same or different performances are respectively assembled in the above-mentioned smartphone and device. In this case, if the device decodes to the upper layer in this stream, the user can view and listen to a high-definition moving image after returning home. Thus, the encoding device 100 does not need to generate multiple streams with the same content but different picture qualities, and can reduce the processing load.

[0256] Furthermore, the enhancement layer can also include meta-information such as based on the statistical information of the image. It can also be that the decoding device 200 generates a high-definition moving image by super-resolution of the pictures in the base layer based on the meta-information. Super-resolution can be either improving the signal-to-noise ratio at the same resolution or expanding the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least squares operation used in the super-resolution process, etc.

[0257] Alternatively, the picture can also be divided into tiles, etc. according to the meaning of each object, etc. within the picture. In this case, the decoding device 200 can decode only a partial area in the picture by selecting the tile to be decoded. Moreover, the attributes of the object (person, car, ball, etc.) and the position within the picture (coordinate position in the same picture, etc.) can be saved as meta-information. In this case, the decoding device 200 can determine the position of the desired object based on the meta-information and decide on the tile containing the object. For example, as Figure 6 shown, a data storage structure different from the image data, such as SEI in HEVC, can also be used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.

[0258] In addition, the meta-information may be saved in units composed of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding device 200 can obtain the time when a specific person appears in the moving image, etc. By using this time and the information of the picture unit, the picture in which the target exists and the position of the target in that picture can be determined.

[0259] [Encoding device]

[0260] Next, the encoding device 100 of the embodiment will be described. Figure 7 FIG. is a block diagram showing an example of the functional structure of the encoding device 100 of the embodiment. The encoding device 100 encodes an image in units of blocks.

[0261] As Figure 7 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. In addition, the intra prediction unit 124 and the inter prediction unit 126 are each configured as part of a prediction processing unit.

[0262] [Installation example of the encoding device]

[0263] Figure 8 FIG. is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 7 as shown, the multiple components of the encoding device 100 are implemented by the Figure 8 shown processor a1 and memory a2.

[0264] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes an image. The processor a1 may also be a processor such as a CPU. In addition, the processor a1 may also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 may also function as Figure 7 the multiple components of the encoding device 100 shown, excluding the components for storing information.

[0265] Memory a2 is a dedicated or general-purpose memory that stores information for the processor a1 to encode an image. Memory a2 can be either an electronic circuit or connected to the processor a1. Additionally, memory a2 can be included in the processor a1. Moreover, memory a2 can also be an aggregate of multiple electronic circuits. Additionally, memory a2 can be a magnetic disk, an optical disc, etc., or can be represented as a storage or a recording medium, etc. Additionally, memory a2 can be either a non-volatile memory or a volatile memory.

[0266] For example, memory a2 can store the encoded image or the stream corresponding to the encoded image. Additionally, a program for the processor a1 to encode an image can also be stored in memory a2.

[0267] Additionally, for example, memory a2 can also serve as Figure 7 the component for storing information among the multiple components of the encoding device 100 shown. Specifically, memory a2 can serve as Figure 7 the block memory 118 and the frame memory 122 shown. More specifically, reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.) can be stored in memory a2.

[0268] Additionally, in the encoding device 100, not all of the multiple components shown Figure 7 need to be installed, and not all of the above-mentioned multiple processes need to be performed. Figure 7 A part of the multiple components shown

[0269] can be included in other devices, or a part of the above-mentioned multiple processes can be performed by other devices.

[0270] [Overall Process of Encoding Processing]

[0271] Figure 9 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0272] First, the segmentation unit 102 of the encoding device 100 divides the pictures included in the original image into multiple blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block with the fixed size into multiple blocks that constitute the selected segmentation pattern. Then, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 for each of the multiple blocks.

[0273] The prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126, and the prediction control unit 128 generate a prediction image of the current block (step Sa_3). In addition, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.

[0274] Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual is also referred to as a prediction error.

[0275] Next, the transform unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by transforming and quantizing the prediction image (step Sa_5).

[0276] Next, the entropy encoding unit 110 generates a bitstream by encoding the plurality of quantization coefficients and prediction parameters related to the generation of the prediction image (specifically, entropy encoding) (step Sa_6).

[0277] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore the prediction residual by inverse quantizing and inverse transforming the plurality of quantization coefficients (step Sa_7).

[0278] Next, the addition unit 116 reconstructs the current block by adding the restored prediction residual to the prediction image (step Sa_8). Thus, a reconstructed image is generated. In addition, the reconstructed image is also referred to as a reconstructed block. In particular, the reconstructed image generated by the encoding apparatus 100 is also referred to as a locally decoded block or a locally decoded image.

[0279] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).

[0280] Then, the encoding apparatus 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and when it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.

[0281] In addition, in the above example, the encoding apparatus 100 selects one segmentation style for blocks of a fixed size and encodes each block according to the segmentation style, but each block may also be encoded according to each of a plurality of segmentation styles. In this case, the encoding apparatus 100 may evaluate the cost for each of the plurality of segmentation styles, and for example, may select the bitstream obtained by encoding according to the segmentation style with the minimum cost as the finally output bitstream.

[0282] In addition, the processing of these steps Sa_1 to Sa_10 may be sequentially performed by the encoding apparatus 100, a part of the plurality of processes among these processes may be performed in parallel, or the order may be changed.

[0283] The encoding process of such an encoding device 100 uses hybrid encoding that combines predictive encoding and transform encoding. In addition, the predictive encoding is performed through an encoding loop, which is composed of a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. That is, the prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126 constitutes a part of the encoding loop.

[0284] [Segmentation Unit]

[0285] The segmentation unit 102 divides each picture included in the original image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the segmentation unit 102 first divides the picture into blocks of a fixed size (e.g., 128×128 pixels). Such blocks of the fixed size are sometimes referred to as coding tree units (CTUs). And, the segmentation unit 102 divides each block of the fixed size into blocks of a variable size (e.g., 64×64 pixels or less) based on recursive quadtree and / or binary tree block segmentation. That is, the segmentation unit 102 selects a segmentation pattern. Such blocks of the variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture can also be used as the processing unit for CUs, PUs, or TUs.

[0286] Figure 10 is a diagram showing an example of the block segmentation of the embodiment. In Figure 10 the solid lines represent the block boundaries based on quadtree block segmentation, and the dashed lines represent the block boundaries based on binary tree block segmentation.

[0287] Here, block 10 is a square block of 128×128 pixels. This block 10 is first divided into 4 square blocks of 64×64 pixels (quadtree block segmentation).

[0288] The upper left 64×64 pixel square block is further vertically divided into 2 rectangular blocks each composed of 32×64 pixels, and the left 32×64 pixel rectangular block is further vertically divided into 2 rectangular blocks each composed of 16×64 pixels (binary tree block segmentation). As a result, the upper left 64×64 pixel square block is divided into 2 rectangular blocks 11 and 12 of 16×64 pixels and a rectangular block 13 of 32×64 pixels.

[0289] The upper right 64×64 pixel square block is horizontally divided into 2 rectangular blocks 14 and 15 each composed of 64×32 pixels (binary tree block segmentation).

[0290] The 64×64 pixel square block in the lower left is divided into 4 square blocks (quad-tree block division), each consisting of 32×32 pixels. The upper left and lower right blocks among the 4 square blocks, each consisting of 32×32 pixels, are further divided. The upper left 32×32 pixel square block is vertically divided into 2 rectangular blocks, each consisting of 16×32 pixels, and the right rectangular block consisting of 16×32 pixels is further horizontally divided into 2 square blocks, each consisting of 16×16 pixels (binary tree block division). The lower right 32×32 pixel square block is horizontally divided into 2 rectangular blocks, each consisting of 32×16 pixels (binary tree block division). As a result, the 64×64 pixel square block in the lower left is divided into 16 rectangular blocks of 16×32 pixels, 2 square blocks 17 and 18 of 16×16 pixels each, 2 square blocks 19 and 20 of 32×32 pixels each, and 2 rectangular blocks 21 and 22 of 32×16 pixels each.

[0291] The block 23 consisting of 64×64 pixels in the lower right is not divided.

[0292] As described above, in Figure 10 , the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.

[0293] In addition, in Figure 10 , 1 block is divided into 4 or 2 blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, 1 block can also be divided into 3 blocks (ternary tree division). Divisions including such ternary tree division are sometimes called MBT (multi type tree) division.

[0294] Figure 11 is a diagram showing an example of the functional structure of the division unit 102. As Figure 11 shown, the division unit 102 may also include a block division determination unit 102a. As an example, the block division determination unit 102a may perform the following processing.

[0295] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the above division pattern based on this block information. The division unit 102 divides the original image according to this division pattern and outputs one or more blocks obtained by this division to the subtraction unit 104.

[0296] In addition, the block segmentation determination unit 102a outputs, for example, parameters indicating the above-described segmentation patterns to the transformation unit 106, the inverse transformation unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy encoding unit 110. The transformation unit 106 can transform the prediction residual based on the parameters, and the intra prediction unit 124 and the inter prediction unit 126 can generate a prediction image based on the parameters. In addition, the entropy encoding unit 110 can also perform entropy encoding on the parameters.

[0297] As an example, parameters related to the segmentation pattern can also be written into the stream as follows.

[0298] Figure 12 FIG. is an example showing a segmentation pattern. Examples of the segmentation pattern include: four-way split (QT), in which a block is split into two in the horizontal and vertical directions respectively; three-way split (HT or VT), in which a block is split in the same direction at a ratio of 1:2:1; two-way split (HB or VB), in which a block is split in the same direction at a ratio of 1:1; and no split (NS).

[0299] In addition, in the case of four-way split and no split, the segmentation pattern does not have a block segmentation direction, and in the case of two-way split and three-way split, the segmentation pattern has segmentation direction information.

[0300] Figure 13A and Figure 13B FIG. is an example of a syntax tree showing a segmentation pattern. In the example of Figure 13A , first, there is information indicating whether to perform segmentation (S: Split flag), and then there is information indicating whether to perform four-way split (QT: QT flag). Next, there is information indicating whether to perform three-way split or two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the segmentation direction (Ver: Vertical flag or Hor: Horizontal flag). In addition, for each of one or more blocks obtained by such segmentation based on the segmentation pattern, the same process can be further repeatedly applied for segmentation. That is, as an example, it is also possible to recursively perform determination of whether to perform segmentation, whether to perform four-way split, whether the segmentation method is horizontal or vertical, and whether to perform three-way split or two-way split, and encode the determination results implemented into the stream in the encoding order disclosed in the syntax tree shown in Figure 13A .

[0301] In addition, in the syntax tree shown in Figure 13A , these information are arranged in the order of S, QT, TT, Ver, but they can also be arranged in the order of S, QT, Ver, BT. That is, in Figure 13BIn the example, first, there is information indicating whether to perform splitting (S: Split flag), then there is information indicating whether to perform four-way splitting (QT: QT flag). Next, there is information indicating the splitting direction (Ver: Vertical flag or Hor: Horizontal flag), and finally there is information indicating whether to perform two-way splitting or three-way splitting (BT: BT flag or TT: TT flag).

[0302] In addition, the splitting pattern described here is an example. A splitting pattern other than the described one can be used, or only a part of the described splitting pattern can be used.

[0303] [Subtraction unit]

[0304] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units input from and split by the splitting unit 102. That is, the subtraction unit 104 calculates the prediction residual of the current block. And the subtraction unit 104 outputs the calculated prediction residual to the transformation unit 106.

[0305] The original image is an input signal of the encoding device 100, for example, a signal representing an image of each picture constituting a moving image (for example, a luma signal and two chroma signals).

[0306] [Transformation unit]

[0307] The transformation unit 106 transforms the prediction residual in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction residual in the spatial domain.

[0308] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types, use a transformation basis function corresponding to the selected transformation type, and transform the prediction residual into transformation coefficients. Such a transformation is called EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform) in some cases. In addition, the transformation basis function is sometimes simply referred to as a basis.

[0309] The multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. In addition, these transformation types can be described as DCT2, DCT5, DCT8, DST1, and DST7 respectively. Figure 14is a table representing transform basis functions corresponding to respective transform types. In Figure 14 , N represents the number of input pixels. The selection of a transform type from among these multiple transform types can depend, for example, on the type of prediction (intra prediction, inter prediction, etc.) or on the intra prediction mode.

[0310] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. In addition, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0311] In addition, the transform unit 106 may also perform a re - transform on the transform coefficients (i.e., the transform result). Such a re - transform may be a case called AST (adaptive secondary transform) or NSST (non - separable secondary transform). For example, the transform unit 106 performs a re - transform for each sub - block (e.g., a 4×4 pixel sub - block) included in a block of transform coefficients corresponding to an intra - prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are generally signaled at the CU level. In addition, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0312] In the transform unit 106, a separable transform and a non - separable transform may also be applied. A separable transform is a method of performing multiple transforms separately for each direction according to the number of input dimensions, and a non - separable transform is a method of treating two or more dimensions as one dimension and performing a transform together when the input is multi - dimensional.

[0313] For example, as an example of a non - separable transform, when the input is a 4×4 pixel block, it can be regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.

[0314] In addition, in a further example of a non - separable transform, after regarding a 4×4 pixel input block as a permutation having 16 elements, a transform of performing multiple Givens rotations on this permutation (Hypercube Givens Transform) may also be performed.

[0315] In the transformation in the transformation unit 106, the type of transformation of the transformation basis function to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform).

[0316] Figure 15 It is a diagram showing an example of SVT.

[0317] In SVT, as Figure 15 shown, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The transformation type can be set for each region. For example, DST7 and DCT8 are used. For example, for the region at position 0 among the two regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used. Or, for the region at position 1 among the two regions, DST7 is used. Similarly, for the region at position 0 among the two regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 are used. Or, for the region at position 1 among the two regions, DST7 is used. In such a Figure 15 shown example, only one of the two regions within the CU is transformed, and the other is not transformed, but it is also possible to transform the two regions separately. In addition, the splitting method is not limited to bisecting, but can also be quartering. Furthermore, it can be more flexible, encoding the information indicating the splitting method and performing signaling etc. in the same way as CU splitting. Additionally, SVT is sometimes also referred to as SBT (Sub-block Transform).

[0318] The aforementioned AMT and EMT can also be referred to as MTS (Multiple Transform Selection). In the case of applying MTS, transform types such as DST7 or DCT8 can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. On the other hand, as a process of selecting the transform type used in the orthogonal transform based on the shape of the CU without encoding the index information, there is a process called IMTS (Implicit MTS). In the case of applying IMTS, for example, if the shape of the CU is rectangular, DST7 is used on the short side of the rectangle and DCT2 is used on the long side for orthogonal transformation respectively. Additionally, for example, in the case where the shape of the CU is square, if MTS is effective within the sequence, DCT2 is used for orthogonal transformation, and if MTS is ineffective, DST7 is used for orthogonal transformation. DCT2 and DST7 are just examples, and other transform types can be used, or different combinations of the used transform types can be set. IMTS can be used only in blocks for intra prediction, or can be used together in blocks for intra prediction and blocks for inter prediction.

[0319] As described above, as a selection process for selectively switching the transform type used in the orthogonal transform, the three processes of MTS, SBT, and IMTS have been described. However, all three selection processes can be effective, or only some of the selection processes can be selectively made effective. Regarding whether each selection process is effective, it can be identified by flag information in headers such as SPS. For example, if all three selection processes are effective, one is selected from the three selection processes for orthogonal transformation in units of CUs. Additionally, as long as the selection process for selectively switching the transform type can achieve at least one of the following four functions [1] to [4], a selection process different from the above three selection processes can be used, or the above three selection processes can be replaced with other processes respectively. Function [1] is the function of performing orthogonal transformation on the entire range within the CU and encoding the information indicating the transform type used in the transformation. Function [2] is the function of performing orthogonal transformation on the entire range of the CU, determining the transform type based on a specified rule without encoding the information indicating the transform type. Function [3] is the function of performing orthogonal transformation on a partial region of the CU and encoding the information indicating the transform type used in the transformation. Function [4] is the function of performing orthogonal transformation on a partial region of the CU, not encoding the information indicating the transform type used in the transformation, and determining the transform type based on a specified rule, etc.

[0320] In addition, the presence or absence of the application of MTS, IMTS, and SBT respectively can also be determined for each processing unit. For example, it can be determined for each sequence unit, picture unit, tile unit, slice unit, CTU unit, or CU unit.

[0321] In addition, the tool for selectively switching the transformation type in the present disclosure can also be renamed as a method for adaptively selecting the basis used in the transformation process, a selection process, or a process for selecting a basis. In addition, the tool for selectively switching the transformation type can also be renamed as a mode for adaptively selecting the transformation type.

[0322] Figure 16 It is a flowchart showing an example of the process performed by the transformation unit 106.

[0323] For example, the transformation unit 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined to perform an orthogonal transformation (Yes in step St_1), the transformation unit 106 selects the transformation type used for the orthogonal transformation from among multiple transformation types (step St_2). Next, the transformation unit 106 performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Then, the transformation unit 106 outputs the information indicating the selected transformation type to the entropy encoding unit 110 to encode this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transformation (No in step St_1), the transformation unit 106 outputs the information indicating that no orthogonal transformation is performed to the entropy encoding unit 110 to encode this information (step St_5). In addition, the determination of whether to perform an orthogonal transformation in step St_1 can be made, for example, based on the size of the transformation block, the prediction mode applied to the CU, etc. In addition, the information indicating the transformation type used for the orthogonal transformation may not be encoded, and an orthogonal transformation may be performed using a pre-specified transformation type.

[0324] Figure 17 It is a flowchart showing another example of the process performed by the transformation unit 106. In addition, Figure 17 The example shown is the same as the example shown in Figure 16 and is an example of an orthogonal transformation in the case of applying the method of selectively switching the transformation type used for the orthogonal transformation.

[0325] As an example, the first transformation type group may include DCT2, DST7, and DCT8. In addition, as an example, the second transformation type group may include DCT2. In addition, the transformation types included in the first transformation type group and the second transformation type group may partially overlap or may be all different transformation types.

[0326] Specifically, the transformation unit 106 determines whether the transformation size is equal to or less than a specified value (step Su_1). Here, when it is determined that the size is equal to or less than the specified value (yes in step Su_1), the transformation unit 106 orthogonally transforms the prediction residual of the current block using the transformation types included in the first transformation type group (step Su_2). Then, the transformation unit 106 encodes the information indicating which transformation type among one or more transformation types included in the first transformation type group is used by outputting the information to the entropy encoding unit 110 (step Su_3). On the other hand, when it is determined that the transformation size is not equal to or less than the specified value (no in step Su_1), the transformation unit 106 orthogonally transforms the prediction residual of the current block using the second transformation type group (step Su_4).

[0327] In step Su_3, the information indicating the transformation type used for the orthogonal transformation may be information indicating a combination of the transformation type applied to the vertical direction of the current block and the transformation type applied to the horizontal direction. In addition, the first transformation type group may include only one transformation type, and the information indicating the transformation type used for the orthogonal transformation may not be encoded. The second transformation type group may include multiple transformation types, and the information indicating the transformation type used in the orthogonal transformation among one or more transformation types included in the second transformation type group may also be encoded.

[0328] In addition, the transformation type may be determined based only on the transformation size. Furthermore, if the process of determining the transformation type used for the orthogonal transformation is based on the transformation size, it is not limited to the determination of whether the transformation size is equal to or less than the specified value.

[0329] [Quantization unit]

[0330] The quantization unit 108 quantizes the transformation coefficients output from the transformation unit 106. Specifically, the quantization unit 108 scans the multiple transformation coefficients of the current block in a specified scan order, and quantizes the transformation coefficient based on the quantization parameter (QP) corresponding to the scanned transformation coefficient. Then, the quantization unit 108 outputs the quantized multiple transformation coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.

[0331] The specified scan order is the order for quantization / inverse quantization of the transformation coefficients. For example, the specified scan order is defined by the ascending order of frequencies (from low frequency to high frequency) or the descending order (from high frequency to low frequency).

[0332] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the error (quantization error) of the quantization coefficient increases.

[0333] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transformation sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing values sampled at a predetermined interval by associating them with predetermined levels, and in this technical field, expressions such as rounding, truncating, or scaling are sometimes used.

[0334] As methods of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device 100 side and a method of using a default quantization matrix (default matrix). On the encoding device 100 side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix. In addition, instead of directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix used in the quantization of the current block can be generated based on the default quantization matrix or the encoded quantization matrix.

[0335] On the other hand, there is also a method of performing quantization in such a way that the coefficients of high-frequency components and the coefficients of low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix (flat matrix) in which all coefficients are the same value.

[0336] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level.

[0337] When using a quantization matrix, the quantization unit 108 scales, for example, the quantization width obtained according to quantization parameters and the like for each transform coefficient using the value of the quantization matrix. The quantization process without using a quantization matrix can also be a process of quantizing the transform coefficient based on the quantization width obtained according to quantization parameters and the like. In addition, in the quantization process without using a quantization matrix, a predetermined value common to all transform coefficients within the block can also be multiplied by the quantization width.

[0338] Figure 18 It is a block diagram showing an example of the functional structure of the quantization unit 108.

[0339] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108a, a prediction quantization parameter generation unit 108b, a quantization parameter generation unit 108c, a quantization parameter storage unit 108d, and a quantization processing unit 108e.

[0340] Figure 19 It is a flowchart showing an example of the quantization performed by the quantization unit 108.

[0341] As an example, the quantization unit 108 can be based onFigure 19 The flowchart shown performs quantization for each CU. Specifically, the quantization parameter generation unit 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generation unit 108c generates quantization parameters for the current block (step Sv_2) and saves the quantization parameters to the quantization parameter storage unit 108d (step Sv_3).

[0342] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). Then, the predicted quantization parameter generation unit 108b obtains quantization parameters of a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). The predicted quantization parameter generation unit 108b generates predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generation unit 108c and the predicted quantization parameters of the current block generated by the predicted quantization parameter generation unit 108b (step Sv_7). By calculating this difference, differential quantization parameters are generated. The differential quantization parameter generation unit 108a outputs the differential quantization parameters to the entropy encoding unit 110, thereby encoding the differential quantization parameters (step Sv_8).

[0343] In addition, the differential quantization parameters can also be encoded at the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameters can be generated using the initial values of the quantization parameters and the differential quantization parameters.

[0344] In addition, the quantization unit 108 can include multiple quantizers, and dependent quantization that quantizes the transform coefficients using a quantization method selected from multiple quantization methods can also be applied.

[0345] [Entropy Encoding Unit]

[0346] Figure 20 is a block diagram showing an example of the functional structure of the entropy encoding unit 110.

[0347] The entropy encoding unit 110 performs entropy encoding on the quantized coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130, thereby generating a stream. In this entropy encoding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy encoding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization that transforms multi-value signals such as quantized coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, etc. The context control unit 110b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. In the method of deriving this context value, for example, there are bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic coding unit 110c uses the derived context value to perform arithmetic coding on the binarized signal.

[0348] Figure 21 It is a diagram showing the process of CABAC in the entropy encoding unit 110.

[0349] First, in the CABAC in the entropy encoding unit 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding unit 110c and setting of the initial context value are performed. Then, the binarization unit 110a and the binary arithmetic coding unit 110c, for example, sequentially perform binarization and arithmetic coding on the multiple quantized coefficients of the CTU. At this time, the context control unit 110b updates the context value each time arithmetic coding is performed. Then, the context control unit 110b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.

[0350] [Inverse Quantization Unit]

[0351] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a specified scan order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.

[0352] [Inverse Transform Unit]

[0353] The inverse transform unit 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.

[0354] In addition, since information is usually lost due to quantization in the restored prediction residual, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual usually contains a quantization error.

[0355] [Addition unit]

[0356] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 and the predicted image input from the prediction control unit 128. As a result, a reconstructed image is generated. Then, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.

[0357] [Block memory]

[0358] The block memory 118 is, for example, a storage unit for storing blocks referred to in intra prediction and that are blocks within the current picture. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.

[0359] [Frame memory]

[0360] The frame memory 122 is, for example, a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.

[0361] [Loop filter unit]

[0362] The loop filter unit 120 performs a loop filter process on the reconstructed image output from the addition unit 116 and outputs the reconstructed image after the filter process to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).

[0363] Figure 22 It is a block diagram showing an example of the functional structure of the loop filter unit 120.

[0364] For example Figure 22As shown, the loop filter unit 120 includes a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 120b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 120c applies the above-described ALF processing to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is a process for improving image quality by reducing ringing (a phenomenon in which pixel values around an edge are deformed in a fluctuating manner) and correcting pixel value deviations. In this SAO processing, for example, there are edge offset processing and band offset processing. In addition, the loop filter unit 120 may not include Figure 22 all the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 120 may also be structured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 22 .

[0365] [Loop Filter Unit > Adaptive Loop Filter]

[0366] In the ALF, a least-squares error filter used to remove coding distortion is employed. For example, for each 2×2 pixel sub-block within the current block, 1 filter selected from multiple filters is used based on the direction and activity of the locality-based gradient.

[0367] Specifically, first, sub-blocks (e.g., 2×2 pixel sub-blocks) are classified into multiple classes (e.g., 15 or 25 classes). The classification of sub-blocks is performed, for example, based on the direction and activity of the gradient. In a specific example, using the gradient direction value D (e.g., 0 to 2 or 0 to 4) and the gradient activity value A (e.g., 0 to 4), the classification value C (e.g., C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes.

[0368] The gradient direction value D is derived, for example, by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantifying the addition result.

[0369] Based on the result of such classification, the filter to be used for the sub-block is determined from among the multiple filters.

[0370] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 23A to 23C It is a diagram showing multiple examples of the shape of the filter used in the ALF. Figure 23A It represents a 5×5 rhombus-shaped filter.Figure 23B Represents a 7×7 rhombus-shaped filter, Figure 23C Represents a 9×9 rhombus-shaped filter. Information indicating the shape of the filter is typically signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).

[0371] The on / off of ALF can also be determined, for example, at the picture level or CU level. For example, regarding luminance, it can be determined at the CU level whether to use ALF, and regarding chrominance difference, it can be determined at the picture level whether to use ALF. Information indicating the on / off of ALF is typically signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of ALF need not be limited to the picture level or CU level and can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).

[0372] Additionally, as described above, one filter is selected from multiple filters and ALF processing is applied to the sub-block. For each of these multiple filters (e.g., up to 15 or 25 filters), the set of coefficients composed of the multiple coefficients used in the filter is typically signaled at the picture level. Additionally, the signaling of the coefficient set need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0373] [Loop Filter>Cross Component Adaptive Loop Filter (or Cross Component Adaptive Loop Filter)]

[0374] Figure 23D Is a diagram showing an example where the Y sample (first component) is used for the CCALF of Cb and the CCALF of Cr (multiple components different from the first component). Figure 23E Is a diagram showing a rhombus-shaped filter.

[0375] One example of CC-ALF is by using a linear rhombus-shaped filter ( Figure 23D , Figure 23E) It operates on the luminance channels of various color difference components. For example, filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled to have variable block sizes and is notified by a context-encoded flag received for each block of samples. The block size and the CC-ALF enable flag are received at the slice level of each color difference component. The syntax and semantics of CC-ALF are provided in the Appendix. In this paper, block sizes of 16x16, 32x32, 64x64, and 128x128 (in color difference samples) are supported.

[0376] [Loop Filter> Combines the Joint Chroma Cross Component Adaptive Loop Filter]

[0377] Figure 23F is a diagram showing an example of JC-CCALF. Figure 23G is a diagram showing an example of the weight_index candidates of JC-CCALF.

[0378] One example of JC-CCALF uses only one CCALF filter, generates one CCALF filter output as the color difference adjustment signal for only one color component, and applies a properly weighted version of the same color difference adjustment signal to other color components. In this way, the complexity of the existing CCALF is approximately halved.

[0379] The weight value is encoded into a sign flag and a weight index. The weight index (denoted as weight_index) is encoded as 3 bits, specifying the magnitude of the JC-CCALF weight JcCcWeight. It cannot be the same as 0. The magnitude of JcCcWeight is determined as follows.

[0380] · When weight_index is 4 or less, JcCcWeight is equal to weight_index >> 2.

[0381] · In other cases, JcCcWeight is equal to 4 / (weight_index - 4).

[0382] The on / off control at the block level for ALF filtering of Cb and Cr is separate. This is the same as CCALF, and two individual sets of on / off control flags at the block level are encoded. Here, different from CCALF, the on / off control block sizes for Cb and Cr are the same, so only one block size variable is encoded.

[0383] [Loop Filter Section > Deblocking Filter]

[0384] In the deblocking filter process, the loop filter section 120 reduces the distortion generated at the block boundary by filtering the block boundary of the reconstructed image.

[0385] Figure 24 It is a block diagram showing an example of the detailed structure of the deblocking filter processing section 120a.

[0386] The deblocking filter processing section 120a includes, for example, a boundary determination section 1201, a filtering determination section 1203, a filtering processing section 1205, a processing determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.

[0387] The boundary determination section 1201 determines whether there are pixels to be deblocked (i.e., target pixels) near the block boundary. Then, the boundary determination section 1201 outputs its determination result to the switch 1202 and the processing determination section 1208.

[0388] When it is determined by the boundary determination section 1201 that target pixels exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1204. On the contrary, when it is determined by the boundary determination section 1201 that target pixels do not exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1206. In addition, the image before filtering processing is an image composed of target pixels and at least one surrounding pixel located around the target pixel.

[0389] The filtering determination section 1203 determines whether to perform deblocking filter processing on the target pixel based on the pixel values of at least one surrounding pixel located around the target pixel. Then, the filtering determination section 1203 outputs the determination result to the switch 1204 and the processing determination section 1208.

[0390] When it is determined by the filtering determination section 1203 that deblocking filter processing is to be performed on the target pixel, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the filtering processing section 1205. On the contrary, when it is determined by the filtering determination section 1203 that deblocking filter processing is not to be performed on the target pixel, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the switch 1206.

[0391] When the image before filtering processing is obtained via the switches 1202 and 1204, the filtering processing section 1205 performs deblocking filter processing with the filtering characteristics determined by the filtering characteristic determination section 1207 on the target pixel. Then, the filtering processing section 1205 outputs the pixel after the filtering processing to the switch 1206.

[0392] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs pixels that have not been deblocking-filtered and pixels that have been deblocking-filtered by the filtering processing unit 1205.

[0393] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processing determination unit 1208 outputs the deblocking-filtered pixels from the switch 1206. In addition, in other cases than the above, the processing determination unit 1208 outputs the pixels that have not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206. In addition, Figure 24 The structure shown is an example of the structure in the deblocking filtering unit 120a, and the deblocking filtering unit 120a may have other structures.

[0394] Figure 25 is a diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to the block boundary.

[0395] In deblocking filtering, for example, using the pixel value and the quantization parameter, either one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Figure 25 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.

[0396] q’0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0397] q’1 = (p0 + q0 + q1 + q2 + 2) / 4

[0398] q’2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0399] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering are filtering coefficients.

[0400] Furthermore, in the deblocking filter processing, the clipping process may also be performed in such a way that the pixel value after the operation does not change when it exceeds the threshold value. In this clipping process, the threshold value determined according to the quantization parameter is used to clip the pixel value after the operation based on the above formula to "the pixel value before the operation ± 2 × the threshold value". Thereby, excessive smoothing can be prevented.

[0401] Figure 26 FIG. is an example of a block boundary for explaining the deblocking filter processing. Figure 27 FIG. is an example showing the BS value.

[0402] The block boundary for performing the deblocking filter processing is, for example, Figure 26 the boundary of the CU, PU, or TU of an 8×8 pixel block shown in FIG. The deblocking filter processing is performed, for example, in units of 4 rows or 4 columns. First, for Figure 26 the blocks P and Q shown in FIG., the Bs (Boundary Strength) value is determined as Figure 27 shown.

[0403] According to Figure 27 the Bs value, it can be determined whether to perform deblocking filter processing with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filter processing for the chrominance signal is performed. When the Bs value is 1 or more and satisfies a specified condition, deblocking filter processing for the luminance signal is performed. In addition, the determination condition of the Bs value is not limited to Figure 27 the condition shown in FIG., and it can also be determined based on other parameters.

[0404] [Prediction unit (intra prediction unit / inter prediction unit / prediction control unit)]

[0405] Figure 28 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the encoding apparatus 100. In addition, as an example, the prediction unit is composed of all or a part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.

[0406] The prediction unit generates a prediction image of the current block (step Sb_1). In addition, in the prediction image, there are, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). Specifically, the prediction unit uses the reconstructed image that has already been obtained by generating a prediction image of another block, generating a prediction residual, generating quantization coefficients, restoring the prediction residual, and adding the prediction images, and generates a prediction image of the current block.

[0407] The reconstructed image can be, for example, an image referring to a reference picture, or an image including the current block, i.e., an encoded block within the current picture (i.e., the other blocks described above). The encoded blocks within the current picture are, for example, adjacent blocks of the current block.

[0408] Figure 29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device 100.

[0409] The prediction unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such prediction methods, the above-described reconstructed image can also be used.

[0410] Next, the prediction unit evaluates the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit evaluates these prediction images by calculating a cost C for each of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c and comparing the costs C of these prediction images. In addition, the cost C is calculated by an equation of the R-D optimization model, such as C = D + λ × R. In this equation, D is the coding distortion of the prediction image and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. Further, R is the bit rate of the stream. In addition, λ is, for example, the undetermined multiplier of Lagrange.

[0411] Next, the prediction unit selects one of the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects a method or mode for obtaining the final prediction image. For example, the prediction unit selects the prediction image with the minimum cost C based on the cost C calculated for these prediction images. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 can also be performed based on the parameters used in the encoding process. The encoding device 100 can signal an information signal for determining the selected prediction image, method, or mode as a stream. This information can be, for example, a flag or the like. Thereby, the decoding device 200 can generate a prediction image based on this information in accordance with the method or mode selected in the encoding device 100. In addition, in Figure 29 the example shown, after generating the prediction images by each method, the prediction unit selects any one of the prediction images. However, before generating these prediction images, the prediction unit can select a method or mode based on the parameters used for the above-described encoding process, and can generate a prediction image according to this method or mode.

[0412] For example, the first mode and the second mode are intra prediction and inter prediction, respectively, and the prediction unit can select the final prediction image for the current block from the prediction images generated according to these prediction modes.

[0413] Figure 30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding apparatus 100.

[0414] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a), and generates a prediction image by inter prediction (step Sd_1b). In addition, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0415] Next, the prediction unit evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above-mentioned cost C can also be used in this evaluation. Then, the prediction unit can select the prediction image that calculates the minimum cost C from the intra prediction image and the inter prediction image as the final prediction image for the current block (step Sd_3). That is, the prediction method or mode for generating the prediction image of the current block is selected.

[0416] [Intra Prediction Unit]

[0417] The intra prediction unit 124 performs intra prediction (also referred to as intra-frame prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction image of the current block (i.e., an intra prediction image). Specifically, the intra prediction unit 124 generates an intra prediction image by performing intra prediction with reference to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control unit 128.

[0418] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.

[0419] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC standard.

[0420] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may include 32-direction prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 31This is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows indicate 33 directions specified by the H.265 / HEVC standard, and the dashed arrows indicate the additional 32 directions (the 2 non-directional prediction modes are not shown in Figure 31 .

[0421] In various installation examples, in the intra prediction of chrominance blocks, the luminance blocks can also be referred to. That is, the chrominance components of the current block can also be predicted based on the luminance component of the current block. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block (for example, called the CCLM mode) can also be added as one of the intra prediction modes of the chrominance block.

[0422] The intra prediction unit 124 can also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes called PDPC (position dependent intraprediction combination). The information indicating whether PDPC is used (for example, called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).

[0423] Figure 32 This is a flowchart showing an example of the processing performed by the intra prediction unit 124.

[0424] The intra prediction unit 124 selects one intra prediction mode from multiple intra prediction modes (step Sw_1). Then, the intra prediction unit 124 generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM consists of, for example, 6 intra prediction modes. Two of the 6 intra prediction modes can be the Planar prediction mode and the DC prediction mode, and the remaining 4 modes can be directional prediction modes. Then, the intra prediction unit 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0425] Here, when it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5), and generates information representing the selected intra prediction mode in the MPM (step Sw_6). In addition, the MPM flag set to 1 and the information representing the intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110.

[0426] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction unit 124 does not set the MPM flag. Then, the intra prediction unit 124 generates information representing the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). In addition, the MPM flag set to 0 and the information representing the intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110. The information representing the intra prediction mode represents any value from 0 to 60, for example.

[0427] [Inter - prediction unit]

[0428] The inter - prediction unit 126 performs inter - prediction (also called inter - picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a predicted image (inter - prediction image). The inter - prediction is performed in units of the current block or the current sub - block within the current block. A sub - block is included in a block and is a unit smaller than the block. The size of the sub - block can be 4x4 pixels, can be 8x8 pixels, or can be other sizes. The size of the sub - block can also be switched in units of slices, bricks, or pictures, etc.

[0429] For example, the inter - prediction unit 126 performs motion search (motion estimation) within the reference picture for the current block or the current sub - block to find the reference block or sub - block that most closely matches the current block or the current sub - block. And, the inter - prediction unit 126 obtains motion information (such as a motion vector) for compensating the motion or change from the reference block or sub - block to the current block or the current sub - block. The inter - prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, thereby generating an inter - prediction image of the current block or the current sub - block. And, the inter - prediction unit 126 outputs the generated inter - prediction image to the prediction control unit 128.

[0430] The motion information used in motion compensation is signaled as the inter - prediction image in various forms. For example, the motion vector can also be signaled. As another example, the difference between the motion vector and the predicted motion vector (motion vector predictor) can also be signaled.

[0431] [List of reference pictures]

[0432] Figure 33 is a diagram showing an example of each reference picture, Figure 34 is a conceptual diagram showing an example of the list of reference pictures. The list of reference pictures is a list showing one or more reference pictures stored in the frame memory 122. In addition, in Figure 33 , the rectangle represents a picture, the arrow represents the reference relationship of the pictures, the horizontal axis represents time, I, P, and B in the rectangle represent an intra-predicted picture, a single-predicted picture, and a bi-predicted picture respectively, and the numbers in the rectangle represent the decoding order. As Figure 33 shown, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As Figure 34 shown, the list of reference pictures is a list showing candidates for reference pictures. For example, one picture (or slice) can have more than one list of reference pictures. For example, if the current picture is a single-predicted picture, one list of reference pictures is used, and if the current picture is a bi-predicted picture, two lists of reference pictures are used. In the examples of Figure 33 and Figure 34 , the picture B3 as the current picture currPic has two lists of reference pictures, namely the L0 list and the L1 list. When the current picture currPic is the picture B3, the candidates for the reference pictures of this current picture currPic are I0, P1, and B2, and each list of reference pictures (that is, the L0 list and the L1 list) represents these pictures. The inter-frame prediction unit 126 or the prediction control unit 128 specifies whether to actually refer to which picture in each list of reference pictures through the reference picture index refidxLx. In Figure 34 , the reference pictures P1 and B2 are specified through the reference picture indices refIdxL0 and refIdxL1.

[0433] Such a list of reference pictures can be generated in units of sequence, picture, slice, tile, CTU, or CU. In addition, the reference picture indices of the reference pictures shown in the list of reference pictures that are referred to in inter-frame prediction can be encoded at the sequence level, picture level, slice level, tile level, CTU level, or CU level. In addition, in multiple inter-frame prediction modes, a common list of reference pictures can also be used.

[0434] [Basic process of inter-frame prediction]

[0435] Figure 35 is a flowchart showing the basic process of inter-frame prediction.

[0436] The inter-frame prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).

[0437] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter-frame prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Additionally, in the candidate MV list, the previously derived MV can be added as a candidate MV. Further, in the derivation of the MV, the inter-frame prediction unit 126 can also determine the MV of the current block by further selecting at least one candidate MV from at least one candidate MV and determining the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 can determine the MV of the current block by searching the region of the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. Additionally, the action of searching the region of the reference picture can also be referred to as motion search.

[0438] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126, but the processing of, for example, step Se_1 or step Se_2 can also be performed by other components included in the encoding device 100.

[0439] Additionally, a candidate MV list can be created for each process in each inter-frame prediction mode, or a common candidate MV list can be used in multiple inter-frame prediction modes. Additionally, the processes of steps Se_3 and Se_4 respectively correspond to Figure 9 the processes of steps Sa_3 and Sa_4 shown. Additionally, the process of step Se_3 corresponds to Figure 30 the process of step Sd_1b.

[0440] [Flow of MV Derivation]

[0441] Figure 36 is a flowchart showing an example of MV derivation.

[0442] The inter-frame prediction unit 126 can derive the MV of the current block in a mode where motion information (e.g., MV) is encoded. In this case, for example, the motion information can be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.

[0443] Alternatively, the inter prediction unit 126 may derive an MV in a mode in which motion information is not encoded. In this case, the motion information is not included in the stream.

[0444] Here, the modes of MV derivation include the ordinary inter mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes in which motion information is encoded include the ordinary inter mode, the ordinary merge mode, and the affine mode (specifically, the affine inter mode and the affine merge mode), etc. In addition, the motion information may include not only the MV but also the predicted MV selection information, which will be described later. In addition, the modes in which motion information is not encoded include the FRUC mode, etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.

[0445] Figure 37 is a flowchart showing another example of MV derivation.

[0446] The inter prediction unit 126 may derive the MV of the current block in a mode in which the differential MV is encoded. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. The differential MV is the difference between the MV of the current block and its predicted MV. In addition, the predicted MV is a predicted motion vector.

[0447] Alternatively, the inter prediction unit 126 may derive an MV in a mode in which the differential MV is not encoded. In this case, the encoded differential MV is not included in the stream.

[0448] Here, as described above, the modes of MV derivation include the ordinary inter mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes in which the differential MV is encoded include the ordinary inter mode and the affine mode (specifically, the affine inter mode), etc. In addition, the modes in which the differential MV is not encoded include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode), etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.

[0449] [Modes of MV Derivation]

[0450] Figure 38A and Figure 38B is an example of a diagram showing the classification of each mode of MV derivation. For example, as Figure 38AAs shown, according to whether motion information is encoded and whether differential MVs are encoded, the MV derivation modes are classified into three major modes. The three modes are the inter-frame mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter-frame mode is the mode that performs motion search and is the mode that encodes motion information and differential MVs. For example, as Figure 38B shown, the inter-frame mode includes the affine inter-frame mode and the normal inter-frame mode. The merge mode is the mode that does not perform motion search and is the mode that selects an MV from the surrounding encoded blocks and uses that MV to derive the MV of the current block. This merge mode is basically the mode that encodes motion information and does not encode differential MVs. For example, as Figure 38B shown, the merge mode includes the normal merge mode (sometimes also referred to as the usual merge mode or the regular merge mode), the MMVD (Merge with Motion Vector Difference) mode, the CIIP (Combined inter merge / intraprediction) mode, the triangular mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the various modes included in the merge mode, differential MVs are encoded exceptionlessly. In addition, the above-mentioned affine merge mode and affine inter-frame mode are the modes included in the affine mode. The affine mode is the mode that assumes an affine transformation and derives the MV of the current block by using the MVs of the multiple sub-blocks that make up the current block. The FRUC mode is the mode that derives the MV of the current block by searching between encoded regions and is the mode that does not encode either motion information or differential MVs. In addition, the details of these various modes will be described later.

[0451] In addition, Figure 38A and Figure 38B shown, the classification of the various modes is an example and is not limited thereto. For example, when differential MVs are encoded in the CIIP mode, this CIIP mode is classified as the inter-frame mode.

[0452] [MV Derivation > Normal Inter-frame Mode]

[0453] The normal inter-frame mode is an inter-frame prediction mode that derives the MV of the current block by finding a block similar to the image of the current block from the region of the reference picture represented by the candidate MVs. In addition, in this normal inter-frame mode, differential MVs are encoded.

[0454] Figure 39 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.

[0455] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located temporally or spatially around the current block (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0456] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction MV candidates respectively in a predetermined order of priority (step Sg_2). In addition, this order of priority is determined in advance for each of the N candidate MVs.

[0457] Next, the inter-frame prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the selected prediction MV into the stream. That is, the inter-frame prediction unit 126 outputs the prediction MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0458] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.

[0459] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are executed for each block. For example, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. In addition, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. Further, it may be that when the processes of steps Sg_1 to Sg_5 are not executed for all the blocks included in a slice but for some blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, it may be that when the processes of steps Sg_1 to Sg_5 are executed for some blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends.

[0460] In addition, the predicted image is the inter-frame prediction signal described above. Further, information indicating the inter-frame prediction mode (in the above example, the normal inter-frame mode) used in the generation of the predicted image included in the coded signal is coded as, for example, a prediction parameter.

[0461] In addition, the candidate MV list may also be commonly used with lists used in other modes. Further, processing related to the candidate MV list may be applied to processing related to lists used in other modes. Processing related to this candidate MV list is, for example, extracting or selecting a candidate MV from the candidate MV list, rearranging the candidate MVs, or deleting a candidate MV, etc.

[0462] [MV Derivation>Normal Merge Mode]

[0463] The normal merge mode is an inter-frame prediction mode in which a candidate MV is selected from the candidate MV list as the MV of the current block to derive the MV. In addition, the normal merge mode is a narrow sense of the merge mode, and is sometimes simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.

[0464] Figure 40 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode.

[0465] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of coded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0466] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 codes the MV selection information for identifying the selected candidate MV into the stream. That is, the inter-frame prediction unit 126 outputs the MV selection information as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130.

[0467] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processes of steps Sh_1 to Sh_3 are performed, for example, on each block. For example, when the processes of steps Sh_1 to Sh_3 are respectively performed on all the blocks included in a slice, the inter-frame prediction using the ordinary merge mode for that slice ends. Also, when the processes of steps Sh_1 to Sh_3 are respectively performed on all the blocks included in a picture, the inter-frame prediction using the ordinary merge mode for that picture ends. In addition, the processes of steps Sh_1 to Sh_3 may be such that when they are not performed on all the blocks included in a slice but on a part of the blocks, the inter-frame prediction using the ordinary merge mode for that slice ends. Similarly, the processes of steps Sh_1 to Sh_3 may be such that when they are performed on a part of the blocks included in a picture, the inter-frame prediction using the ordinary merge mode for that picture ends.

[0468] In addition, information indicating the inter-frame prediction mode (in the above example, the ordinary merge mode) used in the generation of the predicted image included in the stream is encoded as, for example, prediction parameters.

[0469] Figure 41 is a diagram for explaining an example of the MV derivation process of the current picture based on the ordinary merge mode.

[0470] First, the inter-frame prediction unit 126 generates a candidate MV list in which candidate MVs are registered. As candidate MVs, there are: spatially adjacent candidate MVs, which are MVs possessed by a plurality of encoded blocks located in the spatial vicinity of the current block; temporally adjacent candidate MVs, which are MVs possessed by blocks in the vicinity where the position of the current block in the encoded reference picture is projected; combined candidate MVs, which are MVs generated by combining the MV values of the spatially adjacent candidate MVs and the temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero, etc.

[0471] Next, the inter-frame prediction unit 126 determines one candidate MV as the MV of the current block by selecting one candidate MV from the plurality of candidate MVs registered in the candidate MV list.

[0472] Moreover, in the entropy encoding unit 110, a signal indicating which candidate MV is selected, i.e., merge_idx, is described in the stream and encoded.

[0473] In addition, in Figure 41 the candidate MVs registered in the candidate MV list described are an example, and the number may be different from that in the figure, or the structure may not include some types of candidate MVs in the figure, or the structure may be appended with candidate MVs other than the types of candidate MVs in the figure.

[0474] The MV of the current block exported through the normal merge mode can also be used to perform the subsequent DMVR (dynamic motion vector refreshing) to determine the final MV. In addition, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same way as the normal merge mode, but encodes the differential MV. As Figure 38B shown, such an MMVD can also be classified as a merge mode together with the normal merge mode. In addition, the differential MV in the MMVD mode may not be the same as the differential MV used in the inter-frame mode. For example, the derivation of the differential MV in the MMVD mode may also be a process with a smaller processing amount than the derivation of the differential MV in the inter-frame mode.

[0475] In addition, the predicted image generated in the inter-frame prediction can be made to coincide with the predicted image generated in the intra-frame prediction to perform the CIIP (Combined inter merge / intra prediction) mode for generating the predicted image of the current block.

[0476] In addition, the candidate MV list may also be referred to as a candidate list. In addition, merge_idx is MV selection information.

[0477] [MV Derivation>HMVP Mode]

[0478] Figure 42 is a diagram for explaining an example of the MV derivation process of the current picture based on the HMVP mode.

[0479] In the normal merge mode, one candidate MV is selected from the candidate MV list generated by referring to the encoded block (e.g., CU), thereby determining the MV of, for example, the CU of the current block. Here, other candidate MVs can also be registered in the candidate MV list. The mode of registering such other candidate MVs is called the HMVP mode.

[0480] In the HMVP mode, separately from the candidate MV list of the normal merge mode, a FIFO (First-In First-Out) buffer for HMVP is used to manage the candidate MVs.

[0481] In the FIFO buffer, motion information such as the MV of the blocks processed in the past is sequentially stored in the new FIFO buffer. In the management of this FIFO buffer, whenever one block is processed, the MV of the latest block (i.e., the immediately preceding processed CU) is stored in the FIFO buffer, and instead, the MV of the earliest CU (i.e., the CU that was processed first) in the FIFO buffer is deleted from the FIFO buffer. In Figure 42In the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the earliest block.

[0482] Then, for example, the inter-frame prediction unit 126 sequentially checks, starting from HMVP1, whether each MV managed in the FIFO buffer is an MV different from all the candidate MVs already registered in the candidate MV list in the ordinary merge mode. Also, the inter-frame prediction unit 126 may, when determining that it is different from all the candidate MVs, add the MV managed in the FIFO buffer as a candidate MV to the candidate MV list in the ordinary merge mode. At this time, the candidate MV registered from the FIFO buffer may be one or more.

[0483] In this way, by using the HMVP mode, not only can the MVs of the spatially or temporally adjacent blocks of the current block be added to the candidates, but also the MVs of the blocks processed in the past can be added to the candidates. As a result, by expanding the change of the candidate MVs in the ordinary merge mode, the possibility of improving the coding efficiency becomes higher.

[0484] In addition, the above MV may also be motion information. That is, the information stored in the candidate MV list and the FIFO buffer may include not only the value of the MV, but also information such as the information indicating the reference picture, the reference direction, and the number of pictures. In addition, the above block is, for example, a CU.

[0485] In addition, Figure 42 the candidate MV list and the FIFO buffer are an example, and the candidate MV list and the FIFO buffer may also be lists or buffers of different sizes from Figure 42 or a structure in which candidate MVs are registered in an order different from Figure 42 . In addition, the processing described here is common to both the encoding device 100 and the decoding device 200.

[0486] In addition, the HMVP mode can also be applied to modes other than the ordinary merge mode. For example, motion information such as the MVs of the blocks processed in the affine mode in the past may be sequentially stored from a new FIFO buffer and used as candidate MVs. The mode in which the HMVP mode is applied in the affine mode may be referred to as a historical affine mode.

[0487] [MV Derivation>FRUC Mode]

[0488] Motion information may also be derived on the side of the decoding device 200 instead of being signaled from the side of the encoding device 100. For example, motion information may also be derived by performing a motion search on the side of the decoding device 200. In such a case, a motion search is performed on the side of the decoding device 200 without using the pixel values of the current block. Such a mode of performing a motion search on the side of the decoding device 200 includes a FRUC (frame rate up-conversion) mode, a PMMVD (pattern matched motion vector derivation) mode, or the like.

[0489] Figure 43 An example of FRUC processing is shown. First, with reference to the MVs of each encoded block adjacent to the current block in space or time, a list representing these MVs as candidate MVs is generated (i.e., it is a candidate MV list and may also be common to the candidate MV list in the normal merge mode) (step Si_1). Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected as the best candidate MV based on this evaluation value. And, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. In addition, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and evaluation values in the reference picture may be performed on the peripheral region of the best candidate MV, and if there is an MV with a better evaluation value, the best candidate MV is updated to this MV and used as the final MV for the current block. The update to an MV with a better evaluation value may also not be performed.

[0490] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_1 to Si_5 are performed, for example, for each block. For example, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. In addition, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. In addition, the processes of steps Si_1 to Si_5 may be such that when they are not performed for all the blocks included in a slice but are performed for some blocks, the inter-frame prediction using the FRUC mode for that slice ends. Similarly, the processes of steps Si_1 to Si_5 may be such that when they are performed for some blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends.

[0491] The same processing as that in the above block unit may also be performed in the case of processing in sub-block units.

[0492] The evaluation value can also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the MV is compared with the reconstructed image of a specified region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). Then, the difference in pixel values of the two reconstructed images can also be calculated for use as the evaluation value of the MV. In addition, it may be that other information is used in addition to the difference value to calculate the evaluation value.

[0493] Next, the pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (also called the merge list) is selected as the starting point for the search based on pattern matching. As the pattern matching, the first pattern matching or the second pattern matching can be used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.

[0494] [MV Derivation>FRUC>Bilateral Matching]

[0495] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate MV, a region in another reference picture along the motion trajectory of the current block is used.

[0496] Figure 44This is a diagram illustrating an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along a motion trajectory. As Figure 44 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks that are along the motion trajectory of the current block (Cur block) and are in two different reference pictures (Ref0, Ref1), two MVs (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.

[0497] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional MVs are derived.

[0498] [MV Derivation > FRUC > Template Matching]

[0499] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and a block in the reference picture. Thus, in the second pattern matching, as the specified region for calculating the evaluation value of the above candidate MV, the block adjacent to the current block in the current picture is used.

[0500] Figure 45 This is a diagram illustrating an example of the pattern matching (template matching) between a template in the current picture and a block in the reference picture. As Figure 45 shown, in the second pattern matching, the MV of the current block is derived by searching for the block in the reference picture (Ref0) that most matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the left adjacent and / or upper adjacent encoded region(s) and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.

[0501] Information indicating whether such a representation adopts the FRUC mode (e.g., referred to as the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the method of pattern matching that can be adopted (the first pattern matching or the second pattern matching) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0502] [MV Derivation > Affine Mode]

[0503] The affine mode is a mode that generates MVs using an affine transformation. For example, MVs can also be derived in sub-block units based on the MVs of multiple adjacent blocks. This mode is sometimes referred to as the affine motion compensation prediction mode.

[0504] Figure 46A It is a diagram for illustrating an example of the derivation of MVs in sub-block units based on the MVs of multiple adjacent blocks. In Figure 46A , the current block includes, for example, 16 sub-blocks composed of 4×4 pixels. Here, the motion vector v0 of the upper left control point of the current block is derived based on the MVs of adjacent blocks. Similarly, the motion vector v1 of the upper right control point of the current block is derived based on the MVs of adjacent sub-blocks. Then, according to the following formula (1A), the two motion vectors v0 and v1 are projected, and the motion vectors (v x , v y ) of each sub-block within the current block are derived.

[0505]

Equation 1

[0506]

[0507] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a predetermined weight coefficient.

[0508] Information indicating such an affine mode (e.g., referred to as the affine flag) can be signaled at the CU level. Additionally, the signaling of the information indicating this affine mode does not need to be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0509] In addition, in such an affine mode, several modes with different methods of deriving the MVs of the upper left and upper right control points can also be included. For example, in the affine mode, there are two modes: the affine inter-frame (also referred to as the affine ordinary inter-frame) mode and the affine merge mode.

[0510] Figure 46B This is an example diagram for explaining the derivation of the MV of the sub - block unit in the affine mode using three control points. In Figure 46B , the current block includes 16 sub - blocks of 4×4 pixels. Here, based on the MVs of adjacent blocks, the motion vector v0 of the upper - left control point of the current block is derived. Similarly, based on the MVs of adjacent blocks, the motion vector v1 of the upper - right control point of the current block is derived, and based on the MVs of adjacent blocks, the motion vector v2 of the lower - left control point of the current block is derived. Then, according to the following equation (1B), the three motion vectors v0, v1, and v2 are projected to derive the motion vectors (v x , v y ) of each sub - block within the current block.

[0511]

Equation 2

[0512]

[0513] Here, x and y represent the horizontal position and vertical position of the sub - block center respectively, and w and h represent pre - determined weight coefficients. It can also be that w represents the width of the current block and h represents the height of the current block.

[0514] The affine modes using different numbers of control points (e.g., two and three) can also be switched at the CU level and signaled. In addition, the information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub - block level).

[0515] In addition, in such an affine mode with three control points, several modes with different derivation methods for the MVs of the upper - left, upper - right, and lower - left control points can also be included. For example, in the affine mode with three control points, similar to the affine mode with two control points, there are two modes: the affine inter - frame mode and the affine merge mode.

[0516] In addition, in the affine mode, the size of each sub - block included in the current block is not limited to 4x4 pixels and can also be other sizes. For example, the size of each sub - block can also be 8×8 pixels.

[0517] [MV Derivation > Affine Mode > Control Points]

[0518] Figure 47A , Figure 47B and Figure 47C are conceptual diagrams for explaining an example of the MV derivation of control points in the affine mode.

[0519] In the affine mode, as Figure 47AAs shown, for example, based on a plurality of MVs corresponding to blocks encoded in an affine mode in encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, a predicted MV for each control point of the current block is calculated. Specifically, these blocks are examined in the order of encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and the first valid block encoded in the affine mode is determined. The MV of the control point of the current block is calculated based on the plurality of MVs corresponding to the determined block.

[0520] For example, as Figure 47B shown, in the case where block A adjacent to the left side of the current block is encoded in an affine mode with two control points, motion vectors v3 and v4 projected onto positions at the upper left corner and upper right corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are calculated.

[0521] For example, as Figure 47C shown, when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 projected onto positions at the upper left corner, upper right corner, and lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block are calculated.

[0522] In addition, Figures 47A to 47C the method for deriving the MVs shown can be used for deriving the MVs of each control point of the current block in step Sk_1 described later, and can also be used for deriving the predicted MVs of each control point of the current block in step Sj_1 described later. Figure 50 shown, Figure 51

[0523] Figure 48A And Figure 48B is a conceptual diagram for explaining another example of deriving the control point MVs in the affine mode.

[0524] Figure 48A is a diagram for explaining the affine mode with two control points.

[0525] In this affine mode, as Figure 48AAs shown, the MV selected from the MVs of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block.

[0526] Figure 48B FIG. is for explaining an affine mode with three control points.

[0527] In this affine mode, as Figure 48B shown, the MV selected from the MVs of the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the upper left control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 of the upper right control point of the current block. In addition, the MV selected from the MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 of the lower left control point of the current block.

[0528] In addition, Figure 48A and Figure 48B the MV derivation method shown can be used for the derivation of the MVs of each control point of the current block in step Sk_1 described later, and can also be used for the derivation of the predicted MVs of each control point of the current block in step Sj_1 of Figure 50 shown later. Figure 51

[0529] Here, for example, in the case of signaling with different numbers of control points (e.g., two and three) in the affine mode at the CU level, etc., the number of control points may be different depending on the encoded block and the current block.

[0530] Figure 49A and Figure 49B FIG. is a conceptual diagram for explaining an example of the MV derivation method of the control points in the case where the number of control points is different between the encoded block and the current block.

[0531] For example, as Figure 49A shown, the current block has three control points: upper left, upper right, and lower left, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points. In this case, the motion vectors v3 and v4 projected to the positions of the upper left and upper right corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block are calculated. Furthermore, based on the derived motion vectors v0 and v1, the motion vector v2 of the lower left control point is calculated.

[0532] For example, asFigure As shown, the current block has two control points, namely the upper left corner and the upper right corner. The block A adjacent to the left side of the current block is encoded in an affine mode with three control points. In this case, motion vectors v3, v4, and v5 projected to the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are calculated.

[0533] In addition, ​ and ​ the method for deriving the MV shown can be used for deriving the MV of each control point of the current block in step Sk_1 described later, and can also be used for deriving the predicted MV of each control point of the current block in step Sj_1 described later. ​ shown, ​

[0534] [MV Derivation > Affine Mode > Affine Merge Mode]

[0535] ​ is a flowchart showing an example of the affine merge mode.

[0536] In the affine merge mode, first, the inter-frame prediction unit 126 derives the MV of each control point of the current block (step Sk_1). The control points are, as ​ shown, the upper left corner and the upper right corner points of the current block, or, as ​ shown, the upper left corner, the upper right corner, and the lower left corner points of the current block. At this time, the inter-frame prediction unit 126 can also encode the MV selection information used to identify the two or three derived MVs into the stream.

[0537] For example, in the case of using the method for deriving the MV shown in ​ shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), as ​ shown, and determines the initial valid block encoded in the affine mode.

[0538] The inter-frame prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control point. For example, when block A is determined and block A has two control points, as ​As shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block including block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block.

[0539] Alternatively, in the case where block A is determined and block A has three control points, as ​ shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block including block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block.

[0540] In addition, as shown above ​ shown, block A can be determined, and in the case where block A has two control points, the MVs of three control points can be calculated. Also, as shown above ​ shown, block A can be determined, and in the case where block A has three control points, the MVs of two control points can be calculated.

[0541] Next, the inter-frame prediction unit 126 performs motion compensation for each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 126 calculates the MV of each of the multiple sub-blocks as an affine MV using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation for the sub-block (step Sk_3). When the processes of steps Sk_2 and Sk_3 are respectively performed for all the sub-blocks included in the current block, the process of generating the predicted image using the affine merge mode for the current block ends. That is, motion compensation is performed for the current block, and the predicted image of the current block is generated.

[0542] In addition, in step Sk_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can, for example, also be a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be ​ the MV derivation method shown ​ and​ The method for deriving the MV shown ​ and ​ the method for deriving the MV shown, and any combination of other methods for deriving the MV.

[0543] In addition, the candidate MV list may also include candidate MVs of a mode that performs prediction in units of sub-blocks other than the affine mode.

[0544] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and candidate MVs of an affine merge mode with 3 control points may be generated. Or, a candidate MV list including candidate MVs of an affine merge mode with 2 control points may be generated separately, and a candidate MV list including candidate MVs of an affine merge mode with 3 control points may be generated. Or, a candidate MV list including candidate MVs of one of the affine merge modes with 2 control points and the affine merge mode with 3 control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or the MV of the valid block among these blocks.

[0545] In addition, as the MV selection information, an index indicating which candidate MV in the candidate MV list may be sent.

[0546] [MV Derivation>Affine Mode>Affine Inter-Frame Mode]

[0547] ​ is a flowchart showing an example of the affine inter-frame mode.

[0548] In the affine inter-frame mode, first, the inter-frame prediction unit 126 derives the predicted MVs (v0, v1) or (v0, v1, v2) of each of the 2 or 3 control points of the current block (step Sj_1). As ​ or ​ shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0549] For example, in the case of using the method for deriving the MV shown in ​ and ​ the inter-frame prediction unit 126 derives the predicted MVs (v0, v1) or (v0, v1, v2) of the control points of the current block by selecting the MV of a certain block among the encoded blocks near each control point of the current block shown in ​ or ​ At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the 2 or 3 selected predicted MVs into the stream.

[0550] For example, the inter-frame prediction unit 126 can determine which block's MV among the encoded blocks adjacent to the current block is to be used as the predicted MV of the control point by using cost evaluation or the like, and can describe a flag indicating which predicted MV is selected in the bitstream. That is, the inter-frame prediction unit 126 outputs prediction MV selection information such as a flag as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0551] Next, while respectively updating the predicted MVs selected or derived in step Sj_1 (step Sj_2), the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4). That is, the inter-frame prediction unit 126 uses the MV of each sub-block corresponding to the predicted MV to be updated as the affine MV, and calculates it using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference picture (step Sj_4). Whenever the predicted MV is updated in step Sj_2, the processes of steps Sj_3 and Sj_4 are executed for all blocks within the current block. As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted MV that can obtain the minimum cost as the MV of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0552] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).

[0553] In addition, in step Sj_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can, for example, also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be Figures 47A - 47C the MV derivation method shown, Figure 48A and Figure 48B the MV derivation method shown, Figure 49A and Figure 49B the MV derivation method shown, and any combination of other MV derivation methods.

[0554] In addition, the candidate MV list can also contain candidate MVs of prediction modes that perform prediction in units of sub-blocks other than the affine mode.

[0555] In addition, as a candidate MV list, a candidate MV list including candidate MVs of an affine inter-frame mode with two control points and candidate MVs of an affine inter-frame mode with three control points may also be generated. Alternatively, a candidate MV list including candidate MVs of an affine inter-frame mode with two control points and a candidate MV list including candidate MVs of an affine inter-frame mode with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of a mode of either an affine inter-frame mode with two control points or an affine inter-frame mode with three control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be the MV of the valid block among these blocks.

[0556] In addition, as prediction MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent out.

[0557] [MV Derivation > Triangle Mode]

[0558] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the rectangular current block. However, the inter-frame prediction unit 126 may generate a plurality of prediction images having a shape different from that of the rectangle for the rectangular current block, and generate a final rectangular prediction image by combining these plurality of prediction images. The shape different from that of the rectangle may also be a triangle, for example.

[0559] Figure 52A is a diagram for explaining the generation of two triangular prediction images.

[0560] The inter-frame prediction unit 126 performs motion compensation for the first partition of the triangle within the current block using the first MV of the first partition, thereby generating a triangular prediction image. Similarly, the inter-frame prediction unit 126 performs motion compensation for the second partition of the triangle within the current block using the second MV of the second partition, thereby generating a triangular prediction image. Then, the inter-frame prediction unit 126 combines these prediction images, thereby generating a rectangular prediction image identical to the current block.

[0561] In addition, as the prediction image of the first partition, a first rectangular prediction image corresponding to the current block may also be generated using the first MV. In addition, as the prediction image of the second partition, a second rectangular prediction image corresponding to the current block may also be generated using the second MV. The prediction image of the current block may also be generated by weighted addition of the first prediction image and the second prediction image. In addition, the region where the weighted addition is performed may also be only a partial region sandwiching the boundary between the first partition and the second partition.

[0562] Figure 52BIt is a conceptual diagram showing an example of the first part of the first partition that overlaps with the second partition, and the first sample set and the second sample set that can be weighted as part of the correction process. The first part can be, for example, one-fourth of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero. For example, N can be the integer 2. Figure 52B A rectangular partition representing a rectangular part having a width that is one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example in the center of Figure 52B represents a rectangular partition with a rectangular part having a height that is one-fourth of the height of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example on the right of Figure 52B represents a triangular partition with a polygonal part having a height corresponding to 2 samples. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.

[0563] The first part can be a part of the first partition that overlaps with an adjacent partition. Figure 52C It is a conceptual diagram showing the first part of the first partition that is a part of the first partition overlapping with a part of an adjacent partition. For simplicity of explanation, a rectangular partition with a part overlapping with a spatially adjacent rectangular partition is shown. Partitions with other shapes such as triangular partitions can be used, and the overlapping part can also overlap with spatially or temporally adjacent partitions.

[0564] In addition, an example of generating prediction images for two partitions respectively using inter-frame prediction is shown, but intra-frame prediction can also be used to generate a prediction image for at least one partition.

[0565] Figure 53 It is a flowchart showing an example of a triangular pattern.

[0566] In the triangular pattern, first, the inter-frame prediction unit 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame prediction unit 126 can encode the partition information, which is information related to the division into each partition, as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 can output the partition information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0567] Next, the inter-frame prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially located around the current block (step Sx_2). That is, the inter-frame prediction unit 126 creates a candidate MV list.

[0568] Then, the inter-frame prediction unit 126 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_2 (step Sx_3). At this time, the inter-frame prediction unit 126 may also encode the MV selection information for identifying the selected candidate MV as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 may output the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0569] Next, the inter-frame prediction unit 126 uses the selected first MV and the encoded reference picture to perform motion compensation, thereby generating a first prediction image (step Sx_4). Similarly, the inter-frame prediction unit 126 uses the selected second MV and the encoded reference picture to perform motion compensation, thereby generating a second prediction image (step Sx_5).

[0570] Finally, the inter-frame prediction unit 126 performs weighted addition on the first prediction image and the second prediction image, thereby generating a prediction image of the current block (step Sx_6).

[0571] In addition, in Figure 52A the example shown, the first partition and the second partition are respectively triangles, but they may also be trapezoids, or may be of mutually different shapes. Moreover, in Figure 52A the example shown, the current block is composed of 2 partitions, but it may also be composed of 3 or more partitions.

[0572] In addition, the first partition and the second partition may be repeated. That is, the first partition and the second partition may include the same pixel area. In this case, the prediction image in the first partition and the prediction image in the second partition may also be used to generate the prediction image of the current block.

[0573] In addition, in this example, an example in which prediction images are generated by inter-frame prediction in both 2 partitions is shown, but prediction images may also be generated by intra-frame prediction for at least 1 partition.

[0574] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different, or may be the same candidate MV list.

[0575] In addition, the partition information may include an index indicating a splitting direction for splitting at least the current block into a plurality of partitions. The MV selection information may also include an index indicating the selected first MV and an index indicating the selected second MV. One index may also represent multiple pieces of information. For example, one index summarizing a part or the whole of the partition information and a part or the whole of the MV selection information may be encoded.

[0576] [MV Derivation>ATMVP Mode]

[0577] Figure 54 FIG. is an example of an ATMVP mode for deriving MVs in sub-block units.

[0578] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, candidate MVs in sub-block units are registered in a candidate MV list for a normal merge mode.

[0579] Specifically, in the ATMVP mode, first, as Figure 54 shown, in the coded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, a temporal MV reference block corresponding to the current block is determined. Then, for each sub-block within the current block, the MV used for encoding the region corresponding to the sub-block within the temporal MV reference block is determined. The MVs thus determined are included in the candidate MV list as candidate MVs for the sub-blocks of the current block. When selecting the candidate MVs for each sub-block from the candidate MV list, motion compensation is performed on the sub-block using the candidate MV as the MV of the sub-block. Thereby, a predicted image for each sub-block is generated.

[0580] In addition, in Figure 54 the example shown, as the surrounding MV reference block, the block adjacent to the lower left of the current block is used, but other blocks may also be used. In addition, the size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may also be switched in units of slices, bricks, or pictures, etc.

[0581] [Motion Search>DMVR]

[0582] Figure 55 FIG. is a diagram showing the relationship between the merge mode and DMVR.

[0583] The inter-frame prediction unit 126 derives the MV of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform an MV search, that is, a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in the merge mode.

[0584] On the other hand, when it is determined in step Sl_1 to perform a motion search (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by DMVR.

[0585] Figure 56 It is a conceptual diagram for explaining an example of DMVR for determining MV.

[0586] First, for example, in the merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are determined based on the coded picture in the L0 list, that is, the first reference picture (L0). Similarly, according to the candidate MV (L1), reference pixels are determined based on the coded picture in the L1 list, that is, the second reference picture (L1). A template is generated by taking the average of these reference pixels.

[0587] Next, using this template, the peripheral areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV of the current block. In addition, the cost can be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values, etc.

[0588] Even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV to derive the final MV, any processing can be used.

[0589] Figure 57 It is a conceptual diagram for explaining another example of DMVR for determining MV. Figure 57 The example shown is different from Figure 56 an example of the DMVR shown, and does not generate a template but calculates the cost.

[0590] First, the inter-frame prediction unit 126 searches the periphery of the reference blocks included in the reference pictures of the L0 list and the L1 list based on the candidate MV, that is, the initial MV, obtained from the candidate MV list. For example, as Figure 57As shown, the initial MV corresponding to the reference block of the L0 list is InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is InitMV_L1. In motion search, the inter-frame prediction unit 126 first sets the search position for the reference picture in the L0 list. The differential vector representing this set search position, specifically, the differential vector from the position represented by the initial MV (i.e., InitMV_L0) to this search position is MVd_L0. Then, the inter-frame prediction unit 126 determines the search position in the reference picture of the L1 list. This search position is represented by the differential vector from the position indicated by the initial MV (i.e., InitMV_L1) to this search position. Specifically, the inter-frame prediction unit 126 determines this differential vector as MVd_L1 by mirroring MVd_L0. That is, the inter-frame prediction unit 126 sets the positions that are symmetric from the position represented by the initial MV as the search positions in the reference pictures of the L0 list and the L1 list respectively. The inter-frame prediction unit 126 calculates, for each search position, the sum of the absolute differences (SAD) of the pixel values within the block at this search position as the cost, and finds the search position with the minimum cost.

[0591] Figure 58A FIG. is an example diagram showing motion search in DMVR, Figure 58B is a flowchart showing an example of this motion search.

[0592] First, in Step1, the inter-frame prediction unit 126 calculates the cost of the search position (also referred to as the starting point) represented by the initial MV and the 8 search positions around it. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than the starting point is the minimum. Here, when it is determined that the cost of the search positions other than the starting point is the minimum, the inter-frame prediction unit 126 moves to the search position with the minimum cost and performs the processing of Step2. On the other hand, if the cost of the starting point is the minimum, the inter-frame prediction unit 126 skips the processing of Step2 and performs the processing of Step3.

[0593] In Step2, the inter-frame prediction unit 126 uses the search position moved according to the processing result of Step1 as the new starting point and performs the same search as the processing of Step1. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than this starting point is the minimum. Here, if the cost of the search positions other than the starting point is the minimum, the inter-frame prediction unit 126 performs the processing of Step4. On the other hand, if the cost of the starting point is the minimum, the inter-frame prediction unit 126 performs the processing of Step3.

[0594] In Step4, the inter-frame prediction unit 126 processes the search position of this starting point as the final search position, and determines the difference between the position represented by the initial MV and this final search position as the differential vector.

[0595] In Step 3, the inter-frame prediction unit 126 determines a fractional-precision pixel position with the minimum cost based on the costs at four points above, below, left, and right of the start point in Step 1 or Step 2, and sets this pixel position as the final search position. This fractional-precision pixel position is determined by weighted addition of the vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) at the four points above, below, left, and right, with the costs at the respective search positions of the four points as weights. Then, the inter-frame prediction unit 126 determines the difference between the position represented by the initial MV and this final search position as the differential vector.

[0596] [Motion Compensation > BIO / OBMC / LIC]

[0597] In motion compensation, there is a mode of generating a predicted image and correcting this predicted image. This mode is, for example, BIO, OBMC, and LIC described later.

[0598] Figure 59 It is a flowchart showing an example of the generation of a predicted image.

[0599] The inter-frame prediction unit 126 generates a predicted image (step Sm_1), and corrects this predicted image by any of the above modes (step Sm_2).

[0600] Figure 60 It is a flowchart showing another example of the generation of a predicted image.

[0601] The inter-frame prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image using this MV (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting this predicted image (step Sn_4). Also, in LIC described later, it is also possible to correct the luminance and chrominance differences in step Sn_4. On the other hand, when it is determined not to perform the correction process (no in step Sn_3), the inter-frame prediction unit 126 outputs this predicted image as the final predicted image without correcting this predicted image (step Sn_5).

[0602] [Motion Compensation > OBMC]

[0603] Not only the motion information of the current block obtained through motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame predicted image. Specifically, an inter-frame predicted image can also be generated in units of sub-blocks within the current block by weighted addition of a predicted image based on the motion information obtained through motion search (within a reference picture) and a predicted image based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or the OBMC mode.

[0604] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Additionally, the level at which these pieces of information are signaled need not be limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0605] A more specific description of the OBMC mode is given. Figure 61 And Figure 62 are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC.

[0606] First, as Figure 62 shown, using the MV assigned to the current block, a predicted image (Pred) based on normal motion compensation is obtained. In Figure 62 , the arrow "MV" points to the reference picture and indicates which block in the reference picture the current block in the current picture refers to for obtaining the predicted image.

[0607] Next, the MV (MV_L) that has been derived for the already encoded left adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two predicted images Pred and Pred_L, the first correction of the predicted image is performed. This has the effect of blending the boundaries between adjacent blocks.

[0608] Similarly, the Motion Vector (MV_U) that has been derived for the already - coded upper - adjacent block is applied (re - used) to the current block to obtain a predicted image (Pred_U). The MV (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second - stage correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted images that have been corrected for the first time (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second - stage correction is the final predicted image of the current block where the boundaries with adjacent blocks are blended (smoothed).

[0609] In addition, in the above example, a two - path correction method using the left - adjacent and upper - adjacent blocks is used, but this correction method can also be a three - path or more - path correction method that also uses the right - adjacent and / or lower - adjacent blocks.

[0610] Furthermore, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.

[0611] In addition, the prediction - image correction process of OBMC has been described here. The prediction - image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can be applied to each of the multiple reference pictures. In this case, through the OBMC - based image correction for multiple reference pictures, after obtaining the corrected predicted images from each reference picture, the final predicted image is obtained by further overlapping the multiple obtained corrected predicted images.

[0612] In addition, in OBMC, the unit of the current block can be the PU unit or a sub - block unit obtained by further dividing the PU.

[0613] As a method for determining whether to apply OBMC, for example, there is a method of using a signal indicating whether to apply OBMC, namely obmc_flag. As a specific example, the encoding device 100 can also determine whether the current block belongs to a region with complex motion. When the current block belongs to a region with complex motion, the encoding device 100 sets the obmc_flag value to 1 and applies OBMC for encoding. When the current block does not belong to a region with complex motion, the encoding device 100 sets the obmc_flag value to 0 and does not apply OBMC for block encoding. On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, the decoding is switched according to this value to determine whether to apply OBMC.

[0614] [Motion Compensation > BIO]

[0615] Next, a method for deriving the MV will be described. First, a mode for deriving the MV based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi - directional optical flow) mode. Additionally, the bi - directional optical flow can also be expressed as BDOF instead of BIO.

[0616] Figure 63 is a diagram for explaining a model assuming uniform linear motion. In Figure 63 , (v x , v y ) represents the velocity vector, τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.

[0617] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) holds.

[0618]

Equation 3

[0619]

[0620] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a candidate MV list or the like are corrected in pixel units.

[0621] In addition, the MV can also be derived on the decoding device 200 side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector can also be derived in sub - block units based on the MVs of multiple adjacent blocks.

[0622] Figure 64 is a flowchart showing an example of inter - frame prediction according to BIO. Additionally, Figure 65This is a diagram showing an example of the functional structure of the inter-frame prediction unit 126 that performs inter-frame prediction according to BIO.

[0623] As Figure 65 shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a predicted image correction unit 126f. In addition, the memory 126a may be a frame memory 122.

[0624] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 126 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_1). In addition, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0625] Next, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . In addition, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_2). Here, the interpolation image I 0 is the image included in the reference picture Ref0 derived for the current block, and the interpolation image I 1 is the image included in the reference picture Ref1 derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Alternatively, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each be an image larger than the current block. In addition, the interpolation image I 0 and I 1 may include a predicted image derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1), and a motion compensation filter.

[0626] In addition, the gradient image derivation unit 126c derives the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 , Iy 0 , Iy 1)(Step Sy_3). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivation unit 126c can also derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image only needs to represent the spatial change amount of the pixel values along the horizontal direction or the vertical direction.

[0627] Next, the optical flow derivation unit 126d derives the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block, using the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (Step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and can also be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can be a 4x4 pixel sub-CU. In addition, the derivation of the optical flow can be performed not in units of sub-blocks but in other units such as pixel units.

[0628] Next, the inter-frame prediction unit 126 corrects the prediction image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives the correction value of the values of the pixels included in the current block using the optical flow (vx, vy) (Step Sy_5). Moreover, the prediction image correction unit 126f can also correct the prediction image of the current block using the correction value (Step Sy_6). In addition, the correction value can be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.

[0629] In addition, the processing flow of BIO is not limited to Figure 64 the processing disclosed. It can perform only a part of the processing disclosed in Figure 64 , or can add or replace different processing, or can execute in a different processing order.

[0630] [Motion Compensation > LIC]

[0631] Next, an example of a mode for generating a prediction image (prediction) using LIC (local illumination compensation) will be described.

[0632] Figure 66A is a diagram for explaining an example of a prediction image generation method using a luminance correction process based on LIC. In addition, Figure 66BThis is a flowchart showing an example of a prediction image generation method using this LIC.

[0633] First, the inter-frame prediction unit 126 derives an MV from the encoded reference picture and obtains the reference image corresponding to the current block (step Sz_1).

[0634] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values change between the reference picture and the current picture for the current block (step Sz_2). This extraction is based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position within the reference picture specified by the derived MV. Then, the inter-frame prediction unit 126 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_3).

[0635] The inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image within the reference picture specified by the MV by applying this luminance correction parameter (step Sz_4). That is, the prediction image, which is the reference image within the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, it is possible to correct the luminance or the color difference. That is, it is also possible to calculate a correction parameter for the color difference using the information indicating how the color difference changes and perform a color difference correction process.

[0636] In addition, Figure 66A the shape of the peripheral reference region in

[0637] is an example, and shapes other than this can also be used.

[0638] Furthermore, although the process of generating a prediction image based on one reference picture has been described here, the same applies when generating a prediction image based on multiple reference pictures. It is also possible to generate a prediction image after performing a luminance correction process on the reference images obtained from each reference picture in the same manner as described above.

[0638] As a method for determining whether to apply the LIC, for example, there is a method of using a lic_flag as a signal indicating whether to apply the LIC. As a specific example, in the encoding device 100, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and encoding is performed by applying the LIC. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC. On the other hand, in the decoding device 200, it is also possible to decode the lic_flag described in the stream and switch whether to apply the LIC according to its value for decoding.

[0639] As another method for determining whether to apply LIC, for example, there is also a method of determining based on whether LIC is applied to neighboring blocks. As a specific example, when the current block is processed in the merge mode, the inter-frame prediction unit 126 determines whether the neighboring encoded blocks selected at the time of deriving the MV in the merge mode are encoded with LIC applied. Based on the result, the inter-frame prediction unit 126 switches whether to apply LIC for encoding. Additionally, in the case of this example, the same processing also applies to the decoding device 200 side.

[0640] Use Figure 66A And Figure 66B LIC (luminance correction processing) has been described above. Hereinafter, its detailed content will be described.

[0641] First, the inter-frame prediction unit 126 derives an MV for obtaining a reference image corresponding to the current block from a reference picture that is an encoded picture.

[0642] Next, for the current block, the inter-frame prediction unit 126 uses the luminance pixel values of the left adjacent and upper adjacent encoded peripheral reference regions and the luminance pixel values at the same positions in the reference picture specified by the MV to extract information indicating how the luminance values change between the reference picture and the current picture, and calculates a luminance correction parameter. For example, let the luminance pixel value of a certain pixel in the peripheral reference region within the current picture be p0, and let the luminance pixel value of the pixel at the same position in the peripheral reference region within the reference picture be p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1 + B = p0 as the luminance correction parameter for multiple pixels in the peripheral reference region.

[0643] Next, the inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image in the reference picture specified by the MV using the luminance correction parameter. For example, let the luminance pixel value in the reference image be p2, and let the luminance pixel value of the prediction image after the luminance correction process be p3. The inter-frame prediction unit 126 generates a prediction image after the luminance correction process by calculating A×p2 + B = p3 for each pixel in the reference image.

[0644] In addition, it is also possible to use Figure 66A a part of the peripheral reference region shown. For example, it is also possible to use a region including a specified number of pixels respectively removed at intervals from the upper adjacent pixel and the left adjacent pixel as the peripheral reference region. Additionally, the peripheral reference region is not limited to the region adjacent to the current block, and it can also be a region not adjacent to the current block. Additionally, in Figure 66AIn the example shown, the peripheral reference region in the reference picture is the region specified by the MV of the current picture from the peripheral reference region in the current picture, but it can also be the region specified by other MVs. For example, the other MV can also be the MV of the peripheral reference region in the current picture.

[0645] In addition, the operation of the encoding device 100 has been described here, but the operation of the decoding device 200 is the same.

[0646] Furthermore, LIC can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.

[0647] In addition, LIC can also be applied in units of sub-blocks. For example, correction parameters can be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0648] [Prediction control unit]

[0649] The prediction control unit 128 selects one of the intra-predicted image (pixels or signals output from the intra-prediction unit 124) and the inter-predicted image (pixels or signals output from the inter-prediction unit 126), and outputs the selected predicted image to the subtraction unit 104 and the addition unit 116.

[0650] [Prediction parameter generation unit]

[0651] The prediction parameter generation unit 130 can output information related to the selection of the predicted image in intra-prediction, inter-prediction, and the prediction control unit 128 as prediction parameters to the entropy encoding unit 110. The entropy encoding unit 110 can generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream, and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selected prediction signal (e.g., MV, prediction type, or prediction mode used by the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0652] [Decoding device]

[0653] Next, a decoding device 200 that can decode the stream output from the above encoding device 100 will be described. Figure 67FIG. 0 is a block diagram showing an example of the functional configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes an encoded image, i.e., a stream, in units of blocks.

[0654] As Figure 67 shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. In addition, the intra prediction unit 216 and the inter prediction unit 218 are each configured as part of a prediction processing unit.

[0655] [Installation Example of Decoding Device]

[0656] Figure 68 FIG. 12 is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 67 as shown, a plurality of components of the decoding device 200 are implemented by the Figure 68 shown processor b1 and memory b2.

[0657] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor b1 can also function as Figure 67 a plurality of components of the decoding device 200 shown, etc., excluding the components for storing information.

[0658] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode a stream. The memory b2 can be either an electronic circuit or connected to the processor b1. In addition, the memory b2 can also be included in the processor b1. In addition, the memory b2 can also be an aggregate of multiple electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disk, etc., or can be represented as a storage or a recording medium, etc. In addition, the memory b2 can be either a non-volatile memory or a volatile memory.

[0659] For example, the memory b2 can store an image or a stream. In addition, a program for the processor b1 to decode a stream can also be stored in the memory b2.

[0660] In addition, for example, the memory b2 can also function as Figure 67The functions of the components for storing information among the multiple components of the decoding device 200 shown, etc. Specifically, the memory b2 can function as Figure 67 the block memory 210 and the frame memory 214 shown. More specifically, reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.) can be stored in the memory b2.

[0661] In addition, in the decoding device 200, not all of the multiple components shown, etc. may be installed, and not all of the above-mentioned multiple processes may be performed. Figure 67 A part of the multiple components shown, etc. may be included in other devices, or a part of the above-mentioned multiple processes may be performed by other devices. Figure 67

[0662] Hereinafter, after explaining the overall processing flow of the decoding device 200, each component included in the decoding device 200 will be described. In addition, for the components that perform the same processing as the components included in the encoding device 100 among the components included in the decoding device 200, detailed descriptions will be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, addition unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 included in the decoding device 200 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, addition unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120 included in the encoding device 100, respectively.

[0663] [Overall Flow of Decoding Processing]

[0664] Figure 69 is a flowchart showing an example of the overall decoding processing performed by the decoding device 200.

[0665] First, the segmentation determination unit 224 of the decoding device 200 determines the segmentation style of each of the multiple fixed-size blocks (128×128 pixels) included in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This segmentation style is the segmentation style selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 on each of the multiple blocks constituting this segmentation style.

[0666] The entropy decoding unit 202 decodes the encoded quantization coefficients and prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).

[0667] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore the prediction residual of the current block by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (step Sp_3).

[0668] Next, the prediction processing unit including the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 generates a prediction image of the current block (step Sp_4).

[0669] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction residual to the prediction image (step Sp_5).

[0670] Moreover, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).

[0671] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding has not been completed (No in step Sp_7), the processing from step Sp_1 is repeated.

[0672] In addition, the processing of steps Sp_1 to Sp_7 can be sequentially performed by the decoding device 200, and multiple processes of a part of these processes can be performed in parallel or the order can be changed.

[0673] [Segmentation decision unit]

[0674] Figure 70 FIG. is a diagram showing the relationship between the segmentation decision unit 224 and other components. As an example, the segmentation decision unit 224 can also perform the following processing.

[0675] The segmentation decision unit 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoding unit 202. Moreover, the segmentation decision unit 224 can determine the segmentation pattern of a fixed-size block based on the block information and the parameters. And the segmentation decision unit 224 can also output the information indicating the determined segmentation pattern to the inverse transformation unit 206, the intra prediction unit 216, and the inter prediction unit 218. The inverse transformation unit 206 can perform inverse transformation on the transform coefficients based on the segmentation pattern indicated by the information from the segmentation decision unit 224. The intra prediction unit 216 and the inter prediction unit 218 can generate a prediction image based on the segmentation pattern indicated by the information from the segmentation decision unit 224.

[0676] [Entropy decoding unit]

[0677] Figure 71 FIG. is a block diagram showing an example of the functional structure of the entropy decoding unit 202.

[0678] The entropy decoding unit 202 generates quantization coefficients, prediction parameters, and parameters related to the segmentation pattern, etc. by performing entropy decoding on the stream. For example, CABAC is used in this entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a de-binarization unit 202c. The binary arithmetic decoding unit 202a performs arithmetic decoding on the stream as a binary signal using the context value derived by the context control unit 202b. Similar to the context control unit 110b of the encoding device 100, the context control unit 202b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. The de-binarization unit 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoding unit 202a into a multi-value signal representing the above quantization coefficients, etc. This de-binarization is performed in the same manner as the above binarization.

[0679] The entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the prediction parameters included in the stream (refer to Figure 1 ) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as that performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device 100 side.

[0680] [Entropy Decoding Unit]

[0681] Figure 72 is a diagram showing the process of CABAC in the entropy decoding unit 202.

[0682] First, in the CABAC in the entropy decoding unit 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding unit 202c and setting of the initial context value are performed. Then, the binary arithmetic decoding unit 202c and the de-binarization unit 202c perform arithmetic decoding and de-binarization on the encoded data of the CTU, for example. At this time, the context control unit 202b updates the context value each time arithmetic decoding is performed. Then, the context control unit 202b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.

[0683] [Inverse Quantization Unit]

[0684] The inverse quantization unit 204 inverse-quantizes the quantization coefficients of the current block that is input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 inverse-quantizes each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.

[0685] Figure 73 is a block diagram showing an example of the functional structure of the inverse quantization unit 204.

[0686] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204a, a predicted quantization parameter generation unit 204b, a quantization parameter storage unit 204d, and an inverse quantization processing unit 204e.

[0687] Figure 74 is a flowchart showing an example of the inverse quantization performed by the inverse quantization unit 204.

[0688] As an example, the inverse quantization unit 204 can perform inverse quantization processing for each CU based on Figure 74 the process shown. Specifically, the quantization parameter generation unit 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation unit 204a obtains the differential quantization parameter of the current block from the entropy decoding unit 202 (step Sv_12).

[0689] Next, the predicted quantization parameter generation unit 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). The predicted quantization parameter generation unit 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).

[0690] Then, the quantization parameter generation unit 204a adds the differential quantization parameter of the current block obtained from the entropy decoding unit 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generation unit 204b (step Sv_15). Through this addition, the quantization parameter of the current block is generated. In addition, the quantization parameter generation unit 204a stores the quantization parameter of the current block in the quantization parameter storage unit 204d (step Sv_16).

[0691] Next, the inverse quantization processing unit 204e inverse-quantizes the quantization coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).

[0692] In addition, the differential quantization parameter can also be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can also be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.

[0693] In addition, the inverse quantization unit 204 may include multiple inverse quantizers, and may also perform inverse quantization on the quantization coefficients using an inverse quantization method selected from multiple inverse quantization methods.

[0694] [Inverse transform unit]

[0695] The inverse transform unit 206 restores the prediction residual by performing an inverse transform on the transform coefficients that are the input from the inverse quantization unit 204.

[0696] For example, when the information read from the stream indicates the application of EMT or AMT (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type read.

[0697] In addition, for example, when the information read from the stream indicates the application of NSST, the inverse transform unit 206 applies an inverse re-transform to the transform coefficients.

[0698] Figure 75 It is a flowchart showing an example of the process performed by the inverse transform unit 206.

[0699] For example, the inverse transform unit 206 determines whether there is information indicating that orthogonal transformation is not performed in the stream (step St_11). Here, when it is determined that there is no such information (No in step St_11), the inverse transform unit 206 obtains the information indicating the transform type that has been decoded by the entropy decoding unit 202 (step St_12). Next, the inverse transform unit 206 determines the transform type used in the orthogonal transform of the encoding device 100 based on this information (step St_13). And the inverse transform unit 206 performs inverse orthogonal transformation using the determined transform type (step St_14).

[0700] Figure 76 It is a flowchart showing another example of the process performed by the inverse transform unit 206.

[0701] For example, the inverse transform unit 206 determines whether the transform size is equal to or less than a specified value (step Su_11). Here, when it is determined that the size is equal to or less than the specified value (Yes in step Su_11), the inverse transform unit 206 obtains from the entropy decoding unit 202 information indicating which one of one or more transform types included in the first transform type group is used by the encoding device 100 (step Su_12). Further, such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.

[0702] Based on this information, the inverse transform unit 206 determines the transform type to be used in the orthogonal transform in the encoding device 100 (step Su_13). Then, the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, when it is determined in step Su_11 that the transform size is not equal to or less than the specified value (No in step Su_11), the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the second transform type group (step Su_15).

[0703] In addition, as an example, the inverse orthogonal transform performed by the inverse transform unit 206 can be implemented for each TU according to the Figure 75 or Figure 76 shown process. Alternatively, instead of decoding the information indicating the transform type used in the orthogonal transform, an inverse orthogonal transform may be performed using a predefined transform type. Specifically, the transform type is DST7 or DCT8, etc., and in the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.

[0704] [Addition unit]

[0705] The addition unit 208 reconstructs the current block by adding the prediction residual, which is an input from the inverse transform unit 206, to the prediction image, which is an input from the prediction control unit 220. That is, a reconstructed image of the current block is generated. Then, the addition unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filtering unit 212.

[0706] [Block memory]

[0707] The block memory 210 is a storage unit for storing blocks within the current picture that are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the addition unit 208.

[0708] [Loop filtering unit]

[0709] The loop filtering unit 212 performs loop filtering on the reconstructed image generated by the addition unit 208 and outputs the filtered reconstructed image to the frame memory 214, the display device, and the like.

[0710] When the information indicating the ON / OFF of ALF read from the stream indicates that ALF is ON, one filter is selected from among a plurality of filters based on the direction and activity of the locality-based gradient, and the selected filter is applied to the reconstructed image.

[0711] Figure 77 FIG. is a block diagram showing an example of the functional configuration of the loop filter unit 212. In addition, the loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding device 100.

[0712] The loop filter unit 212, for example, as Figure 77 shown, includes a deblocking filter processing unit 212a, an SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 212b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. In addition, the loop filter unit 212 may not include Figure 77 all of the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 212 may also be configured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 77 FIG.

[0713] [Frame Memory]

[0714] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed image that has been filtered by the loop filter unit 212.

[0715] [Prediction Unit (Intra-Frame Prediction Unit / Inter-Frame Prediction Unit / Prediction Control Unit)]

[0716] Figure 78 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the decoding device 200. In addition, as an example, the prediction unit is composed of all or a part of the constituent elements of the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. The prediction processing unit includes, for example, the intra-frame prediction unit 216 and the inter-frame prediction unit 218.

[0717] The prediction unit generates a predicted image of the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Additionally, among the prediction signals, there are, for example, an intra prediction signal and an inter prediction signal. Specifically, the prediction unit generates a predicted image of the current block using a reconstructed image that has already been obtained by generating a predicted image of other blocks, restoring a prediction residual, and adding the predicted images. The prediction unit of the decoding device 200 generates the same predicted image as the predicted image generated by the prediction unit of the encoding device 100. That is, the methods for generating the predicted images used in these prediction units are mutually common or corresponding.

[0718] The reconstructed image can be, for example, an image of a reference picture, or can be an image of a decoded block within the picture containing the current block, that is, the current picture (i.e., the above-mentioned other blocks). The decoded blocks within the current picture are, for example, adjacent blocks of the current block.

[0719] Figure 79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device 200.

[0720] The prediction unit determines a method or mode for generating the predicted image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.

[0721] When it is determined that the first method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the first method (step Sr_2a). In addition, when it is determined that the second method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the second method (step Sr_2b). In addition, when it is determined that the third method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the third method (step Sr_2c).

[0722] The first method, the second method, and the third method are different methods for generating the predicted image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.

[0723] Figure 80A and Figure 80B It is a flowchart showing another example of the processing performed by the prediction unit in the decoding device 200.

[0724] As an example, the prediction unit can also perform the prediction process according to the Figure 80A and Figure 80B shown process. Additionally, Figure 80A and Figure 80BThe intra-block copy shown belongs to one mode of inter-frame prediction, and it is a mode in which a block included in the current picture is referred to as a reference picture or a reference block. That is, in intra-block copy, a picture different from the current picture is not referred to. In addition, Figure 80A The PCM mode shown belongs to one mode of intra-frame prediction, and it is a mode in which no transformation and quantization are performed.

[0725] [Intra-frame prediction unit]

[0726] The intra-frame prediction unit 216 generates a predicted image (i.e., an intra-frame predicted image) of the current block by performing intra-frame prediction with reference to the blocks within the current picture stored in the block memory 210 based on the intra-frame prediction mode read from the stream. Specifically, the intra-frame prediction unit 216 generates an intra-frame predicted image by referring to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra-frame predicted image to the prediction control unit 220.

[0727] In addition, when an intra-frame prediction mode that refers to the luminance block is selected in the intra-frame prediction of the chrominance difference block, the intra-frame prediction unit 216 may also predict the chrominance difference component of the current block based on the luminance component of the current block.

[0728] Furthermore, when the information read from the stream indicates the application of PDPC, the intra-frame prediction unit 216 corrects the pixel values after intra-frame prediction based on the gradients of the reference pixels in the horizontal / vertical directions.

[0729] Figure 81 It is a diagram showing an example of the processing performed by the intra-frame prediction unit 216 of the decoding device 200.

[0730] The intra-frame prediction unit 216 first determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that an MPM flag indicating 1 exists (Yes in step Sw_11), the intra-frame prediction unit 216 obtains the information indicating the intra-frame prediction mode selected in the encoding device 100 from the entropy decoding unit 202 (step Sw_12). In addition, this information is decoded by the entropy decoding unit 202 and output to the intra-frame prediction unit 216. Next, the intra-frame prediction unit 216 determines the MPM (step Sw_13). The MPM is composed of, for example, six intra-frame prediction modes. Then, the intra-frame prediction unit 216 determines the intra-frame prediction mode indicated by the information obtained in step Sw_12 from among the multiple intra-frame prediction modes included in this MPM (step Sw_14).

[0731] On the other hand, when it is determined in step Sw_11 that there is no MPM flag indicating 1 in the stream (No in step Sw_11), the intra prediction unit 216 acquires information indicating the intra prediction mode selected in the encoding device 100 (step Sw_15). That is, the intra prediction unit 216 acquires from the entropy decoding unit 202 information indicating the intra prediction mode selected in the encoding device 100 among one or more intra prediction modes not included in the MPM. Further, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Then, the intra prediction unit 216 determines, from among one or more intra prediction modes not included in the MPM, the intra prediction mode indicated by the information acquired in step Sw_15 (step Sw_17).

[0732] The intra prediction unit 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).

[0733] [Inter prediction unit]

[0734] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks within the current block. In addition, a sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can also be switched in units of slices, bricks, or pictures, etc.

[0735] For example, the inter prediction unit 218 performs motion compensation using motion information (e.g., MV) read from the stream (e.g., prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction image of the current block or sub-block, and outputting the inter prediction image to the prediction control unit 220.

[0736] When the information read from the stream indicates that the OBMC mode is applied, the inter prediction unit 218 generates an inter prediction image using not only the motion information of the current block obtained through motion search but also the motion information of adjacent blocks.

[0737] In addition, when the information read from the stream indicates that the FRUC mode is applied, the inter prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) read from the stream, thereby deriving motion information. And the inter prediction unit 218 uses the derived motion information for motion compensation (prediction).

[0738] In addition, when the inter-frame prediction unit 218 applies the BIO mode, it derives the MV based on a model assuming a uniform linear motion. In addition, when the information read from the stream indicates that the affine mode is applied, the inter-frame prediction unit 218 derives the MV in sub-block units based on the MVs of a plurality of adjacent blocks.

[0739] [Flow of MV derivation]

[0740] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device 200.

[0741] The inter-frame prediction unit 218 determines, for example, whether to decode motion information (e.g., MV). For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the motion information, the inter-frame prediction unit 218 derives the MV of the current block in the mode of decoding the motion information. On the other hand, when it is determined not to decode the motion information, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the motion information.

[0742] Here, the modes of MV derivation include the ordinary inter-frame mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes of decoding the motion information include the ordinary inter-frame mode, the ordinary merge mode, and the affine mode (specifically, the affine inter-frame mode and the affine merge mode), etc. In addition, the motion information can include not only the MV, but also the predicted MV selection information described later. In addition, the mode of not decoding the motion information includes the FRUC mode, etc. The inter-frame prediction unit 218 selects the mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.

[0743] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device 200.

[0744] The inter-frame prediction unit 218 determines, for example, whether to decode the differential MV. For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the differential MV, the inter-frame prediction unit 218 can derive the MV of the current block in the mode of decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.

[0745] On the other hand, when it is determined not to decode the differential MV, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the differential MV. In this case, the encoded differential MV is not included in the stream.

[0746] Here, as described above, the export modes of the MV include the following ordinary inter-frame mode, ordinary merge mode, FRUC mode, and affine mode. Among these modes, the modes for encoding the differential MV include the ordinary inter-frame mode and the affine mode (specifically, the affine inter-frame mode). In addition, the modes that do not encode the differential MV include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode). The inter-frame prediction unit 218 selects a mode for exporting the MV of the current block from these multiple modes and uses the selected mode to export the MV of the current block.

[0747] [MV Export > Ordinary Inter-Frame Mode]

[0748] For example, when the information read from the stream indicates the application of the ordinary inter-frame mode, the inter-frame prediction unit 218 exports the MV in the ordinary merge mode based on the information read from the stream and uses the MV for motion compensation (prediction).

[0749] Figure 84 It is a flowchart showing an example of inter-frame prediction by the ordinary inter-frame mode in the decoding device 200.

[0750] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially around the current block (step Sg_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0751] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the multiple candidate MVs obtained in step Sg_11 as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined priority order (step Sg_12). In addition, this priority order may also be determined in advance for each of the N prediction MV candidates.

[0752] Next, the inter-frame prediction unit 218 decodes the prediction MV selection information from the input stream and uses the decoded prediction MV selection information to select 1 prediction MV candidate from the N prediction MV candidates as the prediction MV of the current block (step Sg_13).

[0753] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction MV (step Sg_14).

[0754] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are executed for each block. For example, when the processes of steps Sg_11 to Sg_15 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. Also, when the processes of steps Sg_11 to Sg_15 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. In addition, the processes of steps Sg_11 to Sg_15 may be such that when they are not executed for all the blocks included in a slice but for some of the blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, when the processes of steps Sg_11 to Sg_15 are executed for some of the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture may end.

[0755] [MV Derivation>Normal Merge Mode]

[0756] For example, in the case where the information read from the stream indicates the application of the normal merge mode, the inter-frame prediction unit 218 derives an MV in the normal merge mode and performs motion compensation (prediction) using the MV.

[0757] Figure 85 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode in the decoding device 200.

[0758] The inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sh_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0759] Next, the inter-frame prediction unit 218 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_11 (step Sh_12). Specifically, the inter-frame prediction unit 218 obtains, for example, MV selection information included as a prediction parameter in the stream and selects the candidate MV identified by the MV selection information as the MV of the current block.

[0760] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processes of steps Sh_11 to Sh_13 are performed on each block, for example. For example, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a slice respectively, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a picture respectively, the inter-frame prediction using the normal merge mode for that picture ends. In addition, the processes of steps Sh_11 to Sh_13 may be such that when they are not performed on all the blocks included in a slice but on a part of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processes of steps Sh_11 to Sh_13 may be such that when they are performed on a part of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.

[0761] [MV Derivation > FRUC Mode]

[0762] For example, in the case where the information read from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding device 100 side but is derived on the decoding device 200 side. For example, the decoding device 200 may also derive the motion information by performing a motion search. In this case, the decoding device 200 does not use the pixel values of the current block for the motion search.

[0763] Figure 86 It is a flowchart showing an example of inter-frame prediction based on the FRUC mode in the decoding device 200.

[0764] First, the inter-frame prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block, and generates a list representing these MVs as candidate MVs (i.e., it is a candidate MV list and can also be common with the candidate MV list in the normal merge mode) (step Si_11). Next, the inter-frame prediction unit 218 selects the best candidate MV from among the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter-frame prediction unit 218 calculates the evaluation value of each candidate MV included in the candidate MV list, and selects one candidate MV as the best candidate MV based on this evaluation value. Then, the inter-frame prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Additionally, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral area of the position in the reference picture corresponding to the selected best candidate MV. That is, for the area around the best candidate MV, a search using pattern matching and evaluation values in the reference picture is performed. Furthermore, in the case where there is an MV for which the evaluation value becomes a good value, the best candidate MV may also be updated to this MV and used as the final MV for the current block. It is also possible not to perform the update to an MV with a better evaluation value.

[0765] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processing of steps Si_11 to Si_15 is performed for each block, for example. For example, when the processing of steps Si_11 to Si_15 is performed separately for all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. Additionally, when the processing of steps Si_11 to Si_15 is performed separately for all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. It is also possible to perform the processing in the same manner as the above block unit in units of sub-blocks.

[0766] [MV Derivation > Affine Merge Mode]

[0767] For example, in the case where the information read from the stream indicates the application of the affine merge mode, the inter-frame prediction unit 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using this MV.

[0768] Figure 87 It is a flowchart showing an example of inter-frame prediction based on the affine merge mode in the decoding device 200.

[0769] In the affine merge mode, the inter-frame prediction unit 218 first derives the MV for each control point of the current block (step Sk_11). As Figure 46AAs shown, the control points are the points at the upper left and upper right corners of the current block, or as Figure 46B shown, are the points at the upper left, upper right, and lower left corners of the current block.

[0770] For example, in the case of using the Figures 47A - 47C MV export method shown, as Figure 47A shown, the inter-frame prediction unit 218 checks these blocks in the order of the decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and determines the first valid block decoded in the affine mode.

[0771] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to export the MV of the control points. For example, in the case where block A is determined and block A has two control points, as Figure 47B shown, the inter-frame prediction unit 218 projects the motion vectors v3 and v4 of the upper left and upper right corners of the decoded block including block A onto the current block to calculate the motion vector v0 of the upper left control point of the current block and the motion vector v1 of the upper right control point. Thus, the MV of each control point is exported.

[0772] In addition, as Figure 49A shown, in the case where block A is determined and block A has two control points, it is also possible to calculate the MV of three control points, or as Figure 49B shown, determine block A, and in the case where block A has three control points, calculate the MV of two control points.

[0773] In addition, when the stream includes MV selection information as a prediction parameter, the inter-frame prediction unit 218 can also use this MV selection information to export the MV of each control point of the current block.

[0774] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 218 calculates the MV of the sub-block as an affine MV for each of the multiple sub-blocks using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_12). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine MVs and the decoded reference picture (step Sk_13). When the processes of steps Sk_12 and Sk_13 are respectively performed on all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0775] In addition, in step Sk_11, the above-mentioned candidate MV list may also be generated. The candidate MV list may also be, for example, a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods may be Figures 47A - 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in, and any combination of other MV derivation methods.

[0776] In addition, the candidate MV list may also include candidate MVs of a prediction mode that is predicted in units of sub-blocks other than the affine mode.

[0777] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and candidate MVs of an affine merge mode with 3 control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and a candidate MV list including candidate MVs of an affine merge mode with 3 control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes with 2 control points and the affine merge mode with 3 control points may be generated.

[0778] [MV Derivation > Affine Inter-Frame Mode]

[0779] For example, when the information read from the stream indicates the application of the affine inter-frame mode, the inter-frame prediction unit 218 derives an MV in the affine inter-frame mode and performs motion compensation (prediction) using the MV.

[0780] Figure 88 is a flowchart showing an example of inter-frame prediction based on the affine inter-frame mode in the decoding device 200.

[0781] In the affine inter-frame mode, first, the inter-frame prediction unit 218 derives the predicted MVs (v0, v1) or (v0, v1, v2) of each of the 2 or 3 control points of the current block (step Sj_11). The control points are, for example, as Figure 46A or Figure 46B shown, the points at the upper left corner, upper right corner, or lower left corner of the current block.

[0782] The inter-frame prediction unit 218 obtains the predicted MV selection information included as prediction parameters in the stream, and uses the MV identified by the predicted MV selection information to derive the predicted MVs of the respective control points of the current block. For example, when using Figure 48A and Figure 48B the MV derivation methods shown in, the inter-frame prediction unit 218 selects Figure 48Aor Figure 48B The motion vectors (MVs) of the blocks in the decoded blocks near the respective control points of the current block shown, which are identified by the prediction MV selection information, are used to derive the predicted MVs (v0, v1) or (v0, v1, v2) of the control points of the current block.

[0783] Next, the inter-frame prediction unit 218, for example, obtains each differential MV included as a prediction parameter in the stream, and adds the predicted MV of each control point of the current block and the differential MV corresponding to the predicted MV (step Sj_12). Thereby, the MVs of each control point of the current block are derived.

[0784] Next, the inter-frame prediction unit 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. That is, the inter-frame prediction unit 218 calculates the MV of each of the plurality of sub-blocks as an affine MV using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sj_13). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine MVs and the decoded reference picture (step Sj_14). When the processes of steps Sj_13 and Sj_14 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.

[0785] In addition, in step Sj_11, the above-mentioned candidate MV list may be generated in the same manner as in step Sk_11.

[0786] [MV Derivation > Triangular Mode]

[0787] For example, in the case where the information read from the stream indicates the application of the triangular mode, the inter-frame prediction unit 218 derives the MV in the triangular mode and uses the MV for motion compensation (prediction).

[0788] Figure 89 is a flowchart showing an example of inter-frame prediction based on the triangular mode in the decoding device 200.

[0789] In the triangular mode, first, the inter-frame prediction unit 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter-frame prediction unit 218 may obtain, from the stream, partition information, which is information related to the division into each partition, as a prediction parameter. Moreover, the inter-frame prediction unit 218 may divide the current block into a first partition and a second partition according to the partition information.

[0790] Next, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). That is, the inter-frame prediction unit 218 creates a candidate MV list.

[0791] Then, the inter-frame prediction unit 218 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter-frame prediction unit 218 may also obtain MV selection information for identifying the selected candidate MVs from the stream as a prediction parameter. Then, the inter-frame prediction unit 218 may select the first MV and the second MV according to the MV selection information.

[0792] Next, the inter-frame prediction unit 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame prediction unit 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).

[0793] Finally, the inter-frame prediction unit 218 generates a prediction image of the current block by performing weighted addition on the first prediction image and the second prediction image (step Sx_16).

[0794] [Motion Search > DMVR]

[0795] For example, in the case where the information read from the stream indicates the application of DMVR, the inter-frame prediction unit 218 performs motion search by DMVR.

[0796] Figure 90 It is a flowchart showing an example of motion search based on DMVR in the decoding device 200.

[0797] The inter-frame prediction unit 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter-frame prediction unit 218 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is determined by DMVR.

[0798] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device 200.

[0799] First, the inter-frame prediction unit 218 is in Figure 58AIn Step 1 shown above, the search positions (also referred to as start points) of the initial MV representation and the costs of 8 search positions around it are calculated. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, when it is determined that the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 moves to the search position with the minimum cost and performs Figure 58A the processing of Step 2 shown above. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 skips Figure 58A the processing of Step 2 shown above and performs the processing of Step 3.

[0800] In Figure 58A Step 2 shown above, the inter-frame prediction unit 218 uses the search position moved according to the processing result of Step 1 as a new start point and performs the same search as the processing of Step 1. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, if the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 performs the processing of Step 4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 performs the processing of Step 3.

[0801] In Step 4, the inter-frame prediction unit 218 treats the search position of the start point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0802] In Figure 58A Step 3 shown above, the inter-frame prediction unit 218 determines the pixel position with the minimum cost at the fractional precision based on the costs of 4 points above, below, left, and right of the start point in Step 1 or Step 2, and uses this pixel position as the final search position. This pixel position at the fractional precision is determined by weighted addition of the vectors (0, 1), (0, -1), (-1, 0), (1, 0) at the 4 points above, below, left, and right with the costs of the search positions of the 4 points as weights. Then, the inter-frame prediction unit 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0803] [Motion Compensation > BIO / OBMC / LIC]

[0804] For example, in the case where the information read from the stream represents the application of correction of the predicted image, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to the correction mode. This mode is, for example, the above-mentioned BIO, OBMC, and LIC.

[0805] Figure 92 is a flowchart showing an example of generation of a predicted image in the decoding device 200.

[0806] The inter-frame prediction unit 218 generates a prediction image (step Sm_11), and corrects the prediction image by any of the above-described modes (step Sm_12).

[0807] Figure 93 It is a flowchart showing another example of the generation of a prediction image in the decoding apparatus 200.

[0808] The inter-frame prediction unit 218 derives the MV of the current block (step Sn_11). Next, the inter-frame prediction unit 218 generates a prediction image using the MV (step Sn_12), and determines whether to perform a correction process (step Sn_13). For example, the inter-frame prediction unit 218 obtains the prediction parameters included in the stream, and determines whether to perform a correction process based on the prediction parameters. The prediction parameter is, for example, a flag indicating whether to apply any of the above-described modes. Here, when it is determined to perform a correction process (Yes in step Sn_13), the inter-frame prediction unit 218 generates a final prediction image by correcting the prediction image (step Sn_14). In addition, in the LIC, the luminance and color difference of the prediction image can be corrected in step Sn_14. On the other hand, when it is determined not to perform a correction process (No in step Sn_13), the inter-frame prediction unit 218 outputs the prediction image without correcting it as the final prediction image (step Sn_15).

[0809] [Motion Compensation>OBMC]

[0810] For example, when the information read from the stream indicates the application of OBMC, when generating a prediction image, the inter-frame prediction unit 218 corrects the prediction image according to OBMC.

[0811] Figure 94 It is a flowchart showing an example of the correction of a prediction image based on OBMC in the decoding apparatus 200. In addition, Figure 94 The flowchart of Figure 62 shows the process of correcting a prediction image using the current picture and the reference picture shown in

[0812] First, as Figure 62 shown, the inter-frame prediction unit 218 obtains a prediction image (Pred) based on normal motion compensation using the MV assigned to the current block.

[0813] Next, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_L) that has been derived for the decoded left adjacent block to the current block, and obtains a prediction image (Pred_L). Then, the inter-frame prediction unit 218 performs the first correction of the prediction image by overlapping the two prediction images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0814] Similarly, the inter-frame prediction unit 218 applies (re-uses) the motion vector (MV_U) that has been derived for the decoded upper adjacent block to the current block to obtain a predicted image (Pred_U). Then, the inter-frame prediction unit 218 performs a second correction of the predicted image by overlapping the predicted image Pred_U with the predicted images that have undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block that is blended (smoothed) with the boundaries of adjacent blocks.

[0815] [Motion Compensation > BIO]

[0816] For example, in the case where the information read from the stream indicates the application of BIO, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to BIO.

[0817] Figure 95 It is a flowchart showing an example of the correction of the predicted image based on BIO in the decoding device 200.

[0818] As Figure 63 shown, the inter-frame prediction unit 218 uses two reference pictures (Ref0, Ref1) different from the picture (CurPic) containing the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 218 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_11). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.

[0819] Next, the inter-frame prediction unit 218 uses the motion vector M0 and the reference picture L0 to derive the interpolation image I 0 for the current block. Additionally, the inter-frame prediction unit 218 uses the motion vector M1 and the reference picture L1 to derive the interpolation image I 1 (step Sy_12) for the current block. Here, the interpolation image I 0 is the image contained in the reference picture Ref0 that is derived for the current block, and the interpolation image I 1 is the image contained in the reference picture Ref1 that is derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Or, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each also be an image larger than the current block. Additionally, the interpolation image I 0 and I 1It may include a predicted image derived by applying motion vectors (M0, M1) and reference pictures (L0, L1), and a motion compensation filter.

[0820] In addition, the inter-frame prediction unit 218 derives the gradient image (Ix 0 and Ix 1 of the current block from the interpolation image I 0 and Ix 1 Iy 0 and Iy 1 )(step Sy_13). In addition, the gradient image in the horizontal direction is (Ix 0 and Ix 1 ), and the gradient image in the vertical direction is (Ix 0 and Ix 1 ). The inter-frame prediction unit 218 may also derive the gradient image by applying a gradient filter to the interpolation image, for example. The gradient image may be an image representing the spatial variation amount of pixel values along the horizontal direction or the vertical direction.

[0821] Next, the inter-frame prediction unit 218 uses the interpolation image (I 0 and I 1 ) and the gradient image (Ix 0 and Ix 1 Iy 0 and Iy 1 ) to derive the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block (step Sy_14). As an example, the sub-block may be a 4x4 pixel sub-CU.

[0822] Next, the inter-frame prediction unit 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame prediction unit 218 derives a correction value of the pixel values included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter-frame prediction unit 218 may correct the predicted image of the current block using the correction value (step Sy_16). In addition, the correction value may be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.

[0823] In addition, the processing flow of BIO is not limited to Figure 95 the disclosed processing. It may perform only a part of the disclosed processing, or may add or replace different processing, or may execute in a different processing order. Figure 95 the disclosed processing.

[0824] [Motion Compensation > LIC]

[0825] For example, in a case where the information read from the stream indicates the application of LIC, when generating a predicted image, the inter-frame prediction unit 218 corrects the predicted image according to LIC.

[0826] Figure 96 It is a flowchart showing an example of correction of a predicted image based on LIC in the decoding device 200.

[0827] First, the inter-frame prediction unit 218 uses the MV to obtain a reference image corresponding to the current block from the decoded reference pictures (step Sz_11).

[0828] Next, the inter-frame prediction unit 218 extracts information indicating how the luminance values change in the reference picture and the current picture for the current block (step Sz_12). As Figure 66A shown, this extraction is based on the luminance pixel values of the decoded left adjacent reference region (peripheral reference region) and the decoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. Then, the inter-frame prediction unit 218 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_13).

[0829] The inter-frame prediction unit 218 generates a predicted image for the current block by performing a luminance correction process of applying its luminance correction parameter to the reference image in the reference picture specified by the MV (step Sz_14). That is, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance or the chromatic aberration can be corrected.

[0830] [Prediction control unit]

[0831] The prediction control unit 220 selects one of the intra-frame predicted image and the inter-frame predicted image, and outputs the selected predicted image to the adder 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device 200 side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device 100 side.

[0832] [First form]

[0833] Next, the encoding device 100 and the decoding device 200 of the first form will be described.

[0834] The encoding device 100 obtains a first parameter, saves the first parameter indicating whether a syntax element associated with chroma tool offset exists in the bitstream in the bitstream, and when the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, saves one or more second parameters used in the deblocking filter processing of the chroma samples of the image to be processed in the bitstream, (ii) performs deblocking filter processing using one or more second parameters, and (iii) encodes the image to be processed. When the first parameter indicates that the syntax element does not exist in the bitstream, the image to be processed is encoded without saving one or more second parameters in the bitstream.

[0835] In addition, the decoding device 200 obtains a first parameter, where the first parameter indicates whether a syntax element associated with chroma tool offset exists in the bitstream. When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, reads one or more second parameters used in the deblocking filter processing of the chroma samples of the image to be processed from the bitstream, (ii) performs deblocking filter processing using one or more second parameters, and (iii) decodes the image to be processed. When the first parameter indicates that the syntax element does not exist in the bitstream, the image to be processed is decoded without reading one or more second parameters from the bitstream.

[0836] Hereinafter, the decoding device 200 will be described.

[0837] Figure 97 FIG. 1000 is a flowchart showing an example of the internal configuration (so-called internal setting) of the loop filter of the decoding device 200 in the first form. Figure 98 FIG. is an example showing the position of the first parameter in the bitstream. Figure 99 FIG. is an example showing the position of the second parameter in the bitstream. Figure 100 FIG. is an example showing the syntax structure in the PPS (Picture Parameter Set).

[0838] As Figure 97 shown, in step S101, the decoding device 200 derives the first parameter. Here, "derives" means reads, refers to, or obtains.

[0839] As Figure 98 shown, the first parameter may be included in the VPS (Video Parameter Set) in the bitstream (refer to Figure 98 (a)), may be included in the SPS (Sequence Parameter Set) (refer to Figure 98 (b)), may be included in the PPS (Picture Parameter Set) (refer to Figure 98 (c)). In addition, the first parameter may also be included in the PH (Picture Header) (refer to Figure 98(d)) may also be included in the SH (slice header) (see Figure 98 (e)). At this time, the first parameter may also be directly read from the VPS, SPS, PPS, PH, or SH in the bitstream. As Figure 100 shown, as an example, the first parameter may be pps_chroma_tool_offsets_present_flag read from the PPS. Here, pps_chroma_tool_offsets_present_flag specifies that the syntax element associated with the chroma tool offset exists in the PPS.

[0840] The chroma tool offset is a set of offsets for controlling the quantization parameters of the chroma components. Specifically, the chroma tool offset specifies the offset of the luminance component relative to the quantization parameter. The syntax element associated with the chroma tool offset is the syntax part corresponding to the set of offsets that controls the quantization parameters of the chroma components in the bitstream. In the first form, pps_chroma_tool_offsets_present_flag is used not only to determine whether to read the syntax element associated with the chroma tool offset from the bitstream, but also to determine whether to read one or more second parameters used in the deblocking filter process from the bitstream. As another example, the first parameter may be chroma_format_idc read from the SPS. Here, chroma_format_idc specifies the chroma sampling associated with the luminance sampling.

[0841] In addition, the first parameter may also be derived based on the parameters read from the VPS, SPS, PPS, PH, or SH in the bitstream. In this case, the first parameter is derived based on the chroma format. The chroma format includes, for example, a monochrome format (black and white) that specifies an image composed only of the Y component (luminance component), a format composed of the Y component (luminance component) and two chroma components (e.g., Cb component and Cr component) (hereinafter, the format composed of Y / Cb / Cr), and a format composed of the R component (red component), G component (green component), and B component (blue component) (hereinafter, the format composed of R / G / B).

[0842] The parameters indicating the chroma format can be, for example, separate_colour_plane_flag, chroma_format_idc, and ChromaArrayType. For example, separate_colour_plane_flag can distinguish whether the chroma format is a format composed of Y / Cb / Cr or a format composed of R / G / B according to whether its value is 0 or 1. In addition, for example, chroma_format_idc can distinguish whether it is a monochrome format, 4:2:0 format, 4:2:2 format, or 4:4:4 format according to its value. In addition, for example, ChromaArrayType basically has the same meaning as chroma_format_idc, but the difference is that when representing the 4:4:4 (RGB) format, its value is set to 0. Therefore, ChromaArrayType can indicate that three color images (i.e., R image, G image, and B image) composed of respective color components of R / G / B are each subjected to the same processing as one monochrome (black and white) image.

[0843] As an example, the first parameter is derived based on at least one of chroma_format_idc and separate_colour_plane_flag read from the SPS. Here, chroma_format_idc specifies the chroma sampling associated with the luma sampling, and separate_colour_plane_flag specifies whether to encode the three color components (e.g., R component, G component, and B component) of the 4:4:4 chroma format separately. For example, when separate_colour_plane_flag is equal to 0, the first parameter is set to be equal to chroma_format_idc. In other cases, the first parameter is set to 0. In this example, as Figure 101 and Figure 102 shown, the first parameter is ChromaArrayType. Figure 101 is a diagram showing an example of the syntax structure in PH. Figure 102 is a diagram showing an example of the syntax structure in SH.

[0844] Refer again to Figure 97, in step S102, the decoding device 200 determines whether the chrominance of the image to be processed is encoded based on the first parameter derived in step S101. For example, the decoding device 200 determines whether the chrominance of the image to be processed is encoded based on the value of the first parameter. In this case, the decoding device 200 determines whether the chrominance of the image to be processed is encoded based on whether the first parameter is a predetermined value (so-called first value). As an example, the decoding device 200 determines that the chrominance of the image to be processed is encoded when the first parameter is not a predetermined value (for example, the first value = 0). As another example, the decoding device 200 may also determine that the chrominance of the image to be processed is encoded when the first parameter is greater than a predetermined value (for example, the first value = 0). Here, chrominance represents the color component. As an example, chrominance represents the Cb component and the Cr component in the YCbCr color space. As another example, chrominance may represent other components in a color space other than YCbCr.

[0845] In addition, as Figure 100 shown in the example, when the first parameter is pps_chroma_tool_offsets_present_flag, the determination in step S102 can be replaced by a determination of whether the syntax element associated with the chrominance tool offset is encoded. In other words, when pps_chroma_tool_offsets_present_flag is used as the first parameter, the determination in step S102 is replaced by a determination of whether to perform processing using the syntax element associated with the chrominance tool offset.

[0846] Next, when the decoding device 200 uses the first parameter derived in step S101 and determines that the chrominance of the image to be processed is encoded (yes in step S102), it reads one or more second parameters for encoding the chrominance of the image to be processed from the bitstream (step S103). Then, the decoding device 200 decodes the image to be processed using one or more second parameters read in step S103 (step S104).

[0847] As Figure 99 shown, one or more second parameters may be included in the SPS in the bitstream (refer to Figure 99 (a)), may be included in the PPS (refer to Figure 99 (b)), may be included in the PH (refer to Figure 99 (c)), may be included in the SH (refer to Figure 99 (d)), or may also be included in the CTU (Coding Tree Unit) (refer to Figure 99(e)). At this time, one or more second parameters can be read from the SPS, PPS, PH, SH, or CTU in the bitstream.

[0848] One or more second parameters can be used in the deblocking filter process. Here, one or more second parameters represent the deblocking parameter offset of beta (β) and / or tC of the color component. The color component is the so-called chrominance, for example, representing the Cb component and the Cr component in the YCbCr color space. β and tC are parameters for controlling the change amount of the values of the chrominance samples in the deblocking filter process. β and tC are determined by the quantization parameter Q. β and tC can be used as thresholds in the determination of whether to perform the block filtering process and / or the intensity of the deblocking filter process. The determination of the intensity of the deblocking filter process refers to the determination of the type of filter to be applied. The intensity of the filtering process indicates the number of target pixels to be changed and / or the number of taps of the filter in the deblocking filter process. For example, the stronger the intensity of the filter, the more target pixels whose pixel values are to be changed, and the longer the filter tap length.

[0849] For example, as Figure 100 shown, one or more second parameters can represent the deblocking parameter offset of β and tC of the color components (for example, the Cb component and the Cr component) in the PPS. In this case, the deblocking parameter offset of β of the Cb component is pps_cb_beta_offset_div2, the deblocking parameter offset of tC of the Cb component is pps_cb_tc_offset_div2, the deblocking parameter offset of β of the Cr component is pps_cr_beta_offset_div2, and the deblocking parameter offset of tC of the Cr component is pps_cr_tc_offset_div2. pps_cb_beta_offset_div2 and pps_cb_tc_offset_div2 show the values (for example, the values divided by 2) representing the default deblocking parameter offsets of β and tC for the Cb component to be applied to the slices referring to this PPS. pps_cr_beta_offset_div2 and pps_cr_tc_offset_div2 show the values (for example, the values divided by 2) representing the default deblocking parameter offsets of β and tC for the Cr component to be applied to the slices referring to this PPS. That is, in the deblocking filter process, the values obtained by adding the offsets determined by the PPS to β and tC determined by the quantization parameter Q are used. β and tC can also be values further added with the offsets determined in units different from the PPS such as slices.

[0850] In addition, in the figure, if(pps_chroma_tool_offsets_present_flag) can also be replaced with if(pps_chroma_tool_offsets_present_flag > 0) or if(pps_chroma_tool_offsets_present_flag != 0).

[0851] In addition, for example, as Figure 101 shown, one or more second parameters can represent the deblocking parameter offsets of β and tC for color components (e.g., Cb component and Cr component) in PH. In this case, the deblocking parameter offset of β for the Cb component is ph_cb_beta_offset_div2, the deblocking parameter offset of tC for the Cb component is ph_cb_tc_offset_div2, the deblocking parameter offset of β for the Cr component is ph_cr_beta_offset_div2, and the deblocking parameter offset of tC for the Cr component is ph_cr_tc_offset_div2. In addition, if(ChromaArrayType != 0) in the figure can be replaced with if(ChromaArrayType > 0) or if(ChromaArrayType < 0).

[0852] In addition, for example, as Figure 102 shown, one or more second parameters can represent the deblocking parameter offsets of β and tC for color components (e.g., Cb component and Cr component) in SH. In this case, the deblocking parameter offset of β for the Cb component is slice_cb_beta_offset_div2, the deblocking parameter offset of tC for the Cb component is slice_cb_tc_offset_div2, the deblocking parameter offset of β for the Cr component is slice_cr_beta_offset_div2, and the deblocking parameter offset of tC for the Cr component is slice_cr_tc_offset_div2. In addition, if(ChromaArrayType != 0) in the figure can be replaced with if(ChromaArrayType > 0) or if(ChromaArrayType < 0).

[0853] In addition, in Figure 100 、 Figure 101 and Figure 102 examples, different structures are used as the first parameter respectively, but the same structure can also be used for all. For example, in Figure 101 and 102 with Figure 100Similarly, the first parameter can also be set to a structure using the pps_chroma_tool_offsets_present_flag.

[0854] Refer again to Figure 97 , in step S102, when the decoding device 200 determines that the first parameter derived in step S101 is used and the chrominance of the image to be processed is not encoded (No), it does not read one or more second parameters from the bitstream and decodes the image to be processed without using one or more second parameters (step S105). Here, when the image to be processed consists of one component in a separate color plane format, the chrominance of the image to be processed can also be determined as not being encoded.

[0855] [Effect of the first form]

[0856] In the structure of the first form, the encoding device 100 can reduce the processing amount while reducing the size of the bitstream by sharing the first parameter between the process using the syntax element associated with the chrominance tool offset and the deblocking filter process. The decoding device 200 can reduce the process related to the determination of whether one or more second parameters should be read and reduce the storage amount by sharing the first parameter between the process using the syntax element associated with the chrominance tool offset and the deblocking filter process. For example, sharing the first parameter is useful when encoding / decoding an image that is a monochromatic image and has no chrominance component.

[0857] [Combination with other forms]

[0858] This form can be implemented in combination with at least a part of other forms in the present disclosure. In addition, this form can also be implemented by combining a part of the processes shown in the flowcharts of any of the other forms, a part of the structure of any of the devices, a part of the syntax, etc.

[0859] In addition, the above-mentioned processes in the components of the decoding device can also be executed in the same way in the components of the encoding device.

[0860] In addition, not all of the components described in this form are necessarily required, and only a part of the components of the first form can be included.

[0861] [Representative example of processing]

[0862] The following shows representative examples of the structures and processes of the encoding device 100 and the decoding device 200 shown above.

[0863] Figure 103It is a flowchart showing the operations performed by the encoding device 100. For example, the encoding device 100 includes a circuit and a memory connected to the circuit. The circuit and the memory included in the encoding device 100 may also correspond to Figure 8 the processor a1 and the memory a2 shown. During operation, the circuit of the encoding device 100 performs the following operations.

[0864] For example, the circuit of the encoding device 100 stores a first parameter in the bitstream, where the first parameter indicates whether a syntax element associated with chroma tool offset exists in the bitstream (step S201), and determines whether the first parameter indicates that the syntax element exists in the bitstream (step S202). In step S202, when the first parameter indicates that the syntax element exists in the bitstream (yes), the circuit of the encoding device 100 (i) stores one or more second parameters used in the deblocking filter processing of the chroma samples of the processing target image in the bitstream in addition to the syntax element (step S203), (ii) performs deblocking filter processing using one or more second parameters, and (iii) encodes the processing target image (step S204). On the other hand, in step S202, when the first parameter indicates that the syntax element does not exist in the bitstream (no), the circuit of the encoding device 100 encodes the processing target image without storing one or more second parameters in the bitstream (step S205).

[0865] Thereby, by jointly using the first parameter between the processing using the syntax element associated with chroma tool offset and the deblocking filter processing, the encoding device 100 can reduce the size of the bitstream and reduce the processing amount at the same time.

[0866] For example, the first parameter may be pps_chroma_tool_offsets_present_flag set in the picture parameter set.

[0867] Thereby, the circuit of the encoding device 100 can, for example, determine whether one or more second parameters exist in the bitstream without parsing a slice header more than the picture parameter set.

[0868] For example, the chroma samples may represent the Cb component and the Cr component in the YCbCr color space.

[0869] Thereby, since the processing target image has information on luminance (Y component) and two chroma components (Cb component and Cr component), the circuit of the encoding device 100 can reduce the information related to chroma while suppressing a reduction in subjective image quality. Therefore, since the amount of information included in the bitstream can be reduced, the encoding amount can be reduced.

[0870] For example, it may also be that one or more second parameters are used to control the amount of change in the values of the chrominance samples in the deblocking process.

[0871] Thus, the circuit of the encoding device 100 can use one or more second parameters to limit the amount of change in the values of the chrominance samples in the deblocking process.

[0872] For example, it may also be that one or more second parameters respectively represent an offset for at least one of β and tC which are parameters for controlling the amount of change.

[0873] Thus, the circuit of the encoding device 100 can control the deblocking process of the chrominance of the image to be processed based on the values of these deblocking parameter offsets.

[0874] For example, it may also be that when the first parameter indicates that a syntax element exists in the bitstream, the circuit performs processing using the syntax element. In this case, the syntax element may be a quantization offset for controlling the quantization parameter of the chrominance component, and the circuit of the encoding device 100 can use the quantization offset to perform quantization processing.

[0875] Thus, when the syntax element exists in the bitstream, the circuit of the encoding device 100 can use the quantization offset to perform quantization processing.

[0876] Figure 104 It is a flowchart showing the operations performed by the decoding device 200. For example, the decoding device 200 includes a circuit and a memory connected to the circuit. The circuit and the memory included in the decoding device 200 may also correspond to Figure 68 the shown processor b1 and memory b2. During operation, the circuit of the decoding device 200 performs the following operations.

[0877] For example, the circuit of the decoding device 200 obtains a first parameter that indicates whether a syntax element associated with a chrominance tool offset exists in the bitstream (step S301), and determines whether the first parameter indicates that the syntax element exists in the bitstream (step S302). In step S302, when the first parameter indicates that the syntax element exists in the bitstream (Yes), the circuit of the decoding device 200 (i) reads out one or more second parameters used in the deblocking process of the chrominance samples of the image to be processed from the bitstream in addition to the syntax element (step S303), (ii) performs deblocking processing using one or more second parameters, and (iii) decodes the image to be processed (step S304). On the other hand, in step S302, when the first parameter indicates that the syntax element does not exist in the bitstream (No), the circuit of the decoding device 200 decodes the image to be processed without reading out one or more second parameters from the bitstream (step S305).

[0878] Accordingly, by sharing the first parameter between the process using the syntax element associated with the chroma tool offset and the deblocking filter process, the decoding device 200 can reduce the process related to the determination of whether one or more second parameters should be read out, and reduce the storage amount.

[0879] For example, the first parameter may be pps_chroma_tool_offsets_present_flag set in the picture parameter set.

[0880] Accordingly, the circuit of the decoding device 200 c...

Claims

1. An encoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, when the circuit is operating, saving a first parameter in a bitstream, the first parameter indicating whether a syntax element associated with a chroma tool offset exists in the bitstream, when the first parameter indicates that the syntax element exists in the bitstream, (i) saving, in addition to the syntax element, one or more second parameters used in deblocking filtering processing of chroma samples of a processing target image in the bitstream, (ii) performing the deblocking filtering processing using the one or more second parameters, and (iii) encoding the processing target image, when the first parameter indicates that the syntax element does not exist in the bitstream, encoding the processing target image without saving the one or more second parameters in the bitstream.

2. The encoding apparatus according to claim 1, wherein the first parameter is pps_chroma_tool_offsets_present_flag set in a picture parameter set.

3. The encoding apparatus according to claim 1 or 2, wherein the chroma samples represent the Cb component and the Cr component in the YCbCr color space.

4. The encoding apparatus according to claim 1 or 2, wherein the one or more second parameters are used to control a change amount of values of the chroma samples based on the deblocking filtering processing.

5. The encoding apparatus according to claim 4, wherein the one or more second parameters respectively represent an offset for at least one of β and tC which are parameters for controlling the change amount.

6. The encoding apparatus according to claim 1 or 2, wherein when the first parameter indicates that the syntax element exists in the bitstream, the circuit performs processing using the syntax element.

7. The encoding apparatus according to claim 6, wherein the syntax element is a quantization offset for controlling a quantization parameter of a chroma component, and the circuit performs quantization processing using the quantization offset.

8. A decoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, when the circuit is operating, obtaining a first parameter, the first parameter indicating whether a syntax element associated with a chroma tool offset exists in a bitstream, when the first parameter indicates that the syntax element exists in the bitstream, (i) reading one or more second parameters used in deblocking filtering processing of chroma samples of a processing target image from the bitstream in addition to the syntax element, (ii) performing deblocking filtering processing using the one or more second parameters, and (iii) decoding the processing target image, when the first parameter indicates that the syntax element does not exist in the bitstream, decoding the processing target image without reading the one or more second parameters from the bitstream.

9. The decoding apparatus according to claim 8, wherein The first parameter is pps_chroma_tool_offsets_present_flag set in the picture parameter set.

10. The decoding apparatus according to claim 8 or 9, wherein The chroma samples represent the Cb and Cr components in the YCbCr color space.

11. The decoding apparatus according to claim 8 or 9, wherein The one or more second parameters are used to control the amount of change in the values of the chroma samples based on the deblocking filtering process.

12. The decoding apparatus according to claim 11, wherein The one or more second parameters respectively represent an offset for at least one of β and tC which are parameters for controlling the amount of change.

13. The decoding apparatus according to claim 8 or 9, wherein When the first parameter indicates that the syntax element exists in the bitstream, the circuit performs a process using the syntax element.

14. The decoding apparatus according to claim 13, wherein The syntax element is a quantization offset for controlling the quantization parameter of the chroma component, The circuit performs inverse quantization processing using the quantization offset.

15. An encoding method, wherein A first parameter is saved in the bitstream, the first parameter indicating whether a syntax element associated with chroma tool offset exists in the bitstream, When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking filtering process of the chroma samples of the image to be processed are saved in the bitstream, (ii) the deblocking filtering process is performed using the one or more second parameters, and (iii) the image to be processed is encoded, When the first parameter indicates that the syntax element does not exist in the bitstream, the image to be processed is encoded without saving the one or more second parameters in the bitstream.

16. A decoding method, wherein A first parameter is obtained, the first parameter indicating whether a syntax element associated with chroma tool offset exists in the bitstream, When the first parameter indicates that the syntax element exists in the bitstream, (i) in addition to the syntax element, one or more second parameters used in the deblocking filtering process of the chroma samples of the image to be processed are read from the bitstream, (ii) the deblocking filtering process is performed using the one or more second parameters, and (iii) the image to be processed is decoded, When the first parameter indicates that the syntax element does not exist in the bitstream, the image to be processed is decoded without reading the one or more second parameters from the bitstream.

Citation Information

Patent Citations

  • Image decoding method, image encoding method, image decoding device, image encoding device, and image encoding / decoding device

    CN103583048A

  • Encoder-side decisions for sample adaptive offset filtering

    CN105409221A